A document translation method, apparatus, device, and storage medium

By obtaining the document attribute information of the source language document before translation, the trade-off between document translation effect and resource consumption is solved, improving the translation effect and reducing resource consumption.

CN115409045BActive Publication Date: 2026-05-05IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-08-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing document translation solutions fail to strike a balance between translation quality and resource consumption. Directly incorporating contextual sentences leads to enormous resource consumption, while restricting the context to one or a few sentences before and after results in a decline in translation quality.

Method used

Before translating, obtain document attribute information from the source language document, such as tense, ambiguity, domain, and coreference attributes. Use this information to assist in the translation, avoiding the direct introduction of contextual sentences and only introducing key and useful information.

Benefits of technology

It achieves a balance between document translation quality and resource consumption, reducing resource consumption while improving translation quality and avoiding the negative impact of contextual noise on translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115409045B_ABST
    Figure CN115409045B_ABST
Patent Text Reader

Abstract

This invention provides a document translation method, apparatus, device, and storage medium. The document translation method includes: acquiring a source language document to be translated; acquiring document attribute information corresponding to each sentence in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or its context sentences to assist in determining the meaning of words in the corresponding sentence within the source language document; and translating each sentence in the source language document using the document attribute information corresponding to each sentence, thereby obtaining a target language translation of the source language document. The document translation method provided by this invention has good document translation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of translation technology, and in particular to a document translation method, apparatus, device, and storage medium. Background Technology

[0002] Document translation is the process of converting a document in one natural language (source language) into a document in another natural language (target language). To achieve better translation results, contextual sentences are often incorporated into the translation process.

[0003] To significantly improve translation quality, some current document translation solutions introduce rich context (such as numerous surrounding sentences of the sentence to be translated). However, introducing rich context leads to huge resource consumption. To reduce resource consumption, some current document translation solutions constrain the context to one or a few sentences before and after the sentence to be translated. While this strategy reduces resource consumption, it leads to a decrease in document translation quality.

[0004] It is evident that current document translation solutions fail to strike a balance between translation effectiveness and resource consumption, and this inability to balance these factors prevents document translation from being truly implemented. Summary of the Invention

[0005] In view of this, the present invention provides a document translation method, apparatus, device, and storage medium to solve the problem that current document translation solutions cannot achieve a balance between document translation effectiveness and resource consumption. The technical solution is as follows:

[0006] A document translation method, comprising:

[0007] Obtain the source language document to be translated;

[0008] Obtain document attribute information corresponding to each sentence contained in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentences of the corresponding sentence to help determine the meaning of words in the corresponding sentence in the source language document;

[0009] Using the document attribute information corresponding to each sentence in the source language document, the sentences in the source language document are translated to obtain the target language translation of the source language document.

[0010] Optionally, the document attribute information includes one or more of the following attribute information: temporal attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information;

[0011] The temporal attribute information includes: the probability distribution of the corresponding sentence in each set temporal;

[0012] The ambiguity attribute information includes: a set of deambiguous words composed of words obtained from the context of the corresponding sentence to help remove ambiguity from the corresponding sentence;

[0013] The domain attribute information includes: the probability distribution of the corresponding sentence in each defined domain;

[0014] The coreference attribute information includes a set of coreference words composed of words that have a coreference relationship with the sentence and are obtained from the context of the corresponding sentence.

[0015] Optionally, the document attribute information includes: temporal attribute information and domain attribute information;

[0016] The process of translating each sentence in the source language document, supplemented by document attribute information corresponding to each sentence, includes:

[0017] Based on the sentences contained in the source language document and the tense attribute information and domain attribute information corresponding to each sentence, the target feature vector corresponding to each sentence in the source language document is determined, wherein the target feature vector contains the sentence-level text information, tense attribute information and domain attribute information of the corresponding sentence;

[0018] Based on the target feature vectors corresponding to each sentence in the source language document, the target language translation of the source language document is determined.

[0019] Optionally, determining the target feature vector corresponding to each sentence in the source language document based on each sentence and the corresponding tense attribute information and domain attribute information includes:

[0020] For the target sentence in the source language document whose corresponding target feature vector needs to be determined:

[0021] Based on the tense attribute information corresponding to the target sentence and the representation vector of each set tense, the tense attribute representation vector corresponding to the target sentence is determined, and based on the domain attribute information corresponding to the target sentence and the representation vector of each set domain, the domain attribute representation vector corresponding to the target sentence is determined.

[0022] The temporal attribute representation vector corresponding to the target sentence is fused with the domain attribute representation vector corresponding to the target sentence, and the fused vector is used as the control vector corresponding to the target sentence.

[0023] Obtain the sentence-level text representation vector of the target sentence;

[0024] The sentence-level text representation vector of the target sentence is fused with the control vector corresponding to the target sentence, and the fused vector is used as the target vector corresponding to the target sentence.

[0025] Optionally, the document attribute information includes: ambiguous attribute information;

[0026] The step of obtaining the sentence-level text representation vector of the target sentence includes:

[0027] The deambiguity word sequence, composed of all the deambiguities in the deambiguity word set corresponding to the target sentence, is concatenated with the target sentence to obtain the concatenated sentence;

[0028] Determine the representation vector of each word contained in the concatenated sentence;

[0029] Based on the representation vector of each word contained in the concatenated sentence, the sentence-level text representation vector of the target sentence is determined.

[0030] Optionally, determining the representation vector of each word contained in the concatenated sentence includes:

[0031] For each word in the concatenated sentence:

[0032] Obtain the text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located;

[0033] The text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located are fused together, and the fused vector is used as the representation vector of the word.

[0034] Optionally, the document attribute information includes: core reference attribute information;

[0035] The step of determining the target language translation corresponding to the source language document based on the target feature vectors corresponding to each sentence contained in the source language document includes:

[0036] The target vectors corresponding to each sentence in the source language document are decoded in the first pass, and the core reference translations corresponding to each sentence in the language document are cached. The core reference translations are the translations of core references in the core reference set corresponding to the sentence, and the core reference translations are obtained through the first pass of decoding.

[0037] Based on the cached information, the target vectors corresponding to each sentence in the source language document are decoded a second time, and the decoding result of the second decoding is used as the target language translation of the source language document.

[0038] Optionally, the second decoding of the target vectors corresponding to each sentence in the source language document, incorporating cached information, includes:

[0039] For each sentence contained in the source language document:

[0040] Retrieve the translation of the core reference word corresponding to the sentence from the cached information;

[0041] At each decoding moment, the context vector for the current decoding moment is determined based on the target vector corresponding to the sentence, and the decoding result for the current decoding moment is determined based on the core reference translation of the sentence, the decoded word sequence, and the context vector for the current decoding moment.

[0042] The decoding result for this sentence is composed of the decoding results at each decoding time.

[0043] Optionally, determining the decoding result at the current decoding moment based on the translated coreference words corresponding to the sentence, the decoded word sequence, and the context vector at the current decoding moment includes:

[0044] If there are multiple translations of the core reference word corresponding to the sentence, then the multiple translations of the core reference word are sorted according to the order in which the corresponding core reference word appears in the source language document to obtain the sequence of translations of the core reference word corresponding to the sentence;

[0045] Based on the core reference word translation sequence corresponding to the sentence, the decoded word sequence, and the context vector at the current decoding moment, determine the decoding result at the current decoding moment.

[0046] Optionally, the step of translating each sentence in the source language document using the document attribute information corresponding to each sentence in the source language document to obtain the target language translation of the source language document includes:

[0047] A structured document is constructed based on the source language document and the document attribute information corresponding to each sentence contained in the source language document. The structured document includes the structured information corresponding to each sentence contained in the source language document, and the structured information includes the corresponding sentence and the document attribute information corresponding to the corresponding sentence.

[0048] The structured document is processed based on a pre-trained document translation model to obtain the target language translation corresponding to the source language document. The document translation model is trained using a structured document constructed based on the training source language document and the document attribute information corresponding to each sentence in the training source language document, as well as the actual target language translation corresponding to the training source language document.

[0049] A document translation device includes: a document acquisition module, a document understanding module, and a document translation module;

[0050] The document acquisition module is used to acquire source language documents to be translated;

[0051] The document understanding module is used to obtain document attribute information corresponding to each sentence in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentence of the corresponding sentence to help determine the meaning of the words in the corresponding sentence in the source language document;

[0052] The document translation module is used to translate each sentence in the source language document, supplemented by the document attribute information corresponding to each sentence, to obtain the target language translation of the source language document.

[0053] A document translation device includes: a memory and a processor;

[0054] The memory is used to store programs;

[0055] The processor is configured to execute the program to implement each step of the document translation method described above.

[0056] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the document translation method described in any of the preceding claims.

[0057] The document translation method, apparatus, device, and storage medium provided by this invention, after obtaining a source language document, first acquires the document attribute information corresponding to each sentence in the source language document. Then, using the document attribute information, the sentences in the source language document are translated to obtain the target language translation. The document translation method provided by this invention does not directly introduce the context sentences of the sentence to be translated into the translation process. Instead, before actual translation, it determines information used to help determine the meaning of words in the sentence to be translated in the source language document based on the sentence to be translated and / or its context sentences—that is, the document attribute information corresponding to the sentence to be translated. This document attribute information is then used to translate the sentence. Compared to introducing context sentences, introducing document attribute information during actual translation reduces resource consumption. Context sentences contain useful information but also noise; introducing only key and useful information, compared to directly introducing context sentences, improves the document translation effect. The document translation method provided by this invention achieves a trade-off between document translation effect and resource consumption. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0059] Figure 1 A flowchart illustrating the document translation method provided in an embodiment of the present invention;

[0060] Figure 2 This is a flowchart illustrating the process of translating sentences in a source language document to obtain the target language translation of the source language document, supplemented by document attribute information corresponding to each sentence in the source language document, as provided in the embodiments of the present invention.

[0061] Figure 3 This is a flowchart illustrating the process of decoding the target vectors corresponding to each sentence in a source language document using coreference attribute information, as provided in this embodiment of the invention, to obtain the target language translation of the source language document.

[0062] Figure 4 A schematic diagram illustrating the process of implementing document translation based on a document translation model, as provided in an embodiment of the present invention;

[0063] Figure 5 A schematic diagram illustrating an example of a document translation model provided in an embodiment of the present invention;

[0064] Figure 6 A schematic diagram of the document translation device provided in an embodiment of the present invention;

[0065] Figure 7 A schematic diagram of the document translation device provided in an embodiment of the present invention. Detailed Implementation

[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0067] In the process of realizing this case, the inventors discovered that most current document translation solutions are translation model-based solutions. In order to obtain better translation results, the sentences to be translated and their context sentences in the source language document are usually input into the translation model for processing. In some cases, in order to obtain better translation results, a large number of context sentences may be input.

[0068] However, as contextual information increases, the consumption of computational resources and memory increases exponentially. In other words, using more contextual sentences consumes more memory and computational resources. To reduce memory and computational resource consumption, some current strategies constrain the context to one or a few sentences before and after the text. However, constraining the context to one or a few sentences before and after the text leads to a decrease in document translation quality. It is clear that current document translation solutions cannot achieve a balance between translation quality and resource consumption. For document translation to truly become practical, achieving this balance is essential. Furthermore, directly introducing contextual sentences into the translation model places significant pressure on it, as the model needs to perform both document understanding and translation.

[0069] Given the numerous problems with current document translation solutions, the inventors of this case conducted in-depth research and ultimately proposed a more effective document translation method. The basic concept of this method is as follows: before translation, document understanding is performed on the source language document to be translated, obtaining the document attribute information corresponding to each sentence in the source language document. This obtained document attribute information is then introduced into both the source and target translation ends, and combined with the document attribute information corresponding to each sentence in the source language document, the sentences in the source language document are translated. This document translation method perfectly solves many of the problems existing in current document translation solutions.

[0070] Before introducing the document translation method provided by this invention, the hardware architecture involved in this invention will be described first.

[0071] In one possible implementation, the hardware architecture involved in this invention may include: electronic devices and servers.

[0072] For example, an electronic device can be any electronic product that can interact with a user, such as a tablet computer, PDA, mobile phone, translator, etc.

[0073] For example, a server can be a single server, a server cluster consisting of multiple servers, or a cloud computing server center. A server may include a processor, memory, and network interfaces, etc.

[0074] For example, an electronic device can establish a connection and communicate with a server through a wireless communication network; for example, an electronic device can establish a connection and communicate with a server through a wired communication network.

[0075] The electronic device can acquire the source language document to be translated, send the source language document to the server, and the server translates the source language document according to the document translation method provided in this invention, and sends the translation result of the source language document to the electronic device.

[0076] In another possible implementation, the hardware architecture involved in this invention may include an electronic device. The electronic device is an electronic product with strong data processing capabilities, such as a tablet computer, PDA, mobile phone, or translation device. The electronic device translates the source language document to be translated according to the translation method provided by this invention.

[0077] Those skilled in the art should understand that the above-described electronic devices and servers are merely examples, and other existing or future electronic devices or servers that are applicable to this invention should also be included within the scope of protection of this invention, and are hereby incorporated by reference.

[0078] The document translation method provided by the present invention will be described below through the following embodiments.

[0079] First Embodiment

[0080] Please see Figure 1 The diagram illustrates a flowchart of a document translation method provided in an embodiment of the present invention, which may include:

[0081] Step S101: Obtain the source language document to be translated.

[0082] The source language document can be a document in any language and includes several sentences.

[0083] Step S102: Obtain the document attribute information corresponding to each sentence in the source language document.

[0084] Among them, document attribute information is information used to help determine the meaning of words in the corresponding sentence in the source language document, based on the corresponding sentence and / or the context of the corresponding sentence.

[0085] Optionally, the document attribute information includes one or more of the following attributes: tense attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information. For better document translation results, the document attribute information preferably includes tense attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information simultaneously.

[0086] Among them, the tense attribute information is the relevant information of the corresponding sentence in terms of the tense attribute, which can reflect the tense to which the corresponding sentence belongs; the ambiguity attribute information is the relevant information of the corresponding sentence in terms of the ambiguity attribute, which is the contextual information that can help the corresponding sentence eliminate ambiguity; the domain attribute information is the relevant information of the corresponding sentence in terms of the domain attribute, which can reflect the domain to which the corresponding sentence belongs; and the coreference attribute information is the relevant information of the corresponding sentence in terms of the coreference attribute, which is the contextual information that has a coreference relationship with the corresponding sentence.

[0087] Optionally, the tense attribute information may include the probability distribution of the corresponding sentence in various set tenses (e.g., simple past, simple present, simple future, etc.); the ambiguity attribute information may include a set of disambiguation words obtained from the context sentences (e.g., all context sentences) to help disambiguate the corresponding sentence; the domain attribute information may include the probability distribution of the corresponding sentence in various set domains; and the coreference attribute information may include a set of coreference words obtained from the context sentences (e.g., all context sentences) that have a coreference relationship with the sentence. It should be noted that the probability distribution of a sentence in various set tenses refers to the probability that the sentence's tense is in each set tense; similarly, the probability distribution of a sentence in various set domains refers to the probability that the sentence belongs to each set domain.

[0088] Optionally, a pre-trained tense prediction model can be used to predict the probability distribution of each sentence in the source language document across various set tenses. The tense prediction model is trained using training sentences labeled with tenses. It should be noted that this embodiment is not limited to using a tense prediction model to predict the probability distribution of each sentence in the source language document across various set tenses; any method capable of predicting the probability distribution of a sentence across various set tenses is applicable to this invention.

[0089] Optionally, words representing the same entity can be extracted from the source language document based on a document reference resolution model. Based on these extracted words, a set of coreference words is obtained, consisting of words that have a coreference relationship with each sentence in the source language document. After obtaining the set of coreference words for each sentence in the source language document, the position (e.g., sentence number) of each coreference word in the set can be recorded. It should be noted that this embodiment is not limited to extracting words representing the same entity from the source language document based on a document reference resolution model; any method capable of extracting words representing the same entity from the source language document is applicable to this invention.

[0090] Optionally, for each sentence in the source language document, ambiguity can be determined, i.e., whether the sentence is ambiguous. If so, words with high similarity to the ambiguous word (e.g., words with similarity greater than a preset similarity threshold) are extracted from the source language document to form a deambiguity word set. It should be noted that if a sentence in the source language document is unambiguous, its corresponding deambiguity word set is empty.

[0091] Optionally, for each sentence in the source language document, the probability distribution of each sentence in each defined domain can be predicted based on a pre-trained domain prediction model. The domain prediction model is trained using domain-annotated training sentences. It should be noted that this embodiment is not limited to predicting the probability distribution of each sentence in each defined domain based on a domain prediction model; any method capable of predicting the probability distribution of a sentence in each defined domain is applicable to this invention.

[0092] For document translation, sentence translation often relies on contextual information. However, contextual information is often limited to certain sentences or words in the document. Most of the context is noise for sentence translation. Generally, the more context used, the greater the noise tends to be, and the greater the impact on the document translation effect. Therefore, this invention proposes not to directly introduce context into the translation process, but to extract key information, namely document attribute information, by understanding the document before translation. This strategy avoids the impact of noise on the translation effect on the one hand, and avoids understanding the document during the translation process on the other hand, thus alleviating the pressure of the translation stage.

[0093] Step S103: Using the document attribute information corresponding to each sentence in the source language document, translate each sentence in the source language document to obtain the target language translation corresponding to the source language document.

[0094] This invention translates sentences in a source language document by combining the document attribute information corresponding to each sentence.

[0095] The document translation method provided in this invention, after obtaining the source language document, first acquires the document attribute information corresponding to each sentence in the source language document. Then, using the document attribute information, it translates each sentence in the source language document to obtain the target language translation. This method does not directly introduce the context sentences of the sentence to be translated into the translation process. Instead, before actual translation, it determines information to aid in determining the meaning of words in the sentence to be translated within the source language document, based on the sentence to be translated and / or its context sentences. This information is the document attribute information corresponding to the sentence to be translated. The translation is then performed using this document attribute information. Compared to directly introducing context sentences, introducing document attribute information during actual translation reduces resource consumption. Context sentences contain useful information but also noise; introducing only key and useful information improves the document translation effect. This method achieves a balance between document translation effectiveness and resource consumption.

[0096] Second Embodiment

[0097] This embodiment focuses on the specific implementation process of "step S103: using the document attribute information corresponding to each sentence in the source language document to translate each sentence in the source language document to obtain the target language translation corresponding to the source language document" in the above embodiment.

[0098] This embodiment takes document attribute information, including tense attribute information, ambiguity attribute information, domain attribute information, and coreference attribute information, as an example. It introduces the specific implementation process of translating each sentence in the source language document, supplemented by the document attribute information corresponding to each sentence in the source language document, to obtain the target language translation of the source language document.

[0099] Please see Figure 2 This diagram illustrates the process of translating sentences in a source language document to obtain the target language translation, supplemented by document attribute information corresponding to each sentence in the source language document. The process may include:

[0100] Step S201: Based on the sentences contained in the source language document and the temporal and domain attribute information corresponding to each sentence, determine the target feature vector corresponding to each sentence contained in the source language document.

[0101] The target feature vector contains sentence-level text information, temporal attribute information, and domain attribute information of the corresponding sentence.

[0102] Specifically, the process of determining the target feature vectors corresponding to each sentence in the source language document, based on the sentences contained therein and the temporal and domain attribute information corresponding to each sentence, includes the following steps for the target sentences in the source language document whose corresponding target feature vectors are to be determined:

[0103] Step a1: Based on the tense attribute information corresponding to the target sentence and the representation vectors of each set tense, determine the tense attribute representation vector corresponding to the target sentence, and based on the domain attribute information corresponding to the target sentence and the representation vectors of each set domain, determine the domain attribute representation vector corresponding to the target sentence.

[0104] The above embodiments mention that the tense attribute information may include the probability distribution of the corresponding sentence in each set tense, that is, the probability that the tense of the corresponding sentence is each set tense. The probability that the tense of the corresponding sentence is each set tense can be used as a weight to sum the representation vectors of each set tense to obtain the tense attribute representation vector corresponding to the target sentence. :

[0105] (1).

[0106] In equation (1), TL represents the number of time states, and w i e represents the probability that the tense of the corresponding sentence is the i-th specified tense. i This represents the representation vector for the i-th set time state.

[0107] The above embodiments mention that the domain attribute information may include the probability distribution of the corresponding sentence in each set domain, that is, the probability that the domain of the corresponding sentence is each set domain. The probability that the domain of the corresponding sentence is each set domain can be used as a weight to sum the representation vectors of each set domain to obtain the domain attribute representation vector E corresponding to the target sentence. field :

[0108] (2).

[0109] In equation (2), FL represents the number of domains, and w i e represents the probability that the corresponding sentence belongs to the i-th defined domain. i This represents the representation vector of the i-th defined domain.

[0110] because and Since the numbers are continuous real numbers, this invention can achieve finer-grained temporal and domain control compared to directly using temporal and domain representation vectors.

[0111] Step a2: Merge the temporal attribute representation vector corresponding to the target sentence with the domain attribute representation vector corresponding to the target sentence, and use the merged vector as the control vector corresponding to the target sentence.

[0112] There are several ways to fuse the temporal attribute representation vector of the target sentence with the domain attribute representation vector of the target sentence. For example, the temporal attribute representation vector of the target sentence can be weighted and summed with the domain attribute representation vector of the target sentence. Alternatively, the temporal attribute representation vector of the target sentence can be directly summed with the domain attribute representation vector of the target sentence.

[0113] Step a3: Obtain the sentence-level text representation vector of the target sentence.

[0114] There are several ways to obtain the sentence-level text representation vector of the target sentence. In one possible approach, the representation vector of each word in the target sentence can be determined first to obtain the word representation vector sequence corresponding to the target sentence. Then, the word representation vector sequence corresponding to the target sentence can be encoded to obtain the sentence-level text representation vector of the target sentence.

[0115] In the process of realizing this invention, the inventors discovered that single sentences often suffer from ambiguity due to insufficient information, leading to multiple reasonable translations for some words. However, the rich contextual information in documents offers an opportunity to resolve such unconstrained sentence-level ambiguities. However, this rich contextual information typically contains significant noise, and coupled with limitations in current technology and hardware computing power, directly utilizing this rich contextual information to resolve sentence-level ambiguities is ineffective. Considering that deambiguation often only requires some keyword information from the document, this invention proposes extracting deambiguous words (words that aid in deambiguation) from the rich contextual information, thus obtaining the aforementioned ambiguity attribute information. This ambiguity attribute information is then used to resolve sentence-level ambiguities. Following this line of thought, another possible implementation method is proposed in this invention:

[0116] The deambiguity word sequence, composed of all deambiguities in the deambiguity word set corresponding to the target sentence, is concatenated with the target sentence to obtain the concatenated sentence. The representation vector of each word in the concatenated sentence is determined to obtain the word representation vector sequence corresponding to the concatenated sentence. This word representation vector sequence is then encoded to obtain the sentence-level text representation vector of the target sentence. It should be noted that if the target sentence is unambiguous, i.e., the deambiguity word set corresponding to the target sentence is empty, then only the representation vector of each word in the target sentence needs to be determined to obtain the word representation vector sequence corresponding to the target sentence. This word representation vector sequence is then encoded to obtain the sentence-level text representation vector of the target sentence.

[0117] The process of determining the representation vector of a word may include: obtaining the text representation vector of the word, the positional representation vector of the word, and the positional representation vector of the sentence in which the word is located; fusing the text representation vector of the word, the positional representation vector of the word, and the positional representation vector of the sentence in which the word is located, and using the fused vector as the representation vector of the word. It should be noted that the positional representation vector of the word refers to the representation vector of the word's position within the sentence it is located in, while the positional representation vector of the sentence in which the word is located refers to the representation vector of the sentence's position within the source language text.

[0118] Step a4: Merge the sentence-level text representation vector of the target sentence with the control vector corresponding to the target sentence, and use the merged vector as the target vector corresponding to the target sentence.

[0119] There are several ways to fuse the sentence-level text representation vector of the target sentence with the corresponding control vector. For example, the sentence-level text representation vector of the target sentence can be weighted and summed with the corresponding control vector. Alternatively, the sentence-level text representation vector of the target sentence can be directly summed with the corresponding control vector.

[0120] Step S202: Determine the target language translation of the source language document based on the target vectors corresponding to each sentence in the source language document.

[0121] There are several ways to determine the target language translation of a source language document based on the target vectors corresponding to each sentence in the source language document. In one possible approach, the target vectors corresponding to each sentence in the source language document can be directly decoded to obtain the target language translation. In another possible approach, document attribute information includes coreference attribute information, which can be used to decode the target vectors corresponding to each sentence in the source language document to obtain the target language translation.

[0122] Given that the second implementation method described above is more effective, we will now introduce the second implementation method.

[0123] Please see Figure 3 This illustrates a flowchart of the process of decoding the target vectors corresponding to each sentence in a source language document using coreference attribute information to obtain the target language translation of the source language document. The process may include:

[0124] Step S301: Decode the target vectors corresponding to each sentence in the source language document for the first time, and cache the core reference translations corresponding to each sentence in the source language document.

[0125] Among them, the translation of core reference words is the translation of core reference words in the core reference word set corresponding to the sentence. The translation of core reference words is obtained through the first decoding.

[0126] Step S302: Combining the cached information, perform a second decoding on the target vectors corresponding to each sentence in the source language document. The decoding result of the second decoding is used as the target language translation of the source language document.

[0127] Specifically, the process of performing a second decoding of the target vectors corresponding to each sentence in the source language document, based on cached information, can include: for each sentence in the source language document, obtaining the translated coreference words corresponding to that sentence from the cached information; and performing a second decoding of the target vector corresponding to that sentence based on the translated coreference words. More specifically, for each sentence in the source language document: obtaining the translated coreference words corresponding to that sentence from the cached information; at each decoding time, determining the context vector for the current decoding time based on the target vector corresponding to that sentence; determining the decoding result for the current decoding time based on the translated coreference words, the decoded word sequence, and the context vector for the current decoding time; and combining the decoding results at each decoding time to form the decoding result corresponding to that sentence.

[0128] It should be noted that the core reference set corresponding to a sentence may include one core reference word or multiple core reference words. In the case that the core reference set corresponding to a sentence includes multiple core reference words, the translations of multiple core reference words can be sorted according to the order in which the corresponding core reference words appear in the source language document to obtain the core reference word translation sequence corresponding to the sentence. Then, combined with the core reference word translation sequence corresponding to the sentence, the target vector corresponding to the sentence is decoded a second time.

[0129] Third Embodiment

[0130] In one possible implementation, document translation can be achieved based on a document translation model. This embodiment focuses on introducing the process of achieving document translation based on a document translation model.

[0131] Please see Figure 4 The diagram illustrates a workflow for document translation based on a document translation model, which may include:

[0132] Step S401: Obtain the source language document to be translated.

[0133] The source language document can be a document in any language and includes several sentences.

[0134] Step S402: Obtain the document attribute information corresponding to each sentence in the source language document.

[0135] The document attribute information refers to information determined based on the corresponding sentence and / or its context sentences, used to assist in determining the meaning of words in the corresponding sentence within the source language document. Optionally, the document attribute information includes one or more of the following attribute information: tense attribute information, ambiguity attribute information, domain attribute information, and coreference attribute information. Preferably, the document attribute information includes tense attribute information, ambiguity attribute information, domain attribute information, and coreference attribute information. For relevant explanations and methods of obtaining document attribute information, please refer to the relevant sections in the first embodiment; these will not be repeated here.

[0136] Step S403: Construct a structured document based on the source language document and the document attribute information corresponding to each sentence contained in the source language document.

[0137] The structured document includes the structured information corresponding to each sentence in the source language document. The structured information includes the corresponding sentence and the document attribute information corresponding to the corresponding sentence.

[0138] For example, if document attribute information includes tense attribute information, ambiguity attribute information, domain attribute information, and coreference attribute information, then the structured information corresponding to the i-th sentence in the source language document can be in the following form: {{text, X} i}, {temporal attribute information, P i时态}, {Ambiguous attribute information, W i},{Domain attribute information, P i领域}, {common attribute information, C i}}, where X i For the i-th sentence in the source language document, P i时态 For the i-th sentence X i The probability distribution W at each given time state i For example, from the i-th sentence X i The context obtained to help the i-th sentence X i A deambiguous word set composed of unambiguous words, P i领域 For the i-th sentence X i The probability distribution in each defined domain, C i For example, from the i-th sentence X i The context obtained from the i-th sentence X i A set of words that share a common reference.

[0139] It should be noted that the position of each piece of structured information in a structured document is the same as the position of the sentence it contains in the source language document. For example, X i If X is the i-th sentence in the source language document, then it contains X. iThe corresponding structured information is the i-th structured information in the structured document.

[0140] Step S404: Process the structured document based on the pre-trained document translation model to obtain the target language translation corresponding to the source language document.

[0141] Specifically, the structured document is input into the pre-trained document translation model and processed to obtain the target language translation of the source language document output by the document translation model.

[0142] The document translation model is trained using a structured document constructed from the training source language document and the document attribute information corresponding to each sentence in the training source language document, as well as the real target language translation corresponding to the training source language document.

[0143] When training the document translation model, the document translation model is first constructed by taking a structured document as input, which is based on the training source language document and the document attribute information corresponding to each sentence in the training source language document. The output of the document translation model is the translation result corresponding to the training source language document. Then, the prediction loss (such as cross-entropy loss) of the document translation model is determined according to the translation result corresponding to the training source language document and the real target language translation corresponding to the training source language document. Finally, the parameters of the document translation model are updated according to the prediction loss of the document translation model. The document translation model is trained iteratively multiple times in the above manner until the training termination condition is met, such as reaching the preset number of training times or the model converges.

[0144] Specifically, the process of processing structured documents based on document translation models to obtain the target language translation corresponding to the source language document can include: processing the sentence, tense attribute information, domain attribute information, and ambiguous attribute information in each piece of structured information contained in the structured document based on the document translation model to obtain the target feature vector corresponding to each sentence in the structured document. The target feature vector contains the sentence-level text information, tense attribute information, domain attribute information, and ambiguous attribute information of the corresponding sentence; combining the coreference attribute information in the structured information contained in the structured document, processing the obtained target feature vector based on the document translation model to obtain the target language translation corresponding to the source language document.

[0145] Please see Figure 5 This shows an example of a document translation model, and the following will be... Figure 5 Taking the document translation model shown as an example, the process of processing structured documents based on the document translation model to obtain the target language translation corresponding to the source language document will be further introduced.

[0146] like Figure 5As shown, the document translation model may include a control vector determination module 501, an encoding module 502, and a decoding module 503. After obtaining a structured document, the structured document can be input into the document translation model. When inputting the structured document into the document translation model, each piece of structured information in the structured document can be input into the document translation model in sequence for processing. The processing procedure of the document translation model is as follows:

[0147] For the i-th sentence X in the source language document i The corresponding structured information s i :

[0148] On the one hand, s i The domain attribute information and temporal attribute information in X i The corresponding domain attribute information and temporal attribute information are input into the control vector determination module 501. The control vector determination module 501 first determines the control vector based on X. i The corresponding temporal attribute information and the representation vector of each set temporal state are used to determine X. i The corresponding temporal attribute representation vector, based on X i Based on the corresponding domain attribute information and the representation vectors of each domain, X is determined. i The corresponding domain attribute representation vector, then X i The corresponding temporal attribute representation vector and X i The corresponding domain attribute representation vectors are fused and output, and the vector output by the control vector determination module 501 is used as X. i The corresponding control vector.

[0149] On the other hand, s i X is in i The sum is X i The corresponding ambiguous attribute information (set of ambiguous words) is input into the encoding module 502 for processing. The encoding module 502 first converts X... i The deambiguity sequence formed by the corresponding deambiguities in the deambiguity word set is used as a sentence and X i The process involves concatenating the words, determining the representation vector of each word in the concatenated sentence to obtain a sequence of word representation vectors, and finally encoding this sequence to output X. i The sentence-level text representation vector. Specifically, for each word in the concatenated sentence, when determining the representation vector of the word, the encoding module 502 can obtain the text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located. The text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located are then fused together, and the fused vector is used as the representation vector of the word.

[0150] In obtaining X i Sentence-level text representation vectors and Xi After the corresponding control vector, X can be... i Sentence-level text representation vector and X i The corresponding control vectors are fused, and the fused vector is used as X. i The corresponding target vector.

[0151] X i The corresponding target vector is input to the decoding module 503 for the first pass of decoding to obtain x. i The corresponding first decoding result.

[0152] By processing each piece of structured information in the structured document according to the above process, we can obtain the first decoding results corresponding to each sentence in the source language document.

[0153] After obtaining the first decoding result, the coreference translations corresponding to each sentence in the source language document can be cached. The coreference translations are the translations of the coreference words in the coreference set corresponding to each sentence, obtained through the first decoding pass. To model consistency in the document translation process, previously decoded translations can be cached. When encountering similar source texts that need translation, it is encouraged to select previously used translations from the cache, thus improving the consistency of wording throughout the document translation process. However, if the entire cached translation is considered for all source texts, it can easily lead to error propagation and mistranslation. Therefore, this case adopts a decoding mechanism based on group caching (grouping and caching the coreference translations corresponding to each sentence in the source language document; for example, the coreference translations corresponding to a sentence are grouped together). On the one hand, this encourages more consistent translations for words with the same meaning; on the other hand, the group caching mechanism can alleviate the problems of error propagation and mistranslation.

[0154] Finally, combining the cached information, a second decoding is performed on the target vectors corresponding to each sentence in the source language document. The decoding result of the second decoding is used as the target language translation of the source language document. Specifically, for each sentence in the source language document, the core reference translation corresponding to the sentence is obtained from the cached information. The core reference translation and the target vector corresponding to the sentence are input into the decoding module 503. At each decoding time, the decoding module 503 determines the context vector of the current decoding time based on the target vector corresponding to the sentence. Based on the core reference translation, the decoded word sequence, and the context vector of the current decoding time, the decoding result of the current decoding time is determined. For example, at decoding time t, the decoding module 503 determines the context vector c_t based on the target vector corresponding to the sentence. Based on the core reference translation, the decoded word sequence Y_(t-1)=(y_1,y_2,…,y_(t-1)), and the context vector c_t of the current decoding time, the current word y_t is calculated and generated. The decoding results of each decoding time constitute the decoding result corresponding to the sentence.

[0155] Optionally, if a sentence has multiple core reference translations, the multiple core reference translations can be sorted according to the order in which the corresponding core references appear in the source language document to obtain the core reference translation sequence for that sentence. The core reference translations and the target vector for that sentence are then input into the decoding module 503. Accordingly, at each decoding moment, the decoding module 503 determines the context vector for the current decoding moment based on the target vector for that sentence, and determines the decoding result for the current decoding moment based on the core reference translation sequence, the decoded word sequence, and the context vector for the current decoding moment.

[0156] The document translation method provided in this invention does not directly introduce the context sentences of the sentence to be translated from the source language document into the translation process. Instead, before actual translation, it determines information to aid in determining the meaning of words in the sentence to be translated in the source language document based on the sentence to be translated and / or its context sentences. This information is the document attribute information corresponding to the sentence to be translated. Then, it simultaneously introduces document attribute information from both the source and target translation ends, supplementing the translation of the sentence to be translated with this document attribute information. Compared to directly introducing context sentences, introducing document attribute information during actual translation reduces resource consumption. Context sentences contain useful information, but also noise. Introducing only key and useful information, compared to directly introducing context sentences, improves the document translation effect. The document translation method provided in this invention achieves a trade-off between document translation effect and resource consumption.

[0157] Fourth embodiment

[0158] This invention also provides a document translation device. The document translation device provided in this invention will be described below. The document translation device described below can be referred to in correspondence with the document translation method described above.

[0159] Please see Figure 6 The diagram shows a structural schematic of a document translation device provided in an embodiment of the present invention, which may include: a document acquisition module 601, a document understanding module 602, and a document translation module 603.

[0160] The document acquisition module 601 is used to acquire source language documents to be translated.

[0161] The document understanding module 602 is used to obtain the document attribute information corresponding to each sentence in the source language document.

[0162] The document attribute information is information determined based on the corresponding sentence and / or the context of the corresponding sentence, used to help determine the meaning of words in the corresponding sentence in the source language document.

[0163] The document translation module 603 is used to translate each sentence in the source language document, supplemented by the document attribute information corresponding to each sentence in the source language document, to obtain the target language translation of the source language document.

[0164] Optionally, the document attribute information includes one or more of the following attribute information: temporal attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information;

[0165] The temporal attribute information includes: the probability distribution of the corresponding sentence in each set temporal;

[0166] The ambiguity attribute information includes: a set of deambiguous words composed of words obtained from the context of the corresponding sentence to help remove ambiguity from the corresponding sentence;

[0167] The domain attribute information includes: the probability distribution of the corresponding sentence in each defined domain;

[0168] The coreference attribute information includes a set of coreference words composed of words that have a coreference relationship with the sentence and are obtained from the context of the corresponding sentence.

[0169] Optionally, the document attribute information includes: temporal attribute information and domain attribute information;

[0170] When translating the sentences in the source language document using the document attribute information corresponding to each sentence, the document translation module 603 is specifically used for:

[0171] Based on the sentences contained in the source language document and the tense attribute information and domain attribute information corresponding to each sentence, the target feature vector corresponding to each sentence in the source language document is determined, wherein the target feature vector contains the sentence-level text information, tense attribute information and domain attribute information of the corresponding sentence;

[0172] Based on the target feature vectors corresponding to each sentence in the source language document, the target language translation of the source language document is determined.

[0173] When determining the target feature vector corresponding to each sentence in the source language document based on the sentences contained in the source language document and the temporal attribute information and domain attribute information corresponding to each sentence, the document translation module 603 is specifically used for:

[0174] For the target sentence in the source language document whose corresponding target feature vector needs to be determined:

[0175] Based on the tense attribute information corresponding to the target sentence and the representation vector of each set tense, the tense attribute representation vector corresponding to the target sentence is determined, and based on the domain attribute information corresponding to the target sentence and the representation vector of each set domain, the domain attribute representation vector corresponding to the target sentence is determined.

[0176] The temporal attribute representation vector corresponding to the target sentence is fused with the domain attribute representation vector corresponding to the target sentence, and the fused vector is used as the control vector corresponding to the target sentence.

[0177] Obtain the sentence-level text representation vector of the target sentence;

[0178] The sentence-level text representation vector of the target sentence is fused with the control vector corresponding to the target sentence, and the fused vector is used as the target vector corresponding to the target sentence.

[0179] Optionally, the document attribute information includes: ambiguous attribute information;

[0180] When obtaining the sentence-level text representation vector of the target sentence, the document translation module 603 is specifically used for:

[0181] The deambiguity word sequence, composed of all the deambiguities in the deambiguity word set corresponding to the target sentence, is concatenated with the target sentence to obtain the concatenated sentence;

[0182] Determine the representation vector of each word contained in the concatenated sentence;

[0183] Based on the representation vector of each word contained in the concatenated sentence, the sentence-level text representation vector of the target sentence is determined.

[0184] Optionally, when determining the representation vector of each word contained in the concatenated sentence, the document translation module 603 is specifically used for:

[0185] For each word in the concatenated sentence:

[0186] Obtain the text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located;

[0187] The text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located are fused together, and the fused vector is used as the representation vector of the word.

[0188] Optionally, the document attribute information includes: core reference attribute information;

[0189] When determining the target language translation of the source language document based on the target feature vectors corresponding to each sentence in the source language document, the document translation module 603 is specifically used for:

[0190] The target vectors corresponding to each sentence in the source language document are decoded in the first pass, and the core reference translations corresponding to each sentence in the language document are cached. The core reference translations are the translations of core references in the core reference set corresponding to the sentence, and the core reference translations are obtained through the first pass of decoding.

[0191] Based on the cached information, the target vectors corresponding to each sentence in the source language document are decoded a second time, and the decoding result of the second decoding is used as the target language translation of the source language document.

[0192] Optionally, when the document translation module 603 performs a second decoding of the target vectors corresponding to each sentence in the source language document, in conjunction with cached information, it specifically performs the following:

[0193] For each sentence contained in the source language document:

[0194] Retrieve the translation of the core reference word corresponding to the sentence from the cached information;

[0195] At each decoding moment, the context vector for the current decoding moment is determined based on the target vector corresponding to the sentence, and the decoding result for the current decoding moment is determined based on the core reference translation of the sentence, the decoded word sequence, and the context vector for the current decoding moment.

[0196] The decoding result for this sentence is composed of the decoding results at each decoding time.

[0197] Optionally, when the document translation module 603 determines the decoding result at the current decoding moment based on the translated coreferences of the sentence, the decoded word sequence, and the context vector at the current decoding moment, it is specifically used for:

[0198] If there are multiple translations of the core reference word corresponding to the sentence, then the multiple translations of the core reference word are sorted according to the order in which the corresponding core reference word appears in the source language document to obtain the sequence of translations of the core reference word corresponding to the sentence;

[0199] Based on the core reference word translation sequence corresponding to the sentence, the decoded word sequence, and the context vector at the current decoding moment, determine the decoding result at the current decoding moment.

[0200] Optionally, when the document translation module 603 translates each sentence in the source language document using the document attribute information corresponding to each sentence in the source language document to obtain the target language translation corresponding to the source language document, it is specifically used for:

[0201] A structured document is constructed based on the source language document and the document attribute information corresponding to each sentence contained in the source language document. The structured document includes the structured information corresponding to each sentence contained in the source language document, and the structured information includes the corresponding sentence and the document attribute information corresponding to the corresponding sentence.

[0202] The structured document is processed based on a pre-trained document translation model to obtain the target language translation corresponding to the source language document. The document translation model is trained using a structured document constructed based on the training source language document and the document attribute information corresponding to each sentence in the training source language document, as well as the actual target language translation corresponding to the training source language document.

[0203] The document translation apparatus provided in this invention determines, before actual translation, information to aid in determining the meaning of words in the sentence to be translated in the source language document based on the sentence to be translated and / or its context sentences. This information is the document attribute information corresponding to the sentence to be translated. The document attribute information is then simultaneously introduced from both the source and target ends of the translation process, and the translation is performed using this information. Compared to directly introducing context sentences, introducing document attribute information during actual translation reduces resource consumption. Context sentences contain useful information but also noise; introducing only key and useful information, compared to directly introducing context sentences, improves the document translation effect. The document translation apparatus provided in this invention achieves a trade-off between document translation effect and resource consumption.

[0204] Fifth Embodiment

[0205] This invention also provides a document translation device; please refer to [link / reference]. Figure 7 The diagram shows the structure of the document translation device, which may include: at least one processor 701, at least one communication interface 702, at least one memory 703 and at least one communication bus 704;

[0206] In this embodiment of the invention, the number of processor 701, communication interface 702, memory 703 and communication bus 704 is at least one, and processor 701, communication interface 702 and memory 703 communicate with each other through communication bus 704.

[0207] The processor 701 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0208] The memory 703 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0209] The memory stores a program, which the processor can call. The program is used for:

[0210] Obtain the source language document to be translated;

[0211] Obtain document attribute information corresponding to each sentence contained in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentences of the corresponding sentence to help determine the meaning of words in the corresponding sentence in the source language document;

[0212] Using the document attribute information corresponding to each sentence in the source language document, the sentences in the source language document are translated to obtain the target language translation of the source language document.

[0213] Optionally, the refined and extended functions of the program can be found in the description above.

[0214] Sixth Embodiment

[0215] This invention also provides a readable storage medium that stores a program suitable for execution by a processor, the program being used for:

[0216] Obtain the source language document to be translated;

[0217] Obtain document attribute information corresponding to each sentence contained in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentences of the corresponding sentence to help determine the meaning of words in the corresponding sentence in the source language document;

[0218] Using the document attribute information corresponding to each sentence in the source language document, the sentences in the source language document are translated to obtain the target language translation of the source language document.

[0219] Optionally, the refined and extended functions of the program can be found in the description above.

[0220] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0221] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0222] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A document translation method, characterized in that, include: Obtain the source language document to be translated; Obtain document attribute information corresponding to each sentence contained in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentences of the corresponding sentence to help determine the meaning of words in the corresponding sentence in the source language document; Using the document attribute information corresponding to each sentence in the source language document, the sentences in the source language document are translated to obtain the target language translation of the source language document; The document attribute information includes one or more of the following attribute information: temporal attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information; The temporal attribute information includes: the probability distribution of the corresponding sentence in each set temporal; The ambiguity attribute information includes: a set of deambiguous words composed of words obtained from the context of the corresponding sentence to help remove ambiguity from the corresponding sentence; The domain attribute information includes: the probability distribution of the corresponding sentence in each defined domain; The coreference attribute information includes a set of coreference words composed of words that have a coreference relationship with the sentence and are obtained from the context of the corresponding sentence.

2. The document translation method according to claim 1, characterized in that, The document attribute information includes: temporal attribute information and domain attribute information; The process of translating each sentence in the source language document, supplemented by document attribute information corresponding to each sentence, includes: Based on the sentences contained in the source language document and the tense attribute information and domain attribute information corresponding to each sentence, the target feature vector corresponding to each sentence in the source language document is determined, wherein the target feature vector contains the sentence-level text information, tense attribute information and domain attribute information of the corresponding sentence; Based on the target feature vectors corresponding to each sentence in the source language document, the target language translation of the source language document is determined.

3. The document translation method according to claim 2, characterized in that, The step of determining the target feature vector corresponding to each sentence in the source language document based on each sentence and the temporal and domain attribute information corresponding to each sentence includes: For the target sentence in the source language document whose corresponding target feature vector needs to be determined: Based on the tense attribute information corresponding to the target sentence and the representation vector of each set tense, the tense attribute representation vector corresponding to the target sentence is determined, and based on the domain attribute information corresponding to the target sentence and the representation vector of each set domain, the domain attribute representation vector corresponding to the target sentence is determined. The temporal attribute representation vector corresponding to the target sentence is fused with the domain attribute representation vector corresponding to the target sentence, and the fused vector is used as the control vector corresponding to the target sentence. Obtain the sentence-level text representation vector of the target sentence; The sentence-level text representation vector of the target sentence is fused with the control vector corresponding to the target sentence, and the fused vector is used as the target vector corresponding to the target sentence.

4. The document translation method according to claim 3, characterized in that, The document attribute information includes: ambiguous attribute information; The step of obtaining the sentence-level text representation vector of the target sentence includes: The deambiguity word sequence, composed of all the deambiguities in the deambiguity word set corresponding to the target sentence, is concatenated with the target sentence to obtain the concatenated sentence; Determine the representation vector of each word contained in the concatenated sentence; Based on the representation vector of each word contained in the concatenated sentence, the sentence-level text representation vector of the target sentence is determined.

5. The document translation method according to claim 4, characterized in that, Determining the representation vector of each word contained in the concatenated sentence includes: For each word in the concatenated sentence: Obtain the text representation vector of the word, the position representation vector of the word within the sentence, and the position representation vector of the word within the sentence; The text representation vector of the word, the position representation vector of the word, and the position representation vector of the sentence in which the word is located are fused together, and the fused vector is used as the representation vector of the word.

6. The document translation method according to claim 2, characterized in that, The document attribute information includes: core reference attribute information; The step of determining the target language translation corresponding to the source language document based on the target feature vectors corresponding to each sentence contained in the source language document includes: The target vectors corresponding to each sentence in the source language document are decoded in the first pass, and the core reference translations corresponding to each sentence in the language document are cached. The core reference translations are the translations of core references in the core reference set corresponding to the sentence, and the core reference translations are obtained through the first pass of decoding. Based on the cached information, the target vectors corresponding to each sentence in the source language document are decoded a second time, and the decoding result of the second decoding is used as the target language translation of the source language document.

7. The document translation method according to claim 6, characterized in that, The second decoding of the target vectors corresponding to each sentence in the source language document, incorporating cached information, includes: For each sentence contained in the source language document: Retrieve the translation of the core reference word corresponding to the sentence from the cached information; At each decoding moment, the context vector for the current decoding moment is determined based on the target vector corresponding to the sentence, and the decoding result for the current decoding moment is determined based on the core reference translation of the sentence, the decoded word sequence, and the context vector for the current decoding moment. The decoding result for this sentence is composed of the decoding results at each decoding time.

8. The document translation method according to claim 7, characterized in that, The step of determining the decoding result at the current decoding moment based on the translation of the core reference word corresponding to the sentence, the decoded word sequence, and the context vector at the current decoding moment includes: If there are multiple translations of the core reference word corresponding to the sentence, then the multiple translations of the core reference word are sorted according to the order in which the corresponding core reference word appears in the source language document to obtain the sequence of translations of the core reference word corresponding to the sentence; Based on the core reference word translation sequence corresponding to the sentence, the decoded word sequence, and the context vector at the current decoding moment, determine the decoding result at the current decoding moment.

9. The document translation method according to any one of claims 1 to 8, characterized in that, The process of translating each sentence in the source language document using the document attribute information corresponding to each sentence in the source language document to obtain the target language translation of the source language document includes: A structured document is constructed based on the source language document and the document attribute information corresponding to each sentence contained in the source language document. The structured document includes the structured information corresponding to each sentence contained in the source language document, and the structured information includes the corresponding sentence and the document attribute information corresponding to the corresponding sentence. The structured document is processed based on a pre-trained document translation model to obtain the target language translation corresponding to the source language document. The document translation model is trained using a structured document constructed based on the training source language document and the document attribute information corresponding to each sentence in the training source language document, as well as the actual target language translation corresponding to the training source language document.

10. A document translation device, characterized in that, include: Document acquisition module, document comprehension module, and document translation module; The document acquisition module is used to acquire source language documents to be translated; The document understanding module is used to obtain document attribute information corresponding to each sentence in the source language document, wherein the document attribute information is information determined based on the corresponding sentence and / or the context sentence of the corresponding sentence to help determine the meaning of the words in the corresponding sentence in the source language document; The document translation module is used to translate each sentence in the source language document, supplemented by the document attribute information corresponding to each sentence in the source language document, to obtain the target language translation of the source language document. The document attribute information includes one or more of the following attribute information: temporal attribute information, ambiguous attribute information, domain attribute information, and coreference attribute information; The temporal attribute information includes: the probability distribution of the corresponding sentence in each set temporal; The ambiguity attribute information includes: a set of deambiguous words composed of words obtained from the context of the corresponding sentence to help remove ambiguity from the corresponding sentence; The domain attribute information includes: the probability distribution of the corresponding sentence in each defined domain; The coreference attribute information includes a set of coreference words composed of words that have a coreference relationship with the sentence and are obtained from the context of the corresponding sentence.

11. A document translation device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the document translation method as described in any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the various steps of the document translation method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Coreference resolution in an ambiguity-sensitive natural language processing system

    CN101796508A

  • Method and system for translating text

    US20020091509A1