Translation method and device of long text, electronic equipment and readable storage medium

By identifying and matching the easily misunderstood words in the text to be translated, the problem of inconsistency in vocabulary translation in the existing technology is solved, and the translation quality and user experience are improved.

CN120218089APending Publication Date: 2025-06-27XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323470.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing translation systems process long texts, due to the limited window length, they cannot consider all context information at the same time, making it difficult for vocabulary to maintain the consistency of terms, logic and style during segmented translation.

Method used

By identifying the easily misleading words in the text to be translated, obtaining the matching vocabulary pairs, and translating the easily misleading words based on the translation content of the vocabulary pairs, ensuring the consistency of the translation.

Benefits of technology

It effectively overcomes the problem of inconsistency in translation, improves the quality of translation and user experience, and ensures the consistency of translation content in error-prone vocabulary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218089A_ABST
    Figure CN120218089A_ABST
Patent Text Reader

Abstract

The invention provides a long text translation method and device, electronic equipment and a readable storage medium, and the method comprises the steps: recognizing error-prone vocabularies in a to-be-translated text when the to-be-translated text is translated; a vocabulary pair is obtained according to the error-prone vocabulary, the vocabulary pair comprises to-be-translated content and translated content, and the to-be-translated content of the obtained vocabulary pair is matched with the error-prone vocabulary; and according to the obtained translation content of the vocabulary pair, translating the error-prone vocabulary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of text translation, and particularly relates to a translation method, apparatus, electronic device, and readable storage medium for long texts. Background Art

[0002] With the development of natural language processing and machine translation technologies, translation systems based on large-scale pre-trained language models (such as GPT-4) have performed excellently in multilingual translation. These models can handle short texts and single conversations, providing high-quality translations. However, when dealing with long text translation, due to the limited window length, for example, the context window of GPT-4 is 8192 tokens, these models cannot consider all context information simultaneously, resulting in difficulties in maintaining the consistency of terms, logic, and style for some vocabulary during segmented translation. For example, the professional term "API" may be translated as "Application Programming Interface" or "Interface" in different chapters, causing confusion for readers. Summary of the Invention

[0003] The present disclosure provides a translation method, apparatus, electronic device, and readable storage medium for long texts, which can effectively solve the above problems.

[0004] The present disclosure is implemented as follows:

[0005] In a first aspect, the present disclosure provides a translation method for long texts, the method comprising:

[0006] When translating a text to be translated, identifying error-prone vocabulary in the text to be translated;

[0007] Obtaining a vocabulary pair according to the error-prone vocabulary, the vocabulary pair including the content to be translated and the translation content, and the content to be translated of the obtained vocabulary pair matching the error-prone vocabulary;

[0008] Translating the error-prone vocabulary according to the translation content of the obtained vocabulary pair.

[0009] In a second aspect, the present disclosure provides a translation apparatus for long texts, the apparatus comprising:

[0010] An identification module, configured to identify error-prone vocabulary in the text to be translated when translating the text to be translated;

[0011] A matching module, configured to obtain a vocabulary pair according to the error-prone vocabulary, the vocabulary pair including the content to be translated and the translation content, and the content to be translated of the obtained vocabulary pair matching the error-prone vocabulary;

[0012] A translation module, configured to translate the error-prone vocabulary according to the translation content of the obtained vocabulary pair.

[0013] In a third aspect, the present disclosure provides an electronic device, including:

[0014] a memory that stores execution instructions; and

[0015] a processor that executes the execution instructions stored in the memory, such that the processor executes the method described in the first aspect.

[0016] In a fourth aspect, the present disclosure provides a readable storage medium storing execution instructions, which when executed by a processor are used to implement the method described in the first aspect.

[0017] Compared with the prior art, the beneficial effects of the present disclosure are as follows:

[0018] 1. The present disclosure provides a translation method for long texts. By identifying error-prone words that may have inconsistent translations in the text to be translated, matching them with pre-set word pairs, and then translating the error-prone words according to the translation contents of the matching word pairs, the consistency of the translation contents of the error-prone words throughout the translation process of the long text is ensured. This method effectively overcomes the problem of inconsistent translations caused by context limitations in existing translation systems, improving the translation quality and user experience.

[0019] 2. The method determines whether an error-prone word matches a word pair based on the similarity of embedding vectors. By capturing the deep semantic features of words in the context rather than relying on surface string matching, it can achieve the matching of error-prone words with different forms but similar semantics and word pairs, and distinguish homographs in the error-prone words. Moreover, the embedding vectors can be directly reused from the embedding layer of a pre-trained language model (such as BERT).

[0020] 3. The word pairs can be reused across texts, which can improve consistency and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of a long text translation method S100 provided by an embodiment of the present disclosure.

[0023] Figure 2 is a schematic structural diagram of a long text translation device 1000 provided by an embodiment of the present disclosure. Detailed implementation manners

[0024] The following further elaborates on the present disclosure in conjunction with the accompanying drawings and implementation manners. It can be understood that the specific implementation manners described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for ease of description, only parts related to the present disclosure are shown in the accompanying drawings.

[0025] It should be noted that, without conflict, the implementation manners in the present disclosure and the features in the implementation manners can be combined with each other. The technical solutions of the present disclosure will be elaborated in detail below with reference to the accompanying drawings and implementation manners.

[0026] Unless otherwise specified, the exemplary implementation manners / embodiments shown are understood to provide exemplary features of various details of some ways that can implement the technical concept of the present disclosure in practice. Therefore, unless otherwise specified, without departing from the technical concept of the present disclosure, the features of various implementation manners / embodiments can be additionally combined, separated, interchanged, and / or rearranged.

[0027] Embodiment 1

[0028] Please refer to Figure 1 , the embodiment of the present disclosure provides a translation method S100 for long texts.

[0029] Specifically, the method S100 includes:

[0030] S102, when translating the text to be translated, identify the error-prone words in the text to be translated;

[0031] S104, obtain a word pair according to the error-prone words, the word pair includes the content to be translated and the translated content, and the content to be translated of the obtained word pair matches the error-prone words;

[0032] S106, translate the error-prone words according to the translated content of the obtained word pair.

[0033] The long text refers to a text that cannot be translated within a context window, which can be a text with a number of tokens greater than the context window, or a cross-dialogue text. Even if the text of each dialogue is less than the context window, due to the limitation of the time span, there may be a problem of inconsistent translation.

[0034] In some implementation manners, the method further includes:

[0035] Segment the text to be translated to generate multiple text paragraphs.

[0036] In step S102, translating the text to be translated includes:

[0037] Translate each of the text paragraphs one by one.

[0038] The number of tokens in each text paragraph is not greater than the context window length of the large-scale pre-trained language model, so that each time of segmented translation can translate all the content of a text paragraph, and the translation content of a text paragraph has consistency.

[0039] For the text with existing natural paragraph divisions, use punctuation marks and line breaks for segmentation.

[0040] For long paragraphs or unstandardized text, use natural language processing tools (such as spaCy) for segmentation.

[0041] Also use natural language processing tools to perform sentence splitting and word segmentation on the text to be translated, for subsequent establishment of the position index of error-prone words and extraction of context text.

[0042] In step S102, identify the error-prone words in the text to be translated, including:

[0043] Use the named entity recognition (NER) function of the large-scale pre-trained language model to identify the error-prone words in the text to be translated.

[0044] In some embodiments, translate each text paragraph one by one, and while translating the text paragraph, identify the error-prone words in the text paragraph.

[0045] In some embodiments, the error-prone words are at least one of technical terms, abstract nouns, abbreviations, and culture-loaded words.

[0046] The error-prone words can also be added by the user for identification. Correspondingly, the user can also add word pairs that match the error-prone words.

[0047] The word pairs can be stored in a database for retrieval and matching.

[0048] In some embodiments, obtaining a word pair according to the error-prone word includes:

[0049] If the word pair cannot be obtained according to the error-prone word, directly translate the error-prone word and generate the word pair.

[0050] If no word pair matching the error-prone word is retrieved, a word pair matching the error-prone word can be generated according to the translation content of the error-prone word by the model and added to the database for retrieval, matching, and application during subsequent text translation.

[0051] The user can review and correct the generated word pairs.

[0052] In some embodiments, the word pair also records the position information of each error-prone word that matches it in the text to be translated and / or the translated text, so as to facilitate the user to confirm and compare the translation content and whether the matching association is correct.

[0053] The position of the error-prone word in the text to be translated and / or the translated text can be information such as paragraph number, full stop, etc., or directly the content of the original clause.

[0054] In some embodiments, after a word pair is updated, a backpropagation mechanism is triggered to batch correct the content associated with the word pair in the translated text.

[0055] These contents in the translated text can be located according to the position information associated with the word pair. Then, these contents in the translated text are corrected according to the updated translation content of the word pair.

[0056] Furthermore, a change log can be generated for the user to view.

[0057] Some error-prone words have homographs, so these error-prone words may match different word pairs corresponding to their different semantic meanings.

[0058] The word pairs in the database are dynamically added.

[0059] In some embodiments, the word pairs can be extracted by the user providing the text to be translated and its corresponding translated text to expand the database.

[0060] Alternatively, for a text to be translated, the user can select a specific database for the translation of the error-prone words therein.

[0061] Furthermore, the specific database adds word pairs by the user providing the text to be translated and its corresponding translated text.

[0062] By establishing a database in advance, existing translation content can be quickly matched and applied when processing new texts, reducing the workload and time of manual translation. Especially when processing texts in similar professional fields or with similar specific terms, it ensures the consistency and accuracy of translation and can provide higher-quality translation results.

[0063] In some embodiments, the word pair and the error-prone word are cross-text. That is, after adding and improving the word pair through one text, the database is used for the matching and translation of the error-prone words in another text, so as to achieve the consistency of terms, logic, and style between the cross-text translated texts.

[0064] In some embodiments, obtaining a word pair according to the error-prone word includes:

[0065] S1. Obtain the embedding vector of the error-prone word.

[0066] In some embodiments, step S1, obtaining the embedding vector of the error-prone word, includes:

[0067] S11. Obtain the context text of a preset length, where the context text includes the error-prone word.

[0068] The context text can be obtained by intercepting with a window of a predetermined length.

[0069] The context text needs to include the error-prone word.

[0070] For example, with the error-prone word as the center, expand N word segments or tags forward and backward respectively.

[0071] The window length does not exceed the maximum input length of BERT, or use overlapping interception of a sliding window.

[0072] Specifically, the overlapping ratio can be 25%-40%.

[0073] The window interception range can be limited to the clause where the error-prone word is located, or when the error-prone word is at the beginning or end of a clause, expand the window in the opposite direction to the adjacent clause to ensure that the window contains a complete semantic unit.

[0074] S12. Input the context text into the BERT model, and extract the vector representation of the error-prone word in the last layer hidden state of the BERT model.

[0075] In some embodiments, the BERT model is a pre-trained model fine-tuned for domain adaptation, and the domain includes at least one of technical documents, literary texts, or multilingual dialogue scenarios. By fine-tuning on data in a specific domain, the BERT model can extract more comprehensive deep semantic features, thereby generating embedding vectors more in line with the characteristics of the domain.

[0076] In some embodiments, for multi-word terms, such as "Artificial Intelligence", calculate the average of the embedding vectors of its constituent words.

[0077] S13. Perform layer normalization on the vector representation to obtain the embedding vector of the error-prone word.

[0078] The layer normalization process includes:

[0079] S131. Calculate the mean μ and variance σ of the vector representation.

[0080] S132. Through the formula Normalize, where ε is a numerical stability coefficient.

[0081] S133, perform a linear transformation on the normalized result γ and β are trainable parameters.

[0082] S2, according to the embedding vector of the error-prone word, obtain a word pair, and the content to be translated of the word pair matches the error-prone word.

[0083] That is, the word pair also records the embedding vector of its content to be translated.

[0084] At this time, the word pair can be recorded in the form of an example sentence.

[0085] Furthermore, when storing the semantic embedding vector, the position index of the content to be translated in the example sentence is also recorded for context alignment during semantic similarity calculation.

[0086] For example, the word segmentation result of the clause "The API provides programmatic access." is: ["The","API","provides","programmatic","access","."]. The position index corresponding to the target word "API" is (2,2) (counting from 1).

[0087] It is also possible to retrieve the translated text based on the embedding vector of the content to be translated of the word pair and locate the error-prone word that matches it.

[0088] For example, it can be used for content correction in the translated text after the word pair is updated.

[0089] S2, the content to be translated of the word pair matches the error-prone word, including:

[0090] The similarity between the embedding vector of the content to be translated of the word pair and the embedding vector of the error-prone word is greater than or equal to a preset threshold.

[0091] In some embodiments, the similarity is calculated by the cosine similarity algorithm.

[0092] If the similarity between the error-prone word and the existing word pair is low, indicating a large difference in word meaning, the error-prone word is stored as a new record as a homograph.

[0093] In step S104, according to the translation content of the obtained word pair, translate the error-prone word, which can be directly translating the error-prone word into the translation content, or adjusting the word order or sentence pattern according to the expression habit to make the sentence smooth.

[0094] The method provided by the embodiments of the present disclosure enhances the translation consistency of long texts, avoids the vocabulary translation differences caused by context breaks in traditional translations, is applicable to application scenarios such as multilingual translation and professional literature translation, and can improve the overall translation quality and user experience.

[0095] Embodiment 2

[0096] The embodiments of the present disclosure provide a translation device 1000 for long texts.

[0097] The device may include corresponding modules that execute each or several steps in the flowchart of the above-mentioned method S100. Therefore, each step or several steps in the above-mentioned flowchart can be executed by the corresponding modules, and the device may include one or more of these modules. The modules may be one or more hardware modules specifically configured to execute the corresponding steps, or implemented by a processor configured to execute the corresponding steps, or stored in a computer-readable medium for implementation by the processor, or implemented through a certain combination.

[0098] Specifically, as Figure 2 shown, the device 1000 includes:

[0099] An identification module 1002, configured to identify error-prone words in the text to be translated when translating the text to be translated;

[0100] A matching module 1004, configured to obtain a word pair according to the error-prone word, the word pair includes the content to be translated and the translated content, and the content to be translated of the obtained word pair matches the error-prone word;

[0101] A translation module 1006, configured to translate the error-prone word according to the translated content of the obtained word pair.

[0102] The embodiments of the present disclosure further provide an electronic device, including: a memory that stores execution instructions; and a processor or other hardware module, the processor or other hardware module executes the execution instructions stored in the memory, so that the processor or other hardware module executes the above-mentioned method.

[0103] The present disclosure also provides a readable storage medium, in which execution instructions are stored, and when the execution instructions are executed by a processor, they are used to implement the above-mentioned method.

[0104] Among them, the hardware structure adopted by the device implemented based on the hardware implementation method using a processor in the present disclosure can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus connects various circuits including one or more processors, memories, and / or hardware modules together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0105] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.

[0106] Any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong. The processor executes the various methods and processes described above. For example, the method embodiments in the present disclosure can be implemented as software programs, which are tangibly contained in a machine-readable medium, such as a memory. In some embodiments, part or all of the software program can be loaded and / or installed via the memory and / or communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the processor can be configured to execute one of the above methods in any other suitable manner (e.g., by means of firmware).

[0107] The logic and / or steps represented in the flowchart or described in other ways herein can be specifically implemented in any readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices.

[0108] For the purposes of this specification, a "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the readable storage medium include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable read-only memory (CDROM). Additionally, the readable storage medium can even be paper or other suitable medium on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or otherwise processing it as appropriate, and then storing it in a memory.

[0109] It should be understood that various parts of the present disclosure can be implemented in hardware, software, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0110] Those of ordinary skill in the art of the present technology can understand that all or part of the steps of implementing the above-described embodiments of the method can be completed by a program instructing relevant hardware. The program can be stored in a readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0111] Furthermore, in each of the various embodiments of the present disclosure, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above-described integrated modules can be implemented in the form of hardware or in the form of software functional modules. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0112] Those skilled in the art should understand that the above-described embodiments are merely for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations still fall within the scope of the present disclosure.

Claims

1. A method for translating a long text, characterized in that: The method comprises: When translating a text to be translated, identifying easy-to-make mistakes words in the text to be translated; Acquire a vocabulary pair according to the error-prone vocabulary, the vocabulary pair comprising content to be translated and translation content, the content to be translated of the acquired vocabulary pair matching the error-prone vocabulary; The error-prone words are translated according to the obtained translation content of the word pair.

2. The method according to claim 1, characterized in that The error-prone vocabulary is at least one of professional terms, abstract nouns, abbreviations, and culturally loaded words.

3. The method according to claim 1, characterized in that The vocabulary pair and the error-prone vocabulary are across texts.

4. The method according to claim 1, characterized in that Acquiring a vocabulary pair according to the error-prone vocabulary includes: If it fails to obtain the vocabulary pair according to the error-prone vocabulary, the error-prone vocabulary is directly translated to generate the vocabulary pair.

5. The method according to claim 1, characterized in that Acquiring a vocabulary pair according to the error-prone vocabulary includes: Obtaining an embedding vector of the error-prone vocabulary; According to the embedding vector of the error-prone vocabulary, a vocabulary pair is obtained, wherein the content to be translated of the vocabulary pair matches the error-prone vocabulary; The content to be translated of the vocabulary pair is matched with the error-prone vocabulary, including: The similarity between the embedding vector of the to-be-translated content of the vocabulary pair and the embedding vector of the error-prone vocabulary is greater than or equal to a preset threshold.

6. The method according to claim 5, characterized in that Obtaining the embedding vector of the error-prone vocabulary includes: Acquire a context text of a preset length, wherein the context text includes the error-prone vocabulary; Input the context text into a BERT model, and extract the vector representation of the error-prone vocabulary in the last hidden state of the BERT model; The vector representation is subjected to layer normalization processing to obtain an embedding vector of the error-prone vocabulary.

7. The method according to claim 1, characterized in that The method further comprises: Segmenting the text to be translated to generate multiple text paragraphs; Translate the text to be translated, including: The text passages are translated one by one.

8. A long text translation device, characterized in that: The device comprises: A recognition module, used for identifying easy-to-make mistakes in the text to be translated when translating the text to be translated; A matching module, used for acquiring a vocabulary pair according to the error-prone vocabulary, wherein the vocabulary pair includes content to be translated and translation content, and the content to be translated of the acquired vocabulary pair is matched with the error-prone vocabulary; The translation module is used to translate the error-prone vocabulary according to the translation content of the acquired vocabulary pair.

9. An electronic device, characterized in that: include: A memory storing execution instructions; as well as A processor, wherein the processor executes the execution instruction stored in the memory so that the processor executes the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores execution instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.