Document translation processing method and device, equipment and storage medium

By segmenting documents and obtaining contextual content, the problem of translation consistency in document translation is solved, achieving higher quality translation results.

CN120805942APending Publication Date: 2025-10-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510908951.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to ensure the consistency of translation of the same word or sentence throughout the entire document, especially when faced with massive and complex technical texts, making it difficult to meet high-quality translation requirements.

Method used

By segmenting the document to be translated, the contextual content of the segments is obtained, and this content and the original content of the segments are input into the translation model for translation processing. The far-domain and near-domain contextual content are combined to ensure translation consistency and semantic coherence.

Benefits of technology

It improves the consistency and accuracy of document translation, avoids semantic gaps caused by lack of context, and improves translation efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805942A_ABST
    Figure CN120805942A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a document translation processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-translated document, segmenting the to-be-translated document into a plurality of fragments, and obtaining the context content of a first fragment in the to-be-translated document, the context content comprises original text content and translated text content of a first preset number of text units adjacent to a first fragment in front of the to-be-translated document, inputting the context content and the original text content of the first fragment into a translation model, translating by the translation model, and outputting the translated text content of the first fragment, and determining a translation result document based on the translation content of the first fragment. The context content of the original text content and the translation content containing the first preset number of text units adjacent to the first fragment is supported to be obtained, the context content and the original text content of the first fragment are input into the translation model, and the first fragment is translated by combining the context content to obtain the translation content of the first fragment. And the translation consistency of translations is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a document translation processing method, apparatus, device, and storage medium. Background Art

[0002] With the development of global technology communication, the importance of document translation is increasing day by day. However, faced with massive and complex technical texts, relevant document translation technologies are unable to meet the growing demand for high-quality translation. For example, during the translation process, it is impossible to guarantee the same translation of the same word or sentence throughout the entire document.

[0003] Therefore, how to provide a document translation processing method to improve translation consistency has become an important technical problem that needs to be solved urgently. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a document translation processing method, apparatus, device and storage medium.

[0005] In a first aspect, an embodiment of the present disclosure provides a document translation processing method, the method comprising:

[0006] Obtaining a document to be translated, and segmenting the document to be translated into multiple fragments;

[0007] Obtaining contextual content of a first fragment among the multiple fragments in the document to be translated; wherein the contextual content includes original content and translated content of a first preset number of text units preceding and adjacent to the first fragment in the document to be translated;

[0008] Inputting the context content and the original content of the first segment into a translation model, and outputting the translated content of the first segment after the translation model translates the original content of the first segment;

[0009] Based on the translated content of the first segment, a translation result document corresponding to the document to be translated is determined.

[0010] In an optional embodiment, after inputting the context content and the original content of the first segment into a translation model, and after the translation model translates the original content of the first segment, and before outputting the translated content of the first segment, the method further includes:

[0011] In the translated fragment library of the document to be translated, a translated fragment having text similarity to the text of the first fragment satisfying a preset similarity condition is determined as a first translated fragment according to the original content of the first translated fragment, and the original content and the translated content of the first translated fragment are obtained; wherein the original content and the translated content of the translated fragment of the document to be translated are stored in the translated fragment library;

[0012] The method further comprises:

[0013] The method further comprises:

[0014] In an optional embodiment, before the context content and the original content of the first fragment are input into the translation model and the translated content of the first fragment is output after the translation processing of the original content of the first fragment by the translation model, the method further comprises:

[0015] The method further comprises:

[0016] If the word exists in the target word table, the translated content corresponding to the word is obtained.

[0017] The method further comprises:

[0018] The method further comprises:

[0019] In an optional embodiment, after the word existing in the target word table in the original content of the first fragment and the translated content corresponding to the word, the context content and the original content of the first fragment are input into the translation model and the translated content of the first fragment is output after the translation processing of the original content of the first fragment by the translation model, the method further comprises:

[0020] extracting a word pair from the original content and the translated content of the first segment, and storing the word pair into the target vocabulary; wherein the word pair comprises a word and a translation having a corresponding relationship.

[0021] In an optional implementation, after the original content and the translated content of the first translated segment, the context content, and the original content of the first segment are input into the translation model, the translation of the original content of the first segment is processed by the translation model, and the translated content of the first segment is output, the method further comprises:

[0022] storing the original content and the translated content of the first segment into the translated segment library of the document to be translated.

[0023] In an optional implementation, the context content further comprises original content of a second preset number of text units adjacent to the first segment in the document to be translated.

[0024] In an optional implementation, the text units in the context content comprise at least one of a segment, a paragraph, and a sentence.

[0025] In an optional implementation, the splitting of the document to be translated to obtain a plurality of segments comprises:

[0026] splitting the document to be translated to obtain a plurality of segments according to a document structure identifier in the document to be translated; wherein the document structure identifier is used to identify a document structure unit in the document to be translated.

[0027] In an optional implementation, the splitting of the document to be translated to obtain a plurality of segments according to a document structure identifier in the document to be translated comprises:

[0028] splitting the document to be translated into a plurality of chapter units according to a chapter identifier in the document to be translated.

[0029] splitting the chapter units into a plurality of paragraph units according to a paragraph identifier in the chapter unit and a preset length threshold; wherein a length of the paragraph unit is less than the preset length threshold.

[0030] determining the plurality of paragraph units as corresponding segments of the document to be translated respectively.

[0031] In an optional implementation, the method further comprises:

[0032] determine a professional term contained in the original content of the first segment according to a preset term table, wherein the preset term table stores the professional term and a translation of the professional term;

[0033] obtain the translation of the professional term contained in the original content of the first segment from the preset term table;

[0034] The inputting of the context content and the original content of the first segment into the translation model, the translation processing of the original content of the first segment by the translation model, and the outputting of the translation content of the first segment include that:

[0035] The inputting of the professional term contained in the original content of the first segment and the translation of the professional term, the context content, and the original content of the first segment into the translation model, the translation processing of the original content of the first segment by the translation model, and the outputting of the translation content of the first segment.

[0036] In a second aspect, the present disclosure provides a document translation processing apparatus, which comprises:

[0037] The segmentation module is configured to obtain a document to be translated, and segment the document to be translated to obtain a plurality of segments.

[0038] The first obtaining module is configured to obtain context content of a first segment in the plurality of segments in the document to be translated, wherein the context content comprises original content and translation content of a first preset number of text units adjacent to the front of the first segment in the document to be translated.

[0039] The translation processing module is configured to input the context content and the original content of the first segment into a translation model, perform translation processing of the original content of the first segment by the translation model, and output translation content of the first segment.

[0040] The first determining module is configured to determine a translation result document corresponding to the document to be translated based on the translation content of the first segment.

[0041] In a third aspect, the present disclosure provides an electronic device, which comprises a processor, a memory for storing executable instructions of the processor, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the document translation processing method provided by the embodiments of the present disclosure.

[0042] In a fourth aspect, the present disclosure provides a computer-readable storage medium, which stores a computer program for executing the document translation processing method provided by the embodiments of the present disclosure.

[0043] In a fifth aspect, the present disclosure provides a computer program product, which comprises computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the method described above.

[0044] Compared with the prior art, the technical solutions provided by the embodiments of the present disclosure have at least the following advantages:

[0045] In the document translation processing method provided by the embodiments of the present disclosure, the document to be translated is obtained, the document to be translated is segmented to obtain a plurality of segments, then, context content of a first segment in the plurality of segments in the document to be translated is obtained, wherein the context content comprises original content and translated content of a first preset number of text units adjacent to the front of the first segment in the document to be translated. Furthermore, the context content and the original content of the first segment are input into a translation model, and after the translation processing of the original content of the first segment by the translation model, the translated content of the first segment is output. After obtaining the translated content of each segment, based on the translated content of each segment, the translation result document corresponding to the document to be translated is determined.

[0046] In the process of document translation, the embodiments of the present disclosure support obtaining context content comprising original content and translated content of a first preset number of text units adjacent to the front of the first segment, and inputting the context content and the original content of the first segment into a translation model. The first segment is translated to obtain the translated content of the first segment in combination with the context content, which can improve the translation consistency of the translated content. BRIEF DESCRIPTION OF DRAWINGS

[0047] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.

[0048] Figure 1 A flowchart of a document translation processing method provided by an embodiment of the present disclosure is shown in FIG. 1;

[0049] Figure 2 A document structure diagram provided by an embodiment of the present disclosure is shown in FIG. 2;

[0050] Figure 3 A flowchart of another document translation processing method provided by an embodiment of the present disclosure is shown in FIG. 3;

[0051] Figure 4 A process diagram of a document translation processing method provided by an embodiment of the present disclosure is shown in FIG. 4;

[0052] Figure 5 A structure diagram of a document translation processing apparatus provided by an embodiment of the present disclosure is shown in FIG. 5;

[0053] Figure 6 A structural schematic diagram of an electronic device is provided for the embodiments of the present disclosure. DETAILED DESCRIPTION

[0054] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.

[0055] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0056] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related terms are defined in the following description.

[0057] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0058] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly stated in the context.

[0059] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely used for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0060] With the development of global technology dissemination, the importance of document translation work is increasing day by day, however, in the face of massive and complex technical texts, relevant document translation technologies are difficult to meet the growing demand for high-quality translation, for example, the uniform translation of the same word or sentence in the entire document cannot be guaranteed in the translation process.

[0061] Therefore, how to provide a document translation processing method to improve translation consistency has become an important technical problem to be solved at present.

[0062] To this end, the embodiment of the present disclosure provides a document translation processing method, obtaining a to-be-translated document, splitting the to-be-translated document to obtain a plurality of shards, and then obtaining context content of a first shard in the to-be-translated document, wherein the context content includes original content and translated content of a first preset number of text units adjacent to the first shard in the to-be-translated document. Further, the context content and the original content of the first shard are input into a translation model, and after translation processing of the original content of the first shard by the translation model, the translated content of the first shard is output. After obtaining the translated content of each shard, the translation result document corresponding to the to-be-translated document is determined based on the translated content of each shard.

[0063] In the process of document translation, the embodiment of the present disclosure supports obtaining context content containing original content and translated content of a first preset number of text units adjacent to the first shard, and inputting the context content and the original content of the first shard into a translation model. The translated content of the first shard is obtained by combining the context content for translation of the first shard, which can improve the translation consistency of the translated content.

[0064] Specifically, the embodiment of the present disclosure provides a document translation processing method, as shown in Figure 1 The flowchart of the document translation processing method provided by the embodiment of the present disclosure, which can be executed by a document translation processing device. The device can be implemented by software and / or hardware, and can be integrated in an electronic device. The document translation processing method provided by the embodiment of the present disclosure is applied in the document translation processing device. As shown in Figure 1 The method comprises the following steps.

[0065] S101: obtaining a to-be-translated document, and splitting the to-be-translated document to obtain a plurality of shards.

[0066] The to-be-translated document in the embodiment of the present disclosure refers to the original document that needs to be converted from a source language to a target language. Specifically, the to-be-translated document can be a long document with a large number of pages or a short document with a small number of pages. It is particularly suitable for long documents, such as patent documents, technical manuals, and papers. The source language can refer to the original language of the to-be-translated document, and the target language can be the language of the translated document, i.e. the language of the translation result document. For example, when the source language of the patent document is Chinese and the target language is English, the Chinese patent document is the to-be-translated document, and the English text obtained after translating the Chinese patent document is the translation result document.

[0067] In the embodiments of the present disclosure, after obtaining the document to be translated, the document to be translated is split to obtain a plurality of fragments, which can refer to splitting the document to be translated into a plurality of independent text units according to a splitting rule. Each split text unit can be a fragment.

[0068] In an optional embodiment, the document to be translated can be split based on a document structure identifier. The document structure identifier is used to identify a document structure unit in the document to be translated. The document structure unit can be located in the document to be translated by the document structure identifier, and the start and end positions of the document structure unit are defined. The document structure identifier can represent the hierarchical structure of the document to be translated (such as the nesting relationship between chapters, titles, and paragraphs). The document structure identifier can include title identifier, chapter identifier, paragraph identifier, etc. The document structure unit is a content segment obtained by dividing the hierarchical structure of the document to be translated. Specifically, the document structure unit can include title unit, chapter unit, paragraph unit, sentence unit, etc.

[0069] Specifically, after obtaining the document to be translated, the document structure identifier in the document to be translated is first identified. Then, according to the hierarchical relationship between the document structure identifiers, the document to be translated is divided into a plurality of document structure units, and each document structure unit is taken as a fragment of the document to be translated. The fragment can refer to a document to be translated unit obtained by dividing the document to be translated according to the document structure identifier. Specifically, the fragment can include chapter fragment, paragraph fragment, etc.

[0070] On the basis of the above embodiments, the fragments of the document to be translated can be obtained by splitting according to the chapter identifier. Specifically, the chapter identifier in the document to be translated is identified. The document to be translated is split into a plurality of chapter units according to the chapter identifier, and the plurality of chapter units are determined as the corresponding fragments of the document to be translated.

[0071] When translating with chapter units as fragments, the chapter units are relatively long, which leads to the need to process a large amount of text data during translation, resulting in slow translation efficiency. Therefore, on the basis of the above embodiments, the present disclosure can further split the chapter units according to the paragraph identifier to obtain paragraph units as the corresponding fragments of the document to be translated, so as to obtain finer-grained fragments and improve translation efficiency. Specifically, after splitting the document to be translated into a plurality of chapter units according to the chapter identifier, each chapter unit is split into a plurality of paragraph units according to the paragraph identifier in each chapter unit, and the plurality of paragraph units are determined as the corresponding fragments of the document to be translated, i.e., the paragraph units are taken as the corresponding fragments of the document to be translated (also referred to as paragraph fragments).

[0072] Since a paragraph is usually developed around a theme, the way of segmenting the chapter units in the translation document according to the paragraph can preserve the logical relationship within the paragraph, reduce the probability of misinterpretation caused by semantic rupture, that is, by first segmenting the chapter units according to the chapter identifier, and then segmenting the paragraph segments according to the paragraph identifier, the semantic coherence during translation can be ensured, thereby improving the accuracy of translation.

[0073] On the basis of the above embodiment, since the first segment is subsequently translated by using the translation model, the translation model has a limit on the length of the input text, and too long input text will cause the performance of the translation model to decline, therefore, in order to improve the translation performance of the translation model, the length of the paragraph unit can also be limited in the embodiments of the present disclosure, specifically, after the translation document is segmented into multiple chapter units according to the chapter identifier, each chapter unit is segmented into multiple paragraph units according to the paragraph identifier in each chapter unit and a preset length threshold, and the multiple paragraph units are determined as the segments corresponding to the translation document respectively, wherein the length of the paragraph unit is less than the preset length threshold, and the preset length threshold is the maximum length of the text unit set in advance when the document is segmented, that is, the preset length threshold is the length limit of the segment, and the preset length threshold can be set based on the capability of the translation model and the translation demand.

[0074] After the translation document is segmented into multiple segments corresponding to the translation document according to the document structure identifier, a segment object corresponding to each segment can also be constructed, so as to facilitate the assembly of the translation content of each segment according to the arrangement order of the translation document. Specifically, the segment object includes segment identifier, original content, chapter to which the segment belongs, and whether it belongs to a title segment and the like, the segment identifier is used to mark the position of the segment in the translation document; the original content refers to the original content contained in the segment, the chapter to which the segment belongs is used to mark the chapter level to which the segment belongs, and can represent the hierarchical structure between the segments; whether it belongs to a title segment is used to distinguish the segment type (such as a title segment, a paragraph segment, etc.) and define the chapter boundary, the title segment includes a document title segment, a chapter title segment, etc., and the definition of the chapter boundary means that by determining whether the segment belongs to a title segment, the starting and ending positions of different chapters in the translation document can be identified. If a segment is a title segment, the segment can be translated by using other translation processing methods (such as manual translation). The segment object can be constructed in the Json format.

[0075] For example, the slice object A includes the following information: {“ID_1”:“1”,“original content”:“XXXXX”,“slice belonging chapter”:“claims”,“whether it belongs to the title slice”:“yes”}, where ID_1 corresponds to the value 1, indicating that the slice is in the first position in the document to be translated; the original content corresponds to the value of the original content of the slice; the slice chapter corresponds to the value of“claims”, indicating that the slice belongs to the chapter“claims”; whether it belongs to the title slice corresponds to the value of yes, indicating that the slice is the title slice of“claims”.

[0076] It can be seen that by constructing the slice object, the translation result document can be assembled according to the chapter, paragraph and other document structures of the document to be translated when the translated text of the slice is assembled subsequently.

[0077] S102: Obtain the context content of the first slice in the plurality of slices in the document to be translated.

[0078] The context content includes the original content and the translated content of the first preset number of text units adjacent to the first slice in the document to be translated.

[0079] The first slice in the embodiment of the disclosure can refer to any slice that is currently being translated.

[0080] In order to ensure the consistency of slice translation, the embodiment of the disclosure further obtains the context content of the first slice according to its position in the document to be translated when translating the first slice after completing the slice processing of the document to be translated. The context content includes the original content and the translated content of the first preset number of text units adjacent to the first slice in the document to be translated, that is, the context content can include the context content of the first slice, and the context content constitutes the near-field context. The first preset number can be a pre-set value, and the first preset number is used to set the number of text units in the context content adjacent to the first slice. For example, if the first slice is the fourth paragraph, the first preset number is 3, and the context content includes the original content and the translated content of paragraphs 1-3. It can be seen that in the document translation scenario, dynamically setting the number of text units in the context content of the first slice can expand the context range on demand and adapt to different document translation requirements.

[0081] In an optional embodiment, the text units in the context content can include at least one of the slices, paragraphs, and sentences, that is, the context content of the first slice can be at least one of the first preset number of slices, paragraphs, and sentences adjacent to the first slice in the document to be translated.

[0082] On the basis of the above-mentioned embodiments, while improving the translation consistency of the translated text, the context content in the embodiments of the present disclosure can also include the original content of the second preset number of text units adjacent to the first segment in the document to be translated, the second preset number being a preset value, the second preset number being used to set the number of text units in the context content adjacent to the first segment, the first preset number and the second number can be the same or different. For example, if the first segment is the fourth paragraph, the first preset number is 3, and the second preset number is 2, the context content of the first segment includes the original content and the translated content of paragraphs 1-3, and the context content of the first segment includes the original content of paragraphs 5-6. That is, the context content of the first segment includes the context content before and after the first segment, wherein the text units in the context content before the first segment include the original content and the translated content, and the text units in the context content after the first segment include the original content.

[0083] As can be seen, when translating the first segment, the translation consistency of the translated text can be ensured by obtaining the context content thereof. In addition, by combining the context content before and after the current segment to translate the current segment, the semantic coherence of the translated text can be improved, an effective context can be provided for the subsequent text, and the semantic discontinuity caused by the lack of context during translation can be avoided, thereby improving the accuracy of translation.

[0084] In an optional embodiment, when obtaining the context content of the first segment, if the context content before the first segment cannot be obtained in the document to be translated or the number of text units of the context content before the first segment does not meet the first preset number (for example, the first segment is the first segment), only the context content after the first segment can be obtained for subsequent determination of the translated content of the first segment. Correspondingly, if the context content after the first segment cannot be obtained in the document to be translated or the number of text units of the context content after the first segment does not meet the second preset number (for example, the first segment is the last segment), only the context content before the first segment can be obtained for subsequent determination of the translated content of the first segment.

[0085] S103: input the context content and the original content of the first segment into a translation model, and output the translated content of the first segment after translation processing of the original content of the first segment by the translation model.

[0086] In the embodiments of the present disclosure, after the context content of the first segment is obtained, the context content and the original content of the first segment can be input into a translation model to determine the translation content of the first segment. Specifically, the original content and the translation content of the context content and the original content of the first segment can be spliced, and then the spliced context content and the original content of the first segment are input into the translation model. The translation model is a trained model with translation function, and the input parameters of the translation model can include the original content of the first segment and the context content before the first segment. The output parameter of the translation model is the translation content of the first segment.

[0087] Before the original content of the first segment and the context content thereof are input into the translation model as input parameters, the first segment and the context content thereof can also be spliced according to a splicing rule, and then input into the translation model after splicing. The splicing rule is a pre-set rule, and the embodiments of the present disclosure do not limit the splicing rule. The splicing rule can include a rule of splicing according to the position sequence, for example, [original content of the first N text units][translation content of the first N text units][first segment], wherein the text unit can be a segment, a paragraph or a sentence, and N can be set according to requirements. It can be seen that splicing the first segment and the context content thereof according to the splicing rule can further enhance the semantic integrity and improve the semantic coherence of the translation, thereby improving the accuracy of translation.

[0088] S104: determining a translation result document corresponding to the document to be translated based on the translation content of the first segment.

[0089] The translation result document corresponding to the document to be translated refers to a document obtained by assembling the translation content of each segment of the document to be translated.

[0090] Before obtaining the translation result document, the translation content of each segment needs to be obtained, and then the translation content of each segment is spliced according to an assembly condition to obtain the translation result document. The assembly condition can include the position sequence of the segment, the segment identifier, etc. The translation content of each segment is obtained in the same way as the translation content of the first segment, and the translation content of each segment can be obtained by distributed or parallel translation.

[0091] In the document translation processing method provided by the embodiments of the present disclosure, a document to be translated is obtained, the document to be translated is segmented to obtain a plurality of segments, then, context content of a first segment in the plurality of segments in the document to be translated is obtained, wherein the context content includes original content and translated content of a first preset number of text units adjacent to the first segment in the document to be translated. Furthermore, the context content and the original content of the first segment are input into a translation model, and after translation processing of the original content of the first segment by the translation model, translated content of the first segment is output. After obtaining the translated content of each segment, based on the translated content of each segment, a translation result document corresponding to the document to be translated is determined.

[0092] In the process of document translation, the embodiments of the present disclosure support obtaining context content including original content and translated content of a first preset number of text units adjacent to the first segment, and inputting the context content and the original content of the first segment into a translation model. The first segment is translated to obtain translated content of the first segment in combination with the context content, which can improve the translation consistency of the translated content.

[0093] In another optional implementation, before the original content of the first segment and its context content are input into the translation model, and after translation processing of the original content of the first segment by the translation model, the translated content of the first segment is output, the first translated segment can also be retrieved based on text similarity, to determine the translated content of the first segment in combination with the original content and the context content of the first segment, so as to further improve the semantic coherence of the translated content on the basis of ensuring translation consistency.

[0094] The first translated segment constitutes a far-domain context, that is, it is across paragraphs or chapters from the first segment, and is different from the context content (also referred to as a near-domain context) of the first segment. The first translated segment is a translated segment whose text similarity with the first segment meets a preset similarity condition. The preset similarity condition is a similarity threshold value that is set in advance and is used to judge the matching degree between the content in the translated segment and the first segment. That is, if the text similarity between the current translated segment and the first segment meets the similarity threshold value, the current translated segment is taken as the first translated segment, otherwise the next translated segment is compared, until the last translated segment is traversed.

[0095] Specifically, when translating the first segment, the text similarity between the original content of each translated segment and the original content of the first segment is calculated in the translated segment library of the document to be translated. Based on the text similarities, the translated segment whose text similarity with the original content of the first segment meets the preset similarity criteria is selected as the first translated segment, and its original content and translated content are obtained. The translated segment library is a database or storage structure used to store the original content of the translated segments of the document to be translated and their corresponding translations. Each translated segment includes two parts of data: the original content and the translated content. The original content refers to the original text content in the document to be translated.

[0096] In other words, the text similarity between the translated segments and the original content of the first segment is compared, and the translated segments whose text similarity exceeds a preset similarity threshold are selected as the first translated segments, and their original content and translation content are obtained. The text similarity between the original content of the translated segments and the original content of the first segment can be calculated using relevant algorithms based on multiple dimensions such as word frequency, keyword matching, and semantics.

[0097] like Figure 2 The figure shows a schematic diagram of a document structure provided by an embodiment of the present disclosure. As can be seen from the figure, the context content is a preset number of text units directly adjacent to the first fragment, which constitute the near-domain context. The first translated fragment is not adjacent to the first fragment, and the first translated fragment spans a chapter or paragraph, so the first translated fragment constitutes the far-domain context.

[0098] After obtaining the original and translated content of the first translated segment, the original and translated content of the first translated segment, the context, and the original content of the first segment are input into the translation model. After the translation model translates the original content of the first segment, the translated content of the first segment is output. For details, please refer to the above embodiment.

[0099] It can be seen that when translating the first fragment, long-distance (across paragraphs or chapters) semantic associations can be achieved by retrieving the far-domain context, improving the overall logical coherence of the translation and avoiding long-distance semantic gaps. In addition, through far-domain context retrieval, the translation of the same word within the entire document can be accurately matched and reused, ensuring the translation consistency of the same word across chapters and fragments, that is, the same word, the same translation. In other words, the method of retrieving the far-domain context in the embodiment of the present disclosure can improve translation accuracy from the two dimensions of semantic coherence and translation consistency.

[0100] After the original and translated content of the first translated segment, along with the context and the original content of the first segment, is input into the translation model, and after the translation model translates the original content of the first segment and outputs the translated content of the first segment, the original and translated content of the first segment can be stored in a translated segment library to facilitate retrieval of translated segments whose textual similarity with the original content of the segment meets a preset similarity condition when translating other segments. This allows for rapid retrieval of the far-domain context based on textual similarity in subsequent translations, thereby improving translation efficiency.

[0101] In another optional embodiment, before determining the translated content of the first segment by combining the original content of the first segment and its context, the target vocabulary may be retrieved for words contained in the original content of the first segment to determine the translated content of the first segment. The target vocabulary is a database or storage structure for storing words and their corresponding translations. The target vocabulary may be generated in real time during the translation process. Optionally, the target vocabulary may be automatically destroyed after the translation of the entire document is completed to save storage space.

[0102] Specifically, first identify and extract words in the original content of the first segment, and search whether the above words exist in the target vocabulary. When the word in the original content of the first segment is found in the target vocabulary, obtain the corresponding translation of the word from the target vocabulary.

[0103] After obtaining the words in the target vocabulary and the corresponding translations in the original content of the first segment, the words and their translations, the context content, and the original content of the first segment can be input into the translation model. After the translation model translates the original content of the first segment, the translated content of the first segment is output.

[0104] As can be seen, when translating the first segment, searching the target vocabulary for words and their corresponding translations in the original text of the first segment can further increase the reuse rate of the same words during the translation process, ensure translation consistency, and thus improve translation accuracy. Retrieving words from the original text of the segment using the target vocabulary is particularly suitable for long-distance translation scenarios (across paragraphs or chapters).

[0105] After the words in the original content of the first segment that exist in the target vocabulary and their corresponding translations, the context content, and the original content of the first segment are input to the translation model, and after the translation processing of the original content of the first segment by the translation model, the translation content of the first segment is output, in order to facilitate the retrieval of the words in the original content of the segment in the target vocabulary when translating other segments subsequently, the word pairs can also be extracted from the original content and the translation content of the first segment, and the extracted word pairs are stored in the target vocabulary, wherein the word pairs include words and translations having a corresponding relationship, and the word pairs are new word pairs that do not exist in the target vocabulary. That is, the new word pairs that exist in the original content and the translation content of the first segment are stored in the target vocabulary, which can facilitate the reuse of retrieval when translating other segments subsequently, improve the translation efficiency, and also can unify the vocabulary translation method to ensure the consistency of translation and improve the accuracy of translation.

[0106] In order to facilitate the understanding of the content of the above-mentioned embodiments, the disclosure embodiments also provide a document translation processing method, which is described with reference to the above-mentioned embodiments. Figure 3 A flowchart of a document translation processing method provided by the embodiments of the disclosure is shown in the figure, and the method can be executed by a document translation processing device. The device can be implemented by software and / or hardware, and can be integrated in an electronic device.

[0107] As shown in the figure, the method includes the following steps. Figure 3

[0108] S301: Obtain a document to be translated, and divide the document to be translated into a plurality of segments according to document structure identifiers in the document to be translated.

[0109] The document structure identifier is used to identify a document structure unit in the document to be translated.

[0110] The content of S301 can be understood with reference to the content of the above-mentioned embodiments, which will not be repeated here.

[0111] S302: Obtain the context content of the first segment in the document to be translated, which includes the original content and translation content of the first preset number of text units adjacent to the front of the first segment in the document to be translated, and the original content of the second preset number of text units adjacent to the rear of the first segment.

[0112] ​After the completion of the segmentation of the document to be translated, in order to ensure the consistency of the translation, and to enable the first segment to accurately take over the previous semantics, provide an effective context for the subsequent text, and avoid semantic ambiguity caused by the lack of context during translation, the embodiments of the present disclosure further acquire the context content of the first segment according to its position in the document to be translated when translating the first segment. Wherein, the context content constitutes the near-domain context. Wherein, the first preset number and the second preset number are pre-set values.

[0113] In an optional implementation, the text units in the context content can include at least one of segments, paragraphs, and sentences.

[0114] S303: In the translated segment library of the document to be translated, the translated segment whose text similarity with the first segment meets the preset similarity condition is determined according to the original text content of the translated segment, and the translated segment is taken as the first translated segment, and the original text content and the translated content of the first translated segment are acquired.

[0115] Wherein, the original text content and the translated content of the translated segment in the translated segment library are stored.

[0116] The first translated segment in the embodiments of the present disclosure constitutes the far-domain context, that is, it is across paragraphs or chapters with the first segment, which is different from the context content (also referred to as the near-domain context) of the first segment. The first translated segment is a translated segment whose text similarity with the first segment meets the preset similarity condition, and the preset similarity condition is a pre-set similarity threshold, which is used to judge the matching degree between the content in the translated segment and the first segment.

[0117] S304: Extract the words in the original text content of the first segment, and determine whether the words exist in the target word table. Wherein, the target word table stores a plurality of words and a plurality of words respectively corresponding to the translated content.

[0118] S305: If the words in the original text content of the first segment exist in the target word table, the translated content corresponding to the words is acquired.

[0119] The contents of S304-S305 can be understood with reference to the contents in the above embodiments, which will not be repeated here.

[0120] S306: The words in the original text content of the first segment that exist in the target word table and the translated content corresponding to the words, the original text content and the translated content of the first translated segment, the context content, and the original text content of the first segment are input to the translation model. After the translation processing of the original text content of the first segment by the translation model, the translated content of the first segment is output.

[0121] Before the above content is input into the translation model, the words in the first segment of the original content that exist in the target vocabulary and the corresponding translation, the original content and the translation content of the first translated segment, the context content, and the original content of the first segment can also be spliced according to the splicing rule and input into the translation model after splicing. The splicing rule is a pre-set rule, which can be set based on demand. Specifically, the splicing rule can be: the original and translation of the first translated segment, the context content of the first segment, which includes the original content and translation content of the context content and the original content of the subsequent content, the words in the original content of the first segment that exist in the target vocabulary and the corresponding translation, and the original content of the first segment.

[0122] In an optional implementation, before the original content of the first segment is translated, a preset terminology table can also be obtained, and the terminology pairs (professional terms and their corresponding translations) in the preset terminology table are also input into the translation model to assist in translating the first segment and improve the consistency of the full-text terminology translation. The original and translation of the professional terms stored in the preset terminology table are obtained in advance, and the preset terminology table is a pre-set professional terminology table.

[0123] Specifically, first, the professional terms contained in the original content of the first segment are determined according to the preset terminology table, and then the translation of the professional terms contained in the original content of the first segment is obtained from the preset terminology table. After obtaining the translation of the above professional terms, the professional terms contained in the original content of the first segment and the translation of the professional terms, the context content, and the original content of the first segment are input into the translation model, and after the translation model processes the translation of the original content of the first segment, the translation content of the first segment is output.

[0124] S307: Based on the translation content of the first segment, determine the translation result document corresponding to the document to be translated.

[0125] The translation result document corresponding to the document to be translated refers to the document obtained by assembling the translation content of each segment of the document to be translated.

[0126] Before obtaining the translation result document, the translation content of each segment needs to be obtained, and then the translation content of each segment is assembled according to the assembly condition to obtain the translation result document, and the assembly condition can include the position order of the segment, the segment identifier, etc. The translation content of each segment is obtained in the same way as the translation content of the first segment, which can be obtained by distributed or parallel translation.

[0127] In the process of document translation, the embodiment of the present disclosure first performs structured segmentation on the document to be translated according to the document structure identification, so as to improve the semantic integrity of the segmented fragments. In addition, the current fragment is translated by combining the context content (near-domain context) before and after the current fragment, the far-domain context, and the words in the original text content of the first fragment that exist in the target vocabulary and the corresponding translation of the words, that is, the semantic coherence of the translation is improved, and the consistency of the vocabulary translation in the document to be translated is ensured. It can be seen that the embodiment of the present disclosure improves the accuracy of translation from two dimensions of semantic coherence and translation consistency.

[0128] In order to facilitate understanding of the content of the above-mentioned embodiment, the present disclosure will take fragment A as an example to introduce the document translation processing method provided by the present disclosure, as shown in Figure 4 The specific process is as follows:

[0129] Firstly, when translating the original text content of the fragment A to be translated, the near-domain context thereof is retrieved, and specifically, the context content adjacent to the fragment A before and after is obtained from the document to be translated, wherein the text units in the context content before and after include the original text content and the translation content, and the text units in the context content after and before include the original text content.

[0130] After obtaining the near-domain context of the fragment A, the far-domain context thereof is retrieved, and specifically, the text similarity between the original text content of the fragment A and each translated fragment in the translated fragment library is calculated, the translated fragment in the translated fragment library that satisfies the preset similarity condition with the original text content of the fragment A is retrieved, and the original text content and the translation content of the translated fragment are obtained.

[0131] After obtaining the far-domain context of the fragment A, it is determined whether the words contained in the original text content of the fragment A exist in the vocabulary M, and specifically, the words contained in the original text content of the fragment A are extracted, and it is determined whether the extracted words exist in the vocabulary M. If the words contained in the original text content of the fragment A exist in the vocabulary M, the existing words and the corresponding translation are extracted.

[0132] After obtaining the near-domain context, the far-domain context, and the words contained in the original text content of the fragment A and the corresponding translation in the vocabulary M, the above three and the original text content of the fragment A are input into the translation model for translation processing, and the translation content of the fragment A is output.

[0133] After obtaining the translation content of the fragment A, the original text content and the translation content of the fragment A are stored in the translated fragment library and the fragment list, and the word pairs are extracted from the original text content and the translation content of the fragment A. The extracted word pairs are stored in the vocabulary M.

[0134] In the process of document translation, the embodiment of the present disclosure first performs structural segmentation on the document to be translated according to a document structure identification, so as to improve the semantic integrity of the segmented fragments. In addition, the current fragment is translated by combining the context content (near-domain context) before and after the current fragment, the far-domain context, and the words in the target vocabulary and the corresponding translation of the words in the original content of the first fragment, that is, the semantic coherence of the translation is improved, and the consistency of the vocabulary translation in the document to be translated is ensured. It can be seen that the embodiment of the present disclosure improves the accuracy of translation from two dimensions of semantic coherence and translation consistency.

[0135] In order to realize the above-mentioned embodiment, the present disclosure further provides a document translation processing device. Figure 5 A structural schematic diagram of a document translation processing device provided by the embodiment of the present disclosure is shown in the figure. The device can be realized by software and / or hardware, and can be integrated in an electronic device. As shown in the figure, the device comprises: Figure 5

[0136] The segmentation module 501 is configured to obtain a document to be translated, and segment the document to be translated to obtain a plurality of fragments.

[0137] The first acquisition module 502 is configured to acquire context content of a first fragment in the plurality of fragments in the document to be translated. The context content comprises original content and translation content of a first preset number of text units adjacent to the front of the first fragment in the document to be translated.

[0138] The translation processing module 503 is configured to input the context content and the original content of the first fragment into a translation model, and output the translation content of the first fragment after translation processing of the original content of the first fragment by the translation model.

[0139] The first determination module 504 is configured to determine a translation result document corresponding to the document to be translated based on the translation content of the first fragment.

[0140] In an optional embodiment, the device further comprises:

[0141] The second determination module is configured to determine, in a translated fragment library of the document to be translated, a translated fragment that meets a preset similarity condition with the text of the first fragment as a first translated fragment according to the original content, and acquire the original content and the translation content of the first translated fragment. The original content and the translation content of the translated fragment in the translated fragment library are stored.

[0142] The translation processing module is specifically configured to:

[0143] ​input the context content and the first segment of the original content into a translation model, and output the first segment of the translated content after translation processing of the first segment of the original content by the translation model.

[0144] In an optional implementation, the apparatus further includes:

[0145] a third determination module configured to extract a word in the first segment of the original content and determine whether the word exists in a target vocabulary; wherein the target vocabulary stores a plurality of words and translated content corresponding to the plurality of words respectively;

[0146] a second acquisition module configured to acquire the translated content corresponding to the word when the word exists in the target vocabulary;

[0147] The translation processing module is specifically configured to:

[0148] input the word existing in the target vocabulary in the first segment of the original content and the translated content corresponding to the word, the context content, and the first segment of the original content into a translation model, and output the first segment of the translated content after translation processing of the first segment of the original content by the translation model.

[0149] In an optional implementation, the apparatus further includes:

[0150] a first storage module configured to extract a word pair from the first segment of the original content and the translated content, and store the word pair in the target vocabulary; wherein the word pair includes a word and translated content having a corresponding relationship.

[0151] In an optional implementation, the apparatus further includes:

[0152] a second storage module configured to store the first segment of the original content and the translated content in a translated segment library of the document to be translated.

[0153] In an optional implementation, the context content further includes original content of a second preset number of text units adjacent to the first segment in the document to be translated.

[0154] In an optional implementation, the text units in the context content include at least one of a segment, a paragraph, and a sentence.

[0155] In an optional implementation, the segmentation module is specifically configured to:

[0156] According to a document structure identifier in the document to be translated, the document to be translated is segmented into a plurality of segments; wherein the document structure identifier is used to identify a document structure unit in the document to be translated.

[0157] In an alternative embodiment, the segmentation module comprises:

[0158] The first segmentation sub-module is configured to segment the document to be translated into a plurality of chapter units according to chapter identifiers in the document to be translated.

[0159] The second segmentation sub-module is configured to segment the chapter units into a plurality of paragraph units according to paragraph identifiers in the chapter units and a preset length threshold; wherein the length of the paragraph units is less than the preset length threshold.

[0160] The first determination sub-module is configured to determine the plurality of paragraph units as the segments corresponding to the document to be translated, respectively.

[0161] In an alternative embodiment, the apparatus further comprises:

[0162] The fourth determination module is configured to determine professional terms contained in the original content of the first segment according to a preset terminology table; wherein the preset terminology table stores professional terms and translations of the professional terms.

[0163] The third acquisition module is configured to acquire the translations of the professional terms contained in the original content of the first segment from the preset terminology table.

[0164] The translation processing module is specifically configured to:

[0165] input the professional terms contained in the original content of the first segment and the translations of the professional terms, the context content, and the original content of the first segment into a translation model, and output the translation content of the first segment after translation processing of the original content of the first segment by the translation model.

[0166] In the document translation processing apparatus provided by the embodiments of the present disclosure, a document to be translated is acquired, the document to be translated is segmented into a plurality of segments, then, context content of a first segment in the plurality of segments in the document to be translated is acquired, wherein the context content includes original content and translation content of a first preset number of text units adjacent to the first segment in the document to be translated. Furthermore, the context content and the original content of the first segment are input into a translation model, and the translation content of the first segment is output after translation processing of the original content of the first segment by the translation model. After obtaining the translation content of each segment, a translation result document corresponding to the document to be translated is determined based on the translation content of each segment.

[0167] In the document translation process, the embodiment of the present disclosure supports obtaining context content containing original text content and translated text content of a first preset number of text units adjacent to a first split piece, and inputting the context content and the original text content of the first split piece into a translation model to translate the first split piece to obtain translated text content of the first split piece, which can improve the translation consistency of the translated text.

[0168] In addition to the above method and device, the embodiment of the present disclosure further provides a computer readable storage medium, which stores instructions, and when the instructions are run on a terminal device, the terminal device implements the document translation processing method provided by the embodiment of the present disclosure.

[0169] The embodiment of the present disclosure further provides a computer program product, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the document translation processing method provided by the embodiment of the present disclosure.

[0170] In addition, the embodiment of the present disclosure further provides an electronic device, as shown in Figure 6 may include:

[0171] The processor 601, the memory 602, the input device 603 and the output device 604. The number of processors 601 in the electronic device can be one or more, Figure 6 In some embodiments of the present disclosure, the processor 601, the memory 602, the input device 603 and the output device 604 can be connected by a bus or other means, wherein, Figure 6 In some embodiments of the present disclosure, the processor 601, the memory 602, the input device 603 and the output device 604 can be connected by a bus or other means, wherein,

[0172] The memory 602 can be used to store software programs and modules, and the processor 601 executes various functions of the electronic device and the document translation processing by running the software programs and modules stored in the memory 602. The memory 602 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required by a function, etc. In addition, the memory 602 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The input device 603 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the electronic device.

[0173] In particular, in the present embodiment, the processor 601 will load the executable file corresponding to the process of one or more application programs into the memory 602 according to the following instructions, and run the application program stored in the memory 602 by the processor 601, thereby realizing the various functions of the above-mentioned electronic device.

[0174] It has to be noted that, in the present document, relational terms are intended to encompass the various possible relationships between means or components or steps in a process. For example, the phrase "first means and second means" is intended to mean that a first means and a second means can be active together, can be active sequentially, or can be active at the same time, depending upon the context in which the phrase is used. The phrase "first means and second means" is not used to require that both the first and second means be present at the same time, or even that they be present at all, unless the context clearly indicates otherwise. The phrase "first means and / or second means" is intended to mean either the first means alone, the second means alone, or both the first and second means. Similarly, the phrases "at least one of the first means and the second means" and "one or more of the first means and the second means" are each intended to mean either the first means alone, the second means alone, or both the first and second means. Also, the phrase "at least one of the first means and the second means" is intended to mean that the first means or the second means is present, but not necessarily both, unless the context clearly indicates otherwise. The phrase "first means and the second means" is intended to mean either the first means or the second means, but not both, unless the context clearly indicates otherwise. The phrase "one or more of the first means and the second means" is intended to mean either the first means or the second means, but not both, unless the context clearly indicates otherwise. The phrase "one or more of the first means and / or the second means" is intended to mean either the first means alone, the second means alone, or both the first and second means. Similarly, the phrase "one or more of the first means and the second means, and / or one or more other means" is intended to mean either the first means alone, the second means alone, the other means alone, or any combination of the first means, the second means, and the other means. The phrase "at least one of the first means and / or the second means" is intended to mean either the first means alone, the second means alone, or both the first and second means. The phrase "at least one of the first means and the second means" is intended to mean either the first means or the second means, but not both, unless the context clearly indicates otherwise.

[0175] The foregoing is merely illustrative of the various ways and specific embodiments in which the disclosure can be carried out. Numerous modifications can be made to these specific embodiments by those skilled in the art without departing from the spirit and scope of the disclosure. Accordingly, it is intended that all such modifications be included within the scope of the disclosure and the general concepts disclosed herein.

Claims

1. A document translation processing method, characterized in that: The method comprises: Obtaining a document to be translated, and segmenting the document to be translated into multiple fragments; Obtaining contextual content of a first fragment among the multiple fragments in the document to be translated; wherein the contextual content includes original content and translated content of a first preset number of text units preceding and adjacent to the first fragment in the document to be translated; Inputting the context content and the original content of the first segment into a translation model, and outputting the translated content of the first segment after the translation model translates the original content of the first segment; Based on the translated content of the first segment, a translation result document corresponding to the document to be translated is determined.

2. The document translation processing method according to claim 1, characterized in that: After inputting the context content and the original content of the first segment into the translation model, and after the translation model translates the original content of the first segment, and before outputting the translated content of the first segment, the method further includes: In the translated fragment library of the document to be translated, a translated fragment whose text similarity with the first fragment satisfies a preset similarity condition is determined as the first translated fragment based on the original content, and the original content and the translated content of the first translated fragment are obtained; wherein the translated fragment library stores the original content and the translated content of the translated fragment of the document to be translated; The step of inputting the context content and the original content of the first segment into a translation model, and outputting the translated content of the first segment after the translation model translates the original content of the first segment, includes: The original content and translated content of the first translated segment, the context content, and the original content of the first segment are input into a translation model. After the translation model translates the original content of the first segment, the translated content of the first segment is output.

3. The document translation processing method according to claim 1, characterized in that: The method further comprises: inputting the context content and the original content of the first segment into the translation model, and after the translation model translates the original content of the first segment, and before outputting the translated content of the first segment. Extracting a word from the original content of the first fragment and determining whether the word exists in a target word list; wherein the target word list stores a plurality of words and translations corresponding to the plurality of words; If the word exists in the target vocabulary, obtaining the translation corresponding to the word; The step of inputting the context content and the original content of the first segment into a translation model, and outputting the translated content of the first segment after the translation model translates the original content of the first segment, includes: The words in the target vocabulary in the original content of the first segment, the translations corresponding to the words, the context content, and the original content of the first segment are input into the translation model. After the translation model translates the original content of the first segment, the translated content of the first segment is output.

4. The document translation processing method according to claim 3, characterized in that: The method further comprises: Extract word pairs from the original content and the translated content of the first segment, and store the word pairs in the target vocabulary; wherein the word pairs include words and translated texts having a corresponding relationship.

5. The document translation processing method according to claim 2, characterized in that: The method further comprises: The original content and the translated content of the first segment are stored in the translated segment library of the document to be translated.

6. The document translation processing method according to claim 1, characterized in that: The context content also includes original content of a second preset number of text units adjacent to the first segment in the document to be translated.

7. The document translation processing method according to claim 1 or 6, characterized in that: The text unit in the context content includes at least one of a segment, a paragraph, and a sentence.

8. The document translation processing method according to claim 1, characterized in that: The step of segmenting the document to be translated into a plurality of fragments includes: According to the document structure identifier in the document to be translated, the document to be translated is segmented into a plurality of segments; wherein the document structure identifier is used to identify the document structure unit in the document to be translated.

9. The document translation processing method according to claim 8, characterized in that: The step of segmenting the document to be translated into a plurality of fragments according to the document structure identifier in the document to be translated comprises: Dividing the document to be translated into a plurality of chapter units according to chapter identifiers in the document to be translated; According to the paragraph identifier in the chapter unit and a preset length threshold, the chapter unit is divided into a plurality of paragraph units; wherein the length of the paragraph unit is less than the preset length threshold; The multiple paragraph units are respectively determined as fragments corresponding to the document to be translated.

10. The document translation processing method according to claim 1, characterized in that: The method further comprises: Determining the professional terms contained in the original content of the first fragment according to a preset term list; wherein the preset term list stores professional terms and translations of the professional terms; Obtaining translations of professional terms contained in the original content of the first fragment from the preset term list; The step of inputting the context content and the original content of the first segment into a translation model, and outputting the translated content of the first segment after the translation model translates the original content of the first segment, includes: The professional terms contained in the original content of the first segment and the translations of the professional terms, the context content and the original content of the first segment are input into the translation model. After the translation model translates the original content of the first segment, the translated content of the first segment is output.

11. A document translation processing device, characterized in that: The device comprises: A segmentation module is used to obtain a document to be translated and segment the document to be translated into multiple segments; A first acquisition module is configured to acquire contextual content of a first fragment among the multiple fragments in the document to be translated; wherein the contextual content includes original content and translated content of a first preset number of text units preceding and adjacent to the first fragment in the document to be translated; a translation processing module, configured to input the context content and the original content of the first segment into a translation model, and output the translated content of the first segment after the translation model translates the original content of the first segment; The first determining module is configured to determine a translation result document corresponding to the document to be translated based on the translated content of the first segment.

12. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product comprises a computer program / instructions, which implement the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Cited By

  • Medical literature large-scale document self-adaptive block translation method and system

    CN121145890A

  • Method and system for large-scale document adaptive chunking translation of medical literature

    CN121145890B