Document translation method and device, equipment and storage medium

By parsing and inserting placeholder elements, the problem of difficulty in translating formatted documents in the existing technology is solved, and accurate translation and format preservation of multimodal documents are achieved.

CN120805943APending Publication Date: 2025-10-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909142.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty in accurately translating formatted articles and multimodal documents, which affects the accuracy of document translation.

Method used

By parsing documents, it generates parsed content of various content types, including text, images, and tables, and inserts translation content based on placeholder elements to generate translated text, ensuring format consistency.

Benefits of technology

Improves the accuracy and format consistency of document translation, ensuring that content in different modalities remains in the correct format after translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805943A_ABST
    Figure CN120805943A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a document translation method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: acquiring a first document to be translated; the method comprises the following steps: analyzing a first document to generate a plurality of analyzed contents corresponding to a plurality of content types, the types of information contained in the analyzed contents being determined based on the corresponding content types, and the plurality of content types comprising a text type and at least one additional type; based on first analysis content corresponding to the text type in the multiple items of analysis content, generating a translated text corresponding to the text type, the translated text comprising placeholder elements; on the basis of the placeholder element, translation content corresponding to the at least one additional type is inserted into the translation text to generate a second document corresponding to the first document, and the translation content is generated on the basis of second analysis content corresponding to the at least one additional type. According to the embodiment of the invention, the accuracy of document translation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device and computer readable storage medium for document translation. BACKGROUND

[0002] Translation technology can be assisted or automated by means of computer software, artificial intelligence, corpus, and other tools and methods. However, with the development of computer technology, people have higher and higher requirements for document translation, not just word or paragraph translation, but also accurate translation of articles with formats, multi-modal documents, and other content. SUMMARY

[0003] In a first aspect of the present disclosure, a method for document translation is provided. The method comprises: obtaining a first document to be translated; generating, by parsing the first document, a plurality of parsed contents corresponding to a plurality of content types, wherein the type of information contained in the parsed content is determined based on the corresponding content type, and the plurality of content types includes a text type and at least one additional type; generating, based on first parsed content corresponding to the text type in the plurality of parsed contents, a translated text corresponding to the text type, the translated text including a placeholder element; inserting, based on the placeholder element, translation content corresponding to the at least one additional type into the translated text to generate a second document corresponding to the first document, the translation content being generated based on second parsed content corresponding to the at least one additional type.

[0004] In a second aspect of the present disclosure, an apparatus for document translation is provided. The apparatus comprises: an obtaining module configured to obtain a first document to be translated; a parsing module configured to generate, by parsing the first document, a plurality of parsed contents corresponding to a plurality of content types, wherein the type of information contained in the parsed content is determined based on the corresponding content type, and the plurality of content types includes a text type and at least one additional type; a translation module configured to generate, based on first parsed content corresponding to the text type in the plurality of parsed contents, a translated text corresponding to the text type, the translated text including a placeholder element; a generation module configured to insert, based on the placeholder element, translation content corresponding to the at least one additional type into the translated text to generate a second document corresponding to the first document, the translation content being generated based on second parsed content corresponding to the at least one additional type.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, the computer program being executable by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer- executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] It should be understood that nothing in the Summary is intended to limit the scope of the embodiments of the disclosure or the claims which follow, either explicitly or by implication. Additional aspects of the disclosure will be apparent to those of ordinary skill in the art in view of the detailed description that follows. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, aspects, and advantages of embodiments of the disclosure will become more apparent from the following detailed description in conjunction with the accompanying drawings, in which like reference numerals denote like elements, and in which:

[0010] FIG. 1 A schematic diagram illustrating an example environment in which embodiments according to the disclosure can be implemented is shown;

[0011] FIG. 2 A first flow diagram illustrating an example process for document translation according to some embodiments of the disclosure is shown;

[0012] FIG. 3A A second flow diagram illustrating an example process for document translation according to some embodiments of the disclosure is shown;

[0013] FIG. 3B A third flow diagram illustrating an example process for document translation according to some embodiments of the disclosure is shown;

[0014] FIG. 3C A fourth flow diagram illustrating an example process for document translation according to some embodiments of the disclosure is shown;

[0015] FIG. 3D A fourth flow diagram illustrating an example process for document translation according to some embodiments of the disclosure is shown;

[0016] FIG. 4 A schematic block diagram of an example apparatus for document translation according to some embodiments of the disclosure is shown; and

[0017] FIG. 5 A block diagram of an electronic device capable of implementing various embodiments of the disclosure is shown. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It is understood that the drawings of the present disclosure and the embodiments are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] It should be noted that the titles of any sections / sub-sections provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / sub-section. Furthermore, embodiments described in any section / sub-section can be combined with any other embodiment described in the same section / sub-section and / or a different section / sub-section in any manner.

[0020] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicitly and implicitly recited definitions can also be found in the following text. The terms "first," "second," etc. can refer to different or the same objects. Other explicit and implicit definitions can also be found in the following text.

[0021] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the present disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all collection, acquisition, processing, processing, forwarding, use, etc. of data are performed on the premise that users are aware of and confirm. Accordingly, when implementing embodiments of the present disclosure, the type of data or information that can be involved, the range of use, the scenario of use, etc. should be informed to users and authorized by users in a proper manner according to relevant laws and regulations. The specific informing and / or authorization manner can vary according to actual situations and application scenarios, and the scope of the present disclosure is not limited in this aspect.

[0022] In the present specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. Users refuse to process personal information other than the necessary information required for basic functions, which will not affect the user's use of basic functions.

[0023] As mentioned above, current document translation generally targets plain text without formatting requirements, but cannot process documents containing other modal content or documents with specific formatting requirements, which affects the accuracy of document translation.

[0024] An embodiment of the present disclosure proposes a document translation solution. The solution includes: obtaining a first document to be translated; parsing the first document to generate multiple parsed contents corresponding to multiple content types, wherein the type of information contained in the parsed contents is determined based on the corresponding content type, and the multiple content types include a text type and at least one additional type; based on a first parsed content corresponding to the text type in the multiple parsed contents, generating a translation text corresponding to the text type, the translation text including a placeholder element; based on the placeholder element, inserting a translation content corresponding to the at least one additional type into the translation text to generate a second document corresponding to the first document, the translation content generated based on the second parsed content corresponding to the at least one additional type.

[0025] In this way, the embodiments of the present disclosure can translate the content of different modalities separately, thereby applying different translation strategies to translate the content of different modalities. On the one hand, it can ensure the accuracy of the content translation, and on the other hand, it can ensure that the content of different modalities still has the correct format after translation.

[0026] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.

[0027] Example Environment

[0028] FIG. 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. FIG. 1 As shown, the example environment 100 may include an electronic device 110. In some embodiments, the electronic device 110 may obtain a first document to be translated; by parsing the first document, generate multiple parsed contents corresponding to multiple content types, wherein the type of information contained in the parsed contents is determined based on the corresponding content type, and the multiple content types include a text type and at least one additional type; based on the first parsed content corresponding to the text type in the multiple parsed contents, generate a translation text corresponding to the text type, and the translation text includes a placeholder element; based on the placeholder element, insert translation content corresponding to at least one additional type into the translation text to generate a second document corresponding to the first document, and the translation content is generated based on the second parsed content corresponding to the at least one additional type. The model 120 for document translation can be deployed on the electronic device 110, and can also be deployed on other devices, which will not be elaborated here.

[0029] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) terminal, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combinations of these, including accessories and peripherals of these devices, or any combinations thereof. In some embodiments, the electronic device 110 can also support any type of interface to the target user (such as “wearable” circuitry, etc.).

[0030] The electronic device 110 can also be a standalone physical server, a server cluster or distributed system of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services of big data and artificial intelligence platforms, etc. The electronic device 110 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.

[0031] It should be understood that the structure and functionality of the various embedded elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0032] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0033] Example Process

[0034] FIG. 2 A flowchart of an example process 200 of document translation is shown in accordance with some embodiments of the present disclosure. The process 200 can be implemented at the electronic device 110. The process 200 is described below with reference to FIG. 1 .

[0035] As shown in FIG. 2 , at block 210, the electronic device 110 can obtain a first document to be translated.

[0036] In embodiments of the present disclosure, the first document is a document to be translated. For example, the first document can be in Chinese and needs to be translated into English; or the first document can be in English and needs to be translated into Chinese, and the like. Moreover, the format of the first document is not limited herein, i.e., the first document can be a WORD document, a PDF document, or the like. In addition, the first document can include a specific structure, e.g., the first document can include multiple chapters, and each chapter has a specific format and writing specification.

[0037] In some embodiments, an initialization step can be performed before the document translation. A dynamically adjustable configuration set can be established, which can include text segmentation parameters, concurrency parameters, and terminology library parameters. The text segmentation parameters can include a preset length when the document is segmented, and a context window size corresponding to a model. The concurrency parameters can include a table translation concurrency degree, an image translation concurrency degree, and the like. Referring to FIG. 3A After the translation task is started, the task status can be updated, and a translation configuration snapshot can be saved to avoid configuration changes during the translation process, which can affect the final translation quality. The terminology library parameters can include multiple terms and correct translations corresponding to the terms, so as to ensure translation accuracy during the translation process.

[0038] At block 220, the electronic device 110 can generate multiple items of parsed content corresponding to multiple content types by parsing the first document.

[0039] Since the first document can include not only text but also image or table content types, different content types can be translated separately, which can improve the final translation effect. Referring to FIG. 3A At blocks 304 and 305, the first document can be downloaded to the memory, and then the first document can be parsed. For example, if the first document includes text type content, image type content, and table type content, the parsed content corresponding to the text type, the parsed content corresponding to the image type, and the parsed content corresponding to the table type can be parsed from the first document. That is, after the first document is parsed, the following JSON format parsed result including multiple items of parsed content can be generated:

[0040]

[0041]

[0042] Through the analysis result, it can be understood that the first document includes the analysis content corresponding to the text type, the analysis content corresponding to the image type, and the analysis content corresponding to the table type. And through the analysis result, it can also be understood that the analysis content corresponds to the style information, for example, the analysis content of the text type can include the style information such as bold, italic, strikethrough, underline, paragraph number, indentation, and the like; the analysis content of the image type can include the style information such as the serial number of the image in the first document, the image format type, the image binary data, and the like; and the analysis content of the table type can include the style information such as the table nesting level, the merged cell coordinates, and the table style.

[0043] In some embodiments, if the first document is in the PDF format, a page including image content can be extracted from the first document, and image binary data can be generated according to the page, and then the image binary data is inserted into the WORD document to be parsed, so as to reuse the image parsing logic in the WORD document parsing for image parsing.

[0044] After the translation of the content in the first document is completed, the translated content can be restored to the style of the content in the first document according to the analysis content, so as to ensure that the style of the content in the second document after translation is the same as that of the corresponding content in the first document.

[0045] In block 230, the electronic device 110 can generate a translated text corresponding to the text type based on the first analysis content corresponding to the text type in the plurality of analysis contents.

[0046] In some embodiments, the translated text includes a placeholder element. For example, a placeholder can be applied to indicate the position of a table in the first document, for example, if a table is included between paragraph 1 and paragraph 2 in the article, the position of the table can be indicated in the form of a placeholder, so that after the translation is completed, the table can be inserted into the corresponding position according to the placeholder, so as to ensure that the positions of the contents in the second document after translation are the same as those in the first document, thereby improving the accuracy of document translation.

[0047] In some embodiments, the placeholder element can also be applied to mark the part of the document that does not need to be translated, for example, a placeholder can be applied to mark some words or paragraphs in the first document, which indicates that these marked words or paragraphs do not need to be translated in the model translation process, so that in the process of applying the model for translation, the part marked by the placeholder can be skipped, and then restored after the translation is completed to obtain the final translation content.

[0048] In some embodiments, since the first document may include specific structural information, such as including multiple chapters or having a specific writing format, the electronic device 110 can segment the first parsed content into multiple text segments based on the structural information of the first document. Then, multiple translation segments corresponding to the multiple text segments are obtained, and then based on the multiple translation segments, the translation text corresponding to the text type is determined. For example, if the first document includes three chapters, the first document can be segmented into a text segment corresponding to the first chapter, a text segment corresponding to the second chapter, and a text segment corresponding to the third chapter, thereby obtaining a translation text.

[0049] In some embodiments, to improve the accuracy of document translation, a model can be applied to the translation. Therefore, during model training, the training samples can be divided into multiple chapters for model training. This allows the model to understand which content belongs to the first chapter and which content belongs to the second chapter during training, and learn which translation strategy should be used.

[0050] This approach allows us to translate the first chapter using a method suitable for translating the first, and the second chapter using a method suitable for translating the second, thereby improving translation accuracy. Furthermore, by segmenting the first document into multiple text segments based on structural information, we ensure that the segmented text segments do not contain content from multiple chapters simultaneously, allowing the model to translate the corresponding text segments using more appropriate strategies.

[0051] Because some chapters in the first document may contain a large amount of content, the translation effect may not be good enough in one go. Therefore, the electronic device 110 may segment the first parsed content into a plurality of initial segments corresponding to the plurality of structural units based on the structural information. The electronic device 110 may then process the plurality of initial segments based on a preset length and a preset segmentation identifier to determine a plurality of text segments, each of which does not exceed a preset length.

[0052] For example, reference FIG. 3A In blocks 308 to 311, the first document is segmented into an initial segment corresponding to the first chapter, an initial segment corresponding to the second chapter, and an initial segment corresponding to the third chapter according to the structural information. The initial segments are then processed based on the preset lengths and the preset segmentation identifiers to determine multiple text segments. For example, the initial segment corresponding to the second chapter may be divided into multiple text segments each having a length less than a preset length. After the first document is segmented into multiple text segments, post-processing may be performed, such as sorting or removing blank lines.

[0053] In some embodiments, the preset segmentation identifier includes a punctuation mark. For example, since the initial segment corresponding to the second chapter can be long, a preset length can be set to 300, i.e., the number of words in each text segment is not more than 300, and a preset segmentation identifier can be determined as a punctuation mark such as a period, an ellipsis, a question mark, or the like, which is used to indicate an end, to ensure that each text segment includes a complete sentence, so as to avoid the problem of splitting a complete sentence into multiple parts, thereby causing semantic inconsistency.

[0054] The first document can include some words or sentences that need to be translated by a specific translation strategy, such as professional terms or words related to the chapter or format of the document. In some embodiments, the electronic device 110 can determine, in response to the first text segment in the plurality of text segments including a target keyword matching a set of keywords, a preset translation word corresponding to the target keyword based on preset translation information corresponding to the set of keywords. Then, based on the preset translation word, a first translation segment corresponding to the first text segment is determined.

[0055] For example, the first document includes some words or sentences that need to be translated by a specific translation strategy, such as the words “neural network”, “specification”, etc. The corresponding translation result can be directly obtained by mapping. Therefore, these target keywords can be annotated with placeholders, and then the model can skip these target keywords during translation. After the translation is completed, the preset keyword corresponding to the target keyword can be directly obtained by mapping, and the first translation segment is determined.

[0056] In some embodiments, in order to improve the accuracy of model translation, the content adjacent to the content to be translated can be input into the model as context, so that the model can obtain more accurate translation content. The electronic device 110 can obtain first context information associated with the second text segment in the plurality of text segments, the first context information being determined based on associated content of the second text segment in the first document. Then, the second text segment and the first context information are provided to the first model to generate a second translation segment corresponding to the second text segment.

[0057] For example, according to the order of the text segments in the first document, three adjacent text segments before the second text segment and three adjacent text segments after the second text segment can be determined, and the first context information can be constructed according to these adjacent text segments. Then, the first context information and the second text segment are input into the model, so that the model outputs a more accurate second translation segment according to the context.

[0058] In some embodiments, based on the first parsed content corresponding to the text type in the plurality of parsed contents, the translation process of generating the translated text corresponding to the text type can fail, and the first context information is determined based on the first parsed content corresponding to the text type in the plurality of parsed contents.FIG. 3B The middle frame 322 to the frame 327, if the translation fails, re-translation, if re-translation more than 10 times, still translation failure, translation task failure, and end the translation task. In order to improve the quality of translation, if the translation is successful, the generated translation content needs to be checked, if the check fails, re-translation, if re-translation more than 5 times, still translation failure, translation task failure, and end the translation task.

[0059] In some embodiments, the condition of the check failure can be set in advance, for example, if the length difference between the elements in the translation content and the corresponding elements in the first document is more than 3 times, the check failure is considered. For example, the number of placeholders does not match the content of the first document. For example, in the Chinese-English translation task, the translation content includes Chinese characters. For example, the number of lines of the translation content and the first document is inconsistent, etc. The check failure condition and the number of re-translation are not limited here, and can be adjusted according to the actual situation.

[0060] In block 240, the electronic device 110 can insert the translation content corresponding to the at least one additional type into the translated text based on the placeholder element to generate a second document corresponding to the first document, the translation content being generated based on the second parsed content corresponding to the at least one additional type.

[0061] In some embodiments, if the first document includes text type, image type and table type, after translating the content of the text type, the content of the image type and the content of the table type, the translation content corresponding to the image type and the translation content corresponding to the table type can be inserted into the translation content of the text type to generate the second document.

[0062] In addition, since the parsed content includes style information, such as bold content in the text, after translation, the translation content can be restored according to the style information, that is, the translated content corresponding to the bold text is also processed to be bold, to ensure that the second document and the first document are consistent in style.

[0063] In some embodiments, since the first document can include images, the electronic device 110 can detect image content corresponding to the image type from the first document, and then generate second parsed content corresponding to the image content, the second parsed content including description information associated with the image content. Determine the first translation content corresponding to the image content to be inserted into the translated text.

[0064] For example, an image in a first document may be determined to be image content of the image type, and then the image content may be parsed to obtain second parsed content. The second parsed content includes descriptive information associated with the image content, such as the image location, image format, and the image sequence number in the first document. The image content is then translated to obtain first translated content, which is then inserted into the translated text.

[0065] In some embodiments, contextual information can be applied to improve the accuracy of image translation. In some embodiments, in response to determining that the image content includes content to be translated, electronic device 110 can obtain second contextual information associated with the image content. The image content and the second contextual information are then provided to a second model to generate first translation content corresponding to the image content.

[0066] For example, the first document may contain descriptive text corresponding to the image content. Therefore, the descriptive text associated with the content to be translated contained in the image content, i.e., the second context information, can be obtained. Both the image content and the second context information are then provided to the model, allowing the model to better translate the content to be translated based on the descriptive text.

[0067] In some embodiments, the electronic device 110 may determine a first reference text segment associated with the image content from the first document, and then generate second context information based on the first reference text segment. FIG. 3A In block 306, optical character recognition (OCR) can be applied to extract the text in the flowchart to obtain the content to be translated. A first reference text segment associated with the content to be translated is then searched from the first document, such as a segment with identical or similar content. Second context information is then generated based on the first reference text segment. This allows the model to translate the image based on these two pieces of content, achieving better translation quality.

[0068] Specifically, if an image includes the four characters "neural network" written vertically, then optical character recognition alone will likely recognize the four characters "神," "经," "网," and "络" separately, but will fail to recognize the word "neural network." By comparing the image with the first document and inputting a similar first reference text fragment into the model as context, the model can recognize that the four characters belong to a single word, thereby ensuring semantic coherence in the translation.

[0069] In some embodiments, in response to the image content corresponding to the formula content, it is determined that the image content does not include content to be translated. FIG. 3DThe middle box 336 to 339, because the image in the first document can include two parts: the part that needs to be translated and the part that does not need to be translated. The part that does not need to be translated can include formula image, image without text, image without Chinese in the Chinese-English translation task, image without English in the English-Chinese translation task, etc. The part of the formula is in the form of a screenshot in the first document, so the image of the formula content does not need to be translated and can be directly inserted into the translated content, so the image of the formula content does not include the content to be translated.

[0070] The formula in the document can include three types, one is the formula inserted by Latex, one is the formula edited by WORD editor, and the last one is the formula displayed in the form of image. The first two types of formulas can be recognized in the analysis stage, but the formula displayed in the form of image cannot be recognized. Therefore, the middle box 306 refers to FIG. 3A The middle box 307 can input the image content and the text extracted from the image into the model, apply the model to identify whether the image is a formula, if the model identifies that the image content is a formula, return "IS_FORMULA" and the coordinates of the image; if the model identifies that the image content is not a formula, return the content to be translated and the coordinates of the image.

[0071] Then, for the image content identified as a formula, the original image can be directly retained in the final translated content without text translation. For the image content identified as non-formula, the text in the image content needs to be translated, and then the translated text is appended to the document position after the image. In addition, all the texts in the image content are spliced with "---" as the delimiter, and if the output translated content does not match the number of lines of the first document, it is not considered as a translation failure.

[0072] In some embodiments, because the first document can also include a table, the electronic device 110 can detect the table content corresponding to the table type from the first document, and then generate the third analysis content corresponding to the table content, which at least indicates the merging relationship of the cells in the table content. The second translation content corresponding to the table content is determined to be inserted into the translated text.

[0073] For example, if the first document includes table content, the third analysis content is obtained after analyzing the table content, and through the third analysis content, the structure of the table, whether there is a merged cell in the table, whether there is a nested table in the table, etc. Information is obtained in order to apply the third analysis content to reconstruct the table structure in the subsequent steps. Then the table content is translated, and the obtained second translation content is inserted into the translated text.

[0074] Specifically, if the table content is a 4x4 table, and the four table elements in the first row are merged into one table element, then during translation, only one table element needs to be translated, instead of four. Through the third parsed content, it can be understood that there are merged cells in the table content, and the location of the merged cells and which table elements do not need to be translated can be determined, so as to facilitate the restoration of the original merged cells after table translation, so as to ensure that the structure of the table before translation is the same as that after translation. And each cell in the table content is spliced by “---” as a separator, and the translated content is also split by “---”.

[0075] For example, there can also be a table nesting situation in the table content, that is, a complete table is inserted into a cell of the table to form a hierarchical structure, for displaying more complex data relationships or classification information in the cell. For this situation, all nested tables in the table content can be extracted first, and then these tables are distributed to independent threads according to the concurrency degree, so that each thread processes one table or sub-table.

[0076] In some embodiments, the electronic device 110 can determine a second reference text segment associated with the table content from the first document, and then generate third context information based on the at least one second reference text segment. The third parsed content and the third context information are provided to the third model to generate the second translation content. For example, the second reference text segment that coincides or is similar to the text in the table content can be queried from the first document according to the text, and the third context information is generated. Then the third context information and the third parsed content are input into the model to generate the second translation content.

[0077] In some embodiments, since the table can also include images, the images can be processed in the process of processing the image content described above, so that the images do not need to be repeatedly translated when translating the table. In addition, after translating the table, the translated content can be checked, and it is checked whether the number of rows of the translated content is the same as that of the table content, if not, re-translation is performed, and if the number of re-translation times exceeds 3, translation is performed according to the cell, and the final translation content is obtained. FIG. 3C

[0078] ​By this method, translation for documents including multiple content types can be realized, and the format of the output translated content is consistent with the format of the input content, thereby ensuring the accuracy of the translation of elements such as text, paragraph numbers, tables, images, etc. By identifying the formulas in the images, translation errors caused by incorrect identification of formulas can be avoided. By identifying the nested structure of the table and merging the cells, etc., it can be ensured that the translated table still retains the original structure. Through the verification step after translation, the quality of translation can be further improved.

[0079] Example Devices and Apparatus

[0080] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 4 A schematic block diagram of an example apparatus 400 for document translation according to certain embodiments of the present disclosure is shown. The apparatus 400 can be implemented as or included in the electronic device 110. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0081] As shown in FIG. 4 The apparatus 400 includes an obtaining module 401 configured to obtain a first document to be translated; a parsing module 402 configured to generate, by parsing the first document, multiple parsed contents corresponding to multiple content types, wherein the type of information contained in the parsed content is determined based on the corresponding content type, and the multiple content types include a text type and at least one additional type; a translation module 403 configured to generate, based on first parsed content corresponding to the text type in the multiple parsed contents, a translated text corresponding to the text type, the translated text including a placeholder element; and a generating module 404 configured to insert, based on the placeholder element, translation content corresponding to the at least one additional type into the translated text to generate a second document corresponding to the first document, the translation content being generated based on second parsed content corresponding to the at least one additional type.

[0082] In some embodiments, the translation module 403 is configured to split the first parsed content into multiple text segments based on structure information of the first document; obtain multiple translated segments corresponding to the multiple text segments; and determine the translated text corresponding to the text type based on the multiple translated segments.

[0083] In some embodiments, the translation module 403 is configured to split the first parsed content into multiple initial segments corresponding to multiple structure units based on the structure information; and process the multiple initial segments based on a preset length and a preset split identifier to determine the multiple text segments, each text segment having a length not exceeding the preset length.

[0084] In some embodiments, the preset split identifier includes a punctuation mark.

[0085] In some embodiments, the translation module 403 is configured to, in response to the first text segment in the plurality of text segments including a target keyword matching the set of keywords, determine a preset translation word corresponding to the target keyword based on preset translation information corresponding to the set of keywords; and determine a first translation segment corresponding to the first text segment based on the preset translation word.

[0086] In some embodiments, obtaining the plurality of translation segments corresponding to the plurality of text segments includes: obtaining first context information associated with a second text segment in the plurality of text segments, the first context information being determined based on associated content of the second text segment in the first document; and providing the second text segment and the first context information to the first model to generate a second translation segment corresponding to the second text segment.

[0087] In some embodiments, the apparatus 400 is configured to detect image content corresponding to an image type from the first document; generate second parsed content corresponding to the image content, the second parsed content including description information associated with the image content; and determine the first translation content corresponding to the image content for insertion into the translated text.

[0088] In some embodiments, the apparatus 400 is configured to, in response to determining that the image content includes the content to be translated, obtain second context information associated with the image content; and provide the image content and the second context information to the second model to generate the first translation content corresponding to the image content.

[0089] In some embodiments, the apparatus 400 is configured to determine a first reference text segment associated with the image content from the first document; and generate the second context information based on the first reference text segment.

[0090] In some embodiments, the apparatus 400 is configured to, in response to the image content corresponding to formula content, determine that the image content does not include the content to be translated.

[0091] In some embodiments, the apparatus 400 is configured to detect table content corresponding to a table type from the first document; generate third parsed content corresponding to the table content, the third parsed content indicating at least a merging relationship of cells in the table content; and determine second translation content corresponding to the table content for insertion into the translated text.

[0092] In some embodiments, the apparatus 400 is configured to determine a second reference text segment associated with the table content from the first document; generate third context information based on at least one second reference text segment; and provide the third parsed content and the third context information to the third model to generate the second translation content.

[0093] AsFIG. 5 As shown, the electronic device 500 is in the form of a general electronic device. Components of the electronic device 500 can include, but are not limited to, at least one processor 510 or processing unit, a memory 520, a storage device 530, one or more communication units 560, one or more input devices 550, and one or more output devices 560. The processor 510 can be a real or virtual processor and is capable of performing various processing according to programs stored in the memory 520. In a multi-processor system, multiple processors perform computer-executable instructions in parallel to improve parallel processing capability of the electronic device 500.

[0094] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable media and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and that can be accessed by the electronic device 500.

[0095] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5 medium (e.g., a "floppy disk") or a removable, non-volatile memory disk (e.g., a "flash drive"), which can be used for storing information and / or data. In these cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0096] The communication unit 540 enables communication with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0097] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown), such as storage devices, display devices, etc., one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices, as desired via the communication unit 540. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0098] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0099] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0100] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0101] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0102] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The computer program product can also be a downloaded file, such as a file downloaded from the Internet or another computer network. Of course, many modifications can be made by those skilled in the art to the inventive concept described herein, which can be practiced in various embodiments and over a variety of applications, and that the scope of the application is limited only by the following claims. For example, a computer program product can be a downloaded file, such as a file downloaded from the Internet or another computer network. Of course, many modifications can be made by those skilled in the art to the inventive concept described herein, which can be practiced in various embodiments and over a variety of applications, and that the scope of the application is limited only by the following claims.

[0103] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations described above have particular applications in the medical field, it should be recognized that the implementations of the disclosure have broad applicability to other fields. Numerous modifications and variations will be apparent to those skilled in the art in light of the above teachings. For example, the above-described implementations can be used in a variety of medical applications, such as in the diagnosis of medical conditions, in the administration of medical treatment, and in the administration of medical care. Any or all of the individual features or operations described above can be used separately or in any combination. By way of non-limiting example, features described above can be used in a medical device, a medical system, a medical method, a medical apparatus, a medical instrument, a medical computer program, a medical computer program product, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program distribution medium, a medical computer program

Claims

1. A document translation method, comprising: Obtaining a first document to be translated; generating, by parsing the first document, a plurality of parsed contents corresponding to a plurality of content types, wherein the type of information contained in the parsed contents is determined based on the corresponding content types, the plurality of content types including a text type and at least one additional type; generating, based on a first parsed content corresponding to the text type among the multiple parsed contents, a translated text corresponding to the text type, wherein the translated text includes a placeholder element; as well as Based on the placeholder element, translation content corresponding to the at least one additional type is inserted into the translated text to generate a second document corresponding to the first document, where the translation content is generated based on second parsed content corresponding to the at least one additional type.

2. The method according to claim 1, wherein generating a translation text corresponding to the text type based on a first parsed content corresponding to the text type among the multiple parsed contents comprises: Based on the structural information of the first document, segmenting the first parsed content into a plurality of text segments; Obtaining multiple translation segments corresponding to the multiple text segments; as well as Based on the multiple translation segments, a translation text corresponding to the text type is determined.

3. The method according to claim 2, wherein the step of segmenting the first parsed content into a plurality of text segments based on the structural information of the first document comprises: Based on the structural information, segmenting the first parsed content into a plurality of initial segments corresponding to a plurality of structural units; as well as Based on a preset length and a preset segmentation identifier, the multiple initial segments are processed to determine the multiple text segments, where the length of each text segment does not exceed the preset length. The method according to claim 3 , wherein the preset segmentation mark includes a punctuation mark.

5. The method according to claim 2, wherein obtaining a plurality of translation segments corresponding to the plurality of text segments comprises: In response to a first text segment among the plurality of text segments including a target keyword matching a keyword set, determining a preset translation word corresponding to the target keyword based on preset translation information corresponding to the keyword set; as well as Based on the preset translation word, a first translation segment corresponding to the first text segment is determined.

6. The method according to claim 2, wherein obtaining a plurality of translation segments corresponding to the plurality of text segments comprises: Obtaining first context information associated with a second text segment among the plurality of text segments, the first context information being determined based on associated content of the second text segment in the first document; as well as The second text segment and the first context information are provided to a first model to generate a second translation segment corresponding to the second text segment.

7. The method of claim 1 , wherein the at least one additional type comprises an image type, the method further comprising: detecting image content corresponding to the image type from the first document; generating second parsed content corresponding to the image content, wherein the second parsed content includes description information associated with the image content; as well as A first translation content corresponding to the image content is determined to be inserted into the translated text.

8. The method of claim 7, wherein determining the first translation content corresponding to the image content comprises: In response to determining that the image content includes content to be translated, obtaining second context information associated with the image content; as well as The image content and the second context information are provided to a second model to generate the first translation content corresponding to the image content.

9. The method according to claim 8, wherein obtaining second context information associated with the image content comprises: determining, from the first document, a first reference text segment associated with the image content; as well as The second context information is generated based on the first reference text segment.

10. The method according to claim 8, further comprising: In response to the image content corresponding to formula content, it is determined that the image content does not include content to be translated.

11. The method of claim 1 , wherein the at least one additional type comprises a table type, the method further comprising: detecting table content corresponding to the table type from the first document; generating a third parsed content corresponding to the table content, wherein the third parsed content at least indicates a merge relationship of cells in the table content; as well as A second translation content corresponding to the table content is determined to be inserted into the translation text.

12. The method according to claim 11, wherein determining the second translation content corresponding to the table content comprises: determining, from the first document, a second reference text segment associated with the table content; generating third context information based on the at least one second reference text segment; as well as The third parsed content and the third context information are provided to a third model to generate the second translated content.

13. A device for document translation, comprising: An acquisition module, configured to acquire a first document to be translated; a parsing module configured to generate a plurality of parsed contents corresponding to a plurality of content types by parsing the first document, wherein the type of information contained in the parsed contents is determined based on the corresponding content type, the plurality of content types including a text type and at least one additional type; a translation module configured to generate a translation text corresponding to the text type based on a first parsed content corresponding to the text type among the multiple parsed contents, the translation text including a placeholder element; A generation module is configured to insert translation content corresponding to the at least one additional type into the translated text based on the placeholder element to generate a second document corresponding to the first document, wherein the translation content is generated based on second parsed content corresponding to the at least one additional type.

14. An electronic device comprising: at least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processor. 15 . A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to claim 1 .

16. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 12.