Method and system for translating PDF (Portable Document Format) text containing complex features

By evaluating the complexity of PDF files and calling the corresponding text detection model, and combining the multi-translation model for translation and layout, the accuracy and format completeness of PDF files' complex content recognition and translation in the existing technology are solved, and efficient and accurate PDF text translation is achieved.

CN120163169APending Publication Date: 2025-06-17AFIRSTSOFT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510102590.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately identify complex content when processing PDF files, resulting in poor accuracy and fluency of translation results, and it is difficult to maintain the format and content integrity of the original document.

Method used

By calculating the complex feature comprehensive score of PDF files, calling the corresponding text detection model for text area recognition, extracting layout information, and translating through collaborative work of multiple translation models, and finally typesetting and output based on layout information.

Benefits of technology

It realizes accurate identification and translation of complex content in PDF files, ensuring the accuracy and fluency of translation results, while retaining the format and content integrity of the original document.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163169A_ABST
    Figure CN120163169A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of text translation, and provides a method and a system for translating a PDF (Portable Document Format) text containing complex features. The method comprises the steps that a PDF analysis engine is initialized, a PDF file is read, and basic information of the PDF file is extracted; judging whether the PDF file is a complex file or not, and recording complex features; preprocessing an image in the complex document, and calling a corresponding text detection model according to the complex features to perform text region identification; and extracting layout information of the text in the text area, translating the text in the text area through a translation model, performing typesetting according to the layout information after a final translation result is obtained, and outputting a target translation file in a user-defined manner. According to the method, the complex content in the PDF document can be intelligently identified, and the original document format is completely reserved in the translation process; meanwhile, multi-language translation is supported, and the accuracy of PDF document translation is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text translation, and particularly relates to a PDF text translation method and system including complex features. Background Art

[0002] PDF documents are one of the main formats for electronic document transmission and storage. When disseminating or sharing information, it is often necessary to translate PDF files. For the translation of PDF files in the prior art, it is usually necessary to convert the PDF into a Word document or other editable format and then use a translation tool for translation, and save the document as a PDF format again after the translation is completed. There are the following defects: First, many current translation tools often have problems when processing PDF files, such as being unable to correctly recognize some texts or even showing garbled characters, resulting in poor accuracy and fluency of the translation results. Second, for PDF files containing a mixture of graphics and text or a large number of non-text elements, the existing PDF translation methods often have problems such as content loss or format disorder when converting the document format, resulting in incomplete content of the translated document and difficulty in restoring to the structure and appearance of the original document.

[0003] Therefore, we need to develop a PDF text translation method and system including complex features that can intelligently recognize the complex content in PDF documents and completely retain the original document format during the translation process; at the same time, it supports multi-language translation to ensure the accuracy of PDF document translation. Summary of the Invention

[0004] The purpose of the present invention is to provide a PDF text translation method and system including complex features to solve the problems in the prior art PDF translation methods, such as difficulty in recognizing the complex content in PDF documents, difficulty in maintaining the original document format, and poor translation quality mentioned in the above background art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: According to one aspect of the present invention, there is provided a PDF text translation method including complex features, and the method includes the following steps: Initialize a PDF parsing engine, read a PDF file and extract the basic information of the PDF file; Judge whether the PDF file is a complex document and record the complex features; Preprocess the images in the complex document and call the corresponding text detection model according to the complex features to identify the text area; Extract the layout information of the text in the text area, and translate the text in the text area through a translation model to obtain the final translation result; Typeset the final translation result according to the layout information, and customize and output the target translation file.

[0006] According to another aspect of the present invention, there is provided a PDF text translation system including complex features, and the system includes: an initialization module, a document judgment module, a text recognition module, a text translation module, and a translation result output module. Among them: The above-mentioned initialization module is used to initialize the PDF parsing engine, read the PDF file, and extract the basic information of the PDF file; The above-mentioned document judgment module is used to judge whether the PDF file is a complex document and record the complex features; The above-mentioned text recognition module is used to preprocess the images in the complex document and call the corresponding text detection model according to the complex features to identify the text area; The above-mentioned text translation module is used to extract the layout information of the text in the text area, translate the text in the text area through a translation model, and obtain the final translation result; The above-mentioned translation result output module is used to typeset the final translation result according to the layout information, and customize and output the target translation file.

[0007] Based on the foregoing solution, the extraction of the basic information of the PDF file specifically includes: Parse the PDF file through the PDF parsing engine to obtain the page tree structure of the PDF file; Traverse each object in each page of the PDF one by one according to the page tree structure, and read the metadata, page information, font information, and image information of the PDF file to obtain the basic information of the PDF file.

[0008] Based on the foregoing solution, the judgment of whether the PDF file is a complex document specifically includes: According to the basic information of the PDF file, count the number of pictures or tables in the PDF file, and detect whether the PDF file is a picture-type PDF; Based on the text density analysis method, evaluate the text density of each page in the PDF file; Detect whether there are non-standard texts, mixed text and graphics, or overlapping paragraphs in the PDF file; Assign weights to the complex features, and calculate the comprehensive score of the complex features in the PDF file; Compare the comprehensive score with a preset threshold. If the comprehensive score exceeds the preset threshold, it is determined that the PDF file is the complex document.

[0009] When determining whether the PDF file is a complex document, it includes calculating and judging on a page-by-page basis, or calculating and judging on the basis of some rectangular areas in the page; specifically, it can be flexibly selected according to the situations of different pages in the PDF file.

[0010] Based on the foregoing solution, the complex features include: text density, picture-type PDF, number of tables, non-standard text, text and picture mixing, and paragraph overlap; among them, the non-standard text includes: text in images, text drawn by paths, text with font embedding, and text containing mixed languages or special symbols.

[0011] Based on the foregoing solution, the preprocessing of the images in the complex document includes: noise removal, image enhancement, and edge detection.

[0012] Based on the foregoing solution, the text detection model includes a lightweight CRAFT model and an improved EAST model; among them, the lightweight CRAFT model uses MobileNetV3 as the backbone network; the improved EAST model adopts a dynamic backbone network selection strategy and modifies the output head of the EAST model to support the detection of rotated rectangles.

[0013] Based on the foregoing solution, the calling of the corresponding text detection model according to the complex features for text area recognition specifically includes: If the complex feature is text density or non-standard text or paragraph overlap, then call the lightweight CRAFT model for text area recognition; If the complex feature is picture-type PDF or number of tables or text and picture mixing, then call the improved EAST model for text area recognition.

[0014] Based on the foregoing solution, the extraction of the layout information of the text in the text area is implemented through a pre-trained LayoutLMv3 model; when translating the text in the text area through a translation model, the translation model includes an optimized mBART model and an optimized T5 model, and the translation process specifically includes: Detect the delimiters in the text area, and segment the text in the text area based on the delimiters; Input the segmented text into the optimized mBART model paragraph by paragraph for translation to obtain a first translation result; Input the first translation result into the optimized T5 model for optimization to obtain a second translation result; Perform post-processing on the second translation result to obtain the final translation result.

[0015] As can be seen from the above technical solutions, compared with the prior art, the present invention has at least the following advantages and positive effects: (1) By calculating the comprehensive score of complex features in the PDF file to evaluate the complexity of the PDF file, and calling the corresponding text detection model according to different complex features for text region recognition, the present invention can accurately identify the text content in the PDF file, ensuring the accuracy of subsequent translation results.

[0016] (2) The present invention uses a lightweight CRAFT model and an improved EAST model to detect text in PDF files containing complex features. It can not only accurately identify the content in complex documents, avoiding the problem of text omission in traditional OCR recognition methods, but also improve text detection efficiency, with high efficiency and applicability.

[0017] (3) By using the LayoutLMv3 model to extract the layout information of the text, the present invention can retain the layout structure of the original document, ensuring the consistency of the translated text and the original text in layout.

[0018] (4) By adopting a multi-translation model collaborative working method, the present invention can not only support multi-modal translation of text, meet diverse translation needs, but also optimize the translation results, improve the accuracy of the translated text, and ensure the translation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description only relate to some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 shows a flowchart of a method for translating PDF text containing complex features provided by an embodiment of the present invention; Figure 2 shows a flowchart of a method for judging the complexity of a PDF file provided by an embodiment of the present invention; Figure 3 shows a flowchart of a method for text region recognition based on a text detection model provided by an embodiment of the present invention; Figure 4 shows a flowchart of a method for text translation provided by an embodiment of the present invention; Figure 5 shows a schematic structural diagram of a system for translating PDF text containing complex features provided by an embodiment of the present invention; Wherein, Figure 5 the reference numerals in the drawings are explained as follows: 500 - A PDF text translation system with complex features; 501 - Initialization module, 5011 - PDF file reading unit, 5012 - Basic information extraction unit; 502 - Document judgment module, 5021 - Complexity assessment unit, 5022 - Complex feature recording unit; 503 - Text recognition module, 5031 - Image preprocessing unit, 5032 - Text area recognition unit; 504 - Text translation module, 5041 - Layout information extraction unit, 5042 - Text translation unit; 505 - Translated text output module, 5051 - Typesetting adjustment unit, 5052 - File output unit. Detailed implementation mode

[0021] In order to more clearly illustrate the purpose, technical solution and advantages of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The example embodiments can be implemented in multiple forms and should not be construed as limited to the examples described herein; on the contrary, providing these embodiments makes the present invention more comprehensive and complete, and conveys the concept of the example embodiments to those skilled in the art in an all-round way.

[0022] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present invention. However, those skilled in the art will realize that the technical solutions of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present invention.

[0023] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0024] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0025] The present invention will be described in detail below in conjunction with specific embodiments.

[0026] Embodiment 1

[0027] As Figure 1 shown, an embodiment of the present invention provides a PDF text translation method including complex features, and the specific steps of this method are as follows: S1: Initialize the PDF parsing engine, read the PDF file and extract the basic information of the PDF file; When the system starts, the client first initializes the PDF parsing engine, reads the PDF file specified by the user, and extracts the basic information of the PDF file. Specifically, in this embodiment, the PDF parsing engine is obtained by secondary development of the open-source library PDFium, and the extraction of the basic information of the PDF file specifically includes: Parse the PDF file through the PDF parsing engine to obtain the page tree structure of the PDF file; Traverse each object in each page of the PDF according to the page tree structure, read the metadata, page information, font information, and image information of the PDF file to obtain the basic information of the PDF file.

[0028] S2: Determine whether the PDF file is a complex document and record complex features; Further, after extracting the basic information of the PDF file, start to evaluate the complexity of the PDF file. As Figure 2 shown, determining whether the PDF file is a complex document specifically includes: S201: According to the basic information of the PDF file, count the number of pictures or tables in the PDF file and detect whether the PDF file is a picture-type PDF; Specifically, set a threshold for the number of pictures or tables in the PDF file according to actual needs. According to the basic information of the PDF file, count the number of elements of type picture or table in each page of the PDF, determine whether the number of pictures or tables in each page exceeds the set threshold, and record it. At the same time, determine whether the PDF file is a picture-type PDF according to the number of pictures in each page of the PDF: if the whole page is a picture (such as a scanned copy) or the number of pictures exceeds the set threshold, then determine that the PDF file is a picture-type PDF and record it.

[0029] S202: Based on the text density analysis method, evaluate the text density of each page in the PDF file; Specifically, the evaluation of the text density on a page includes, but is not limited to: evaluating by the number of characters per unit area (for example, when the number of characters per square centimeter exceeds a certain threshold, it is determined that the text density of the page is relatively high), evaluating by line spacing and character spacing (for example, when the character spacing is less than 10% of the character height or the line spacing is less than 15% of the line height, it is determined that the text density of the page is relatively high), or evaluating by the proportion of the text layout area (for example, when the area occupied by the text on the page exceeds 90% of the page, it is determined that the text density of the page is relatively high); the specific evaluation method can be flexibly selected and adjusted according to actual needs, and this embodiment does not make any restrictions.

[0030] S203: Detect whether there are non-standard texts, text and image mixing, or paragraph overlapping in the PDF file; Preferably, in this embodiment, non-standard texts include: texts in images, texts drawn by paths, texts with font embedding, and texts containing mixed languages or special symbols. Detect whether there are non-standard texts, text and image mixing, or paragraph overlapping in each page of the PDF file through the PDF parsing engine, and record them.

[0031] S204: Assign weights to the complex features, and calculate the comprehensive score of the complex features in the PDF file; Specifically, in this embodiment, the complex features include: picture-type PDF, number of tables, text density, non-standard texts, text and image mixing, and paragraph overlapping; in actual application scenarios, users can assign different weights to different complex features according to the importance of the complex features or the impact of the complex features on the translation of the PDF file, and obtain the comprehensive score of the complex features in the PDF file by calculating the comprehensive weights of each complex feature.

[0032] S205: Compare the comprehensive score with a pre-set threshold. If the comprehensive score exceeds the pre-set threshold, determine that the PDF file is the complex document.

[0033] Furthermore, after obtaining the comprehensive score of the complex features in the PDF file, compare the comprehensive score with a pre-set threshold. If the comprehensive score exceeds the pre-set threshold, determine that the PDF file is the complex document. Preferably, when determining whether the PDF file is a complex document, it includes calculating and judging on a page-by-page basis, or calculating and judging on the basis of some rectangular areas in the page; specifically, it can be flexibly selected according to the situations of different pages in the PDF file. In addition, in specific scenarios, it is also possible to set the determination conditions for simultaneously meeting multiple complex features to improve the reliability of complex document determination; or increase or decrease the detection of some complex features according to the needs of the scenario, and specific adjustments can be made flexibly according to the actual situation, and this embodiment does not make any restrictions.

[0034] S3: Preprocess the images in the complex document, and call the corresponding text detection model according to the complex features for text region recognition; Preferably, in this embodiment, preprocessing the images in the complex document includes: noise removal, image enhancement, and edge detection. Among them, when removing noise, a filtering algorithm (such as Gaussian filtering) is used to remove the noise in the image; image enhancement mainly aims at low-resolution blurred images, including but not limited to using histogram equalization, contrast adjustment, or super-resolution technology to improve the image clarity and make the text in the image more obvious; edge detection mainly uses the Canny edge detection algorithm to accurately extract the text region contour in the image for subsequent text region recognition.

[0035] Furthermore, after preprocessing the images in the complex document, call the corresponding text detection model according to the complex features for text region recognition. Preferably, in this embodiment, the text detection model includes a lightweight CRAFT model and an improved EAST model; among them, the lightweight CRAFT model uses MobileNetV3 as the backbone network; the improved EAST model adopts a dynamic backbone network selection strategy and modifies the output head of the EAST model to support the detection of rotated rectangles.

[0036] In this embodiment, the lightweight CRAFT model uses MobileNetV3 as the backbone network, which can effectively improve the model recognition speed, especially showing better performance on devices with low CPU performance or mobile devices. The improved EAST model adopts a dynamic backbone network selection strategy. For example, for a complex document with a mixture of text and images, a lightweight ResNet-50 can be used as the backbone network to improve the diversity of feature extraction; for a document with relatively simple text and images, MobileNet or ShuffleNet can be used as the backbone network to reduce the model calculation cost; the backbone network of the improved EAST model can be dynamically selected and improved according to the actual scenario requirements, which is not limited in this embodiment. In addition, in this embodiment, the improved EAST model also modifies the output head of the original EAST model to handle inclined text by adding support for rotated rectangles; during training, by predicting both rectangular and polygonal text boxes in the detection head, the adaptability of the improved EAST model to irregular text layouts is improved. The improved EAST model has a smaller computational cost than the lightweight CRAFT model and is suitable for processing pages containing a large number of pictures or a mixture of text and images; at the same time, the improved EAST model supports multi-scale input and is also applicable to the processing of documents with different resolutions.

[0037] Specifically, as Figure 3As shown, the corresponding text detection model is called according to the complex features for text area recognition, which specifically includes the following steps: S301: Determine the complex feature type; S302: If the complex feature is text density or non-standard text or paragraph overlap, call the lightweight CRAFT model; S303: If the complex feature is picture-type PDF or the number of tables or mixed text and graphics, call the improved EAST model; S304: Perform text area recognition through the corresponding text detection model.

[0038] In this embodiment, if the PDF file is not a complex document and does not contain any complex features, that is, the PDF file is a simple document of pure text type, the text can be directly extracted through the PDF parsing engine to improve the processing efficiency. For complex documents containing complex features, the corresponding text detection model is called according to the complex feature type for text area recognition, which can not only accurately identify the content in the complex document, avoid the problem of text missing detection in traditional OCR recognition methods, ensure the accuracy of the subsequent translation content, but also improve the text detection efficiency, and is extremely efficient and applicable.

[0039] S4: Extract the layout information of the text in the text area, and translate the text in the text area through the translation model to obtain the final translation result; Furthermore, after the text area is recognized, the layout information of the text in the text area is extracted. In this embodiment, the extraction of the layout information of the text in the text area of the complex document is realized through the pre-trained LayoutLMv3 model; for a simple document that does not contain complex features, the layout information of the text can be directly extracted through the PDF parsing engine to improve the processing efficiency.

[0040] Furthermore, the text in the text area is translated through the translation model. Preferably, the translation model includes the optimized mBART (Multilingual BART) model and the optimized T5 (Text-to-Text Transfer Transformer) model; wherein, the optimized mBART model and the optimized T5 model are obtained by optimizing the structures of the original mBART model and T5 model, specifically including: For the mBART model, on the one hand, a hierarchical attention enhancement module is introduced. An independent global attention head is introduced for the [CLS] token of each sentence, and the paragraph context is used as prior information. By adding a cross-sentence context attention mechanism in the encoder, not only can the context dependencies in the paragraph be captured, but also the coherence of the mBART model in long text translation can be improved. On the other hand, hierarchical position encoding is added based on the document logical structure. For example, at the paragraph level, block-level position information is added to capture cross-paragraph relationships; at the sentence level, sentence-level hierarchical position information is added to strengthen the in-sentence semantic modeling ability.

[0041] For the T5 model, on the one hand, inter-segment bidirectional modeling is introduced. For the problem of cross-page paragraph translation, a cross-segment bidirectional attention layer is added to transfer forward semantic constraints from the previous segment and introduce reverse consistency constraints in the subsequent segment. On the other hand, a text noise injection mechanism (enhancing the robustness of the T5 model by randomly replacing, inserting, or deleting some words) and a regularization mechanism (avoiding overfitting of the T5 model by increasing the Dropout ratio) are introduced.

[0042] Specifically, as Figure 4 shown, translating the text in the text area through the translation model includes the following steps: S401: Detect the delimiter in the text area and segment the text in the text area based on the delimiter; S402: Input the segmented text into the optimized mBART model for translation based on paragraphs to obtain the first translation result; S403: Input the first translation result into the optimized T5 model for optimization to obtain the second translation result; S404: Perform post-translation processing on the second translation result to obtain the final translation result.

[0043] Among them, the post-translation processing includes but is not limited to: spelling check and correction, duplicate detection and removal, term consistency processing, and cultural adaptation processing; the specific post-processing content can be adjusted routinely according to actual needs, and this embodiment does not make restrictions.

[0044] In this embodiment, by adopting the method of collaborative work of multiple translation models, it can not only support multi-modal translation of text, meet diverse translation needs, but also optimize the translation result, improve the accuracy of the translated text, and ensure the translation quality.

[0045] S5: Typeset the final translation result according to the layout information and customize the output target translation file.

[0046] Further, based on the text layout information extracted in step S4, the final translation result is typeset and adjusted so that the position information of the translated text is consistent with the original document; at the same time, the layout of elements such as images and tables in the translated document is adjusted, so that the finally output translated text file can retain the layout structure of the original document and improve the user's reading experience.

[0047] Further, customize the output target translated text file according to actual needs, including but not limited to PDF, Word, HTML or other document formats.

[0048] The PDF text translation method described in this embodiment can accurately identify the text content in the PDF file and improve the text detection efficiency by evaluating the complexity of the PDF file and correspondingly calling different text detection models for text detection; by adopting the method of multi-translation models working together, it can not only support multi-modal translation of text, meet diverse translation needs, but also optimize the translation result, improve the accuracy of the translated text, and ensure the translation quality; at the same time, restore the layout typesetting of the document content before outputting the translation, ensuring the consistency of the finally output translated text file and the original document in layout, thus improving the user's reading experience and being highly applicable.

[0049] Embodiment 2

[0050] As Figure 4 shown, the embodiment of the present invention provides a PDF text translation system 500 including complex features, and the system includes: an initialization module 501, a document judgment module 502, a text recognition module 503, a text translation module 504, and a translated text output module 505; wherein: The initialization module 501 is used to initialize the PDF parsing engine, read the PDF file and extract the basic information of the PDF file; In this embodiment, the PDF parsing engine is obtained by secondary development of the open-source library PDFium, and the extraction of the basic information of the PDF file specifically includes: Parse the PDF file through the PDF parsing engine to obtain the page tree structure of the PDF file; Traverse each object in each page of the PDF one by one according to the page tree structure, read the metadata, page information, font information, and image information of the PDF file, and obtain the basic information of the PDF file.

[0051] The initialization module 501 includes: a PDF file reading unit 5011 and a basic information extraction unit 5012; wherein: The above PDF file reading unit 5011 is configured to: initialize the PDF parsing engine and read the PDF file specified by the user; The above basic information extraction unit 5012 is configured to extract the basic information of the PDF file through a PDF parsing engine.

[0052] The document judgment module 502 is used to judge whether the PDF file is a complex document and record complex features; Preferably, in this embodiment, the complex features include: picture-type PDF, number of tables, text density, non-standard text, text and picture mixing, and paragraph overlap; judging whether the PDF file is a complex document specifically includes the following steps: According to the basic information of the PDF file, count the number of pictures or tables in the PDF file and detect whether the PDF file is a picture-type PDF; Based on the text density analysis method, evaluate the text density of each page in the PDF file; Detect whether there is non-standard text, text and picture mixing or paragraph overlap in the PDF file; Assign weights to the complex features and calculate the comprehensive score of the complex features in the PDF file; Compare the comprehensive score with a preset threshold. If the comprehensive score exceeds the preset threshold, it is determined that the PDF file is the complex document.

[0053] The document judgment module 502 includes: a complexity evaluation unit 5021 and a complex feature recording unit 5022; where: The above complexity evaluation unit 5021 is configured to detect complex features in the PDF file and judge whether the PDF file is a complex document; The above complex feature recording unit 5022 is configured to receive and record the complex features detected by the complexity evaluation unit 5021.

[0054] The text recognition module 503 is used to preprocess the images in the complex document and call the corresponding text detection model according to the complex features to identify the text area; Preferably, in this embodiment, preprocessing the images in the complex document includes: noise removal, image enhancement, and edge detection. The text detection models include a lightweight CRAFT model and an improved EAST model; among them, the lightweight CRAFT model uses MobileNetV3 as the backbone network; the improved EAST model adopts a dynamic backbone network selection strategy and modifies the output head of the EAST model to support the detection of rotated rectangles.

[0055] Specifically, when identifying the text area by calling the corresponding text detection model according to complex features, if the complex feature is text density or non-standard text or paragraph overlap, the lightweight CRAFT model is called for text area identification; if the complex feature is picture-type PDF or the number of tables or mixed text and graphics, the improved EAST model is called for text area identification.

[0056] The text recognition module 503 includes: an image preprocessing unit 5031 and a text area recognition unit 5032; where: The above-mentioned image preprocessing unit 5031 is configured to: preprocess the images in the complex document; The above-mentioned text area recognition unit 5032 is configured to: call the corresponding text detection model according to complex features for text area identification.

[0057] The text translation module 504 is used to extract the layout information of the text in the text area, and translate the text in the text area through a translation model to obtain the final translation result; Preferably, in this embodiment, the extraction of the text layout information in the complex document text area is implemented by a pre-trained LayoutLMv3 model; the translation model includes an optimized mBART (MultilingualBART) model and an optimized T5 (Text-to-Text Transfer Transformer) model; the translation of the text in the text area through the translation model includes the following steps: Detect the delimiters in the text area, and segment the text in the text area based on the delimiters; Input the segmented text into the optimized mBART model segment by segment for translation to obtain a first translation result; Input the first translation result into the optimized T5 model for optimization to obtain a second translation result; Perform post-processing on the second translation result to obtain the final translation result.

[0058] The text translation module 504 includes: a layout information extraction unit 5041 and a text translation unit 5042; where: The above-mentioned layout information extraction unit 5041 is configured to: extract the layout information of the text in the text area; The above-mentioned text translation unit 5042 is configured to: translate the text in the text area through a translation model to obtain the final translation result.

[0059] The translation output module 505 is used to typeset the final translation result according to the layout information and customize the output of the target translation file.

[0060] Further, based on the text layout information extracted by the layout information extraction unit 5041, the final translation result is typeset and adjusted so that the position information of the translated text is consistent with the original document; at the same time, the layout of elements such as images and tables in the translated document is adjusted, so that the finally output translation file can retain the layout structure of the original document and improve the user's reading experience.

[0061] The translation output module 505 includes: a typesetting adjustment unit 5051 and a file output unit 5052; where: The above-mentioned typesetting adjustment unit 5051 is configured to: typeset the final translation result according to the text layout information extracted by the layout information extraction unit 5041; The above-mentioned file output unit 5052 is configured to: customize the output format of the typeset document and output the target translation file.

[0062] In this embodiment, the complexity of the PDF file is evaluated by the document judgment module 502, and the text recognition module 503 calls different text detection models according to different complex features to detect the text area, which can accurately identify the text content in the PDF file and provide a basis for subsequent translation; the pre-trained LayoutLMv3 model is embedded in the text translation module 504, and a multi-translation model collaborative working method is adopted, which can not only support the multi-modal translation of text and improve the accuracy of the translated text, but also restore the layout typesetting of the document content before outputting the translation, ensuring the consistency of the finally output translation file and the original document in layout, thus improving the user's reading experience and being highly efficient and adaptable.

[0063] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the claims. It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for translating PDF text containing complex features, characterized in that: The steps include: Initialize the PDF parsing engine, read the PDF file and extract the basic information of the PDF file; Determine whether the PDF file is a complex document, and record the complex features; Preprocessing the image in the complex document, and calling a corresponding text detection model to perform text area recognition according to the complex features; Extracting layout information of the text in the text area, and translating the text in the text area through a translation model to obtain a final translation result; The final translation result is typeset according to the layout information, and a target translation file is outputted in a customized manner.

2. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The determining whether the PDF file is a complex document specifically includes: According to the basic information of the PDF file, count the number of pictures or tables in the PDF file to detect whether the PDF file is a picture-type PDF; Based on the text density analysis method, evaluating the text density of each page in the PDF file; Detect whether there is non-standard text, mixed text and graphics, or overlapping paragraphs in the PDF file; Assigning weights to the complex features and calculating a comprehensive score of the complex features in the PDF file; The comprehensive score is compared with a preset threshold value, and if the comprehensive score exceeds the preset threshold value, the PDF file is determined to be the complex document.

3. The method for translating a PDF text containing complex features according to claim 1, characterized in that: When determining whether the PDF file is a complex document, the determination may be performed based on a page as a unit, or based on a partial rectangular area in a page as a unit.

4. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The complex features include: text density, image-type PDF, number of tables, non-standard text, mixed text and image layout, and paragraph overlap; wherein, the non-standard text includes: text in images, text drawn by paths, text embedded in fonts, and text containing mixed languages ​​or special symbols.

5. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The text detection model includes a lightweight CRAFT model and an improved EAST model; wherein the lightweight CRAFT model adopts MobileNetV3 as the backbone network; the improved EAST model adopts a dynamic backbone network selection strategy and modifies the output head of the EAST model to support the detection of rotated rectangles.

6. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The calling of a corresponding text detection model to perform text region recognition according to the complex features specifically includes: If the complex feature is text density or non-standard text or paragraph overlap, a lightweight CRAFT model is called to perform text region recognition; If the complex feature is a picture-type PDF or the number of tables or a mixture of pictures and texts, the improved EAST model is called to perform text area recognition.

7. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The translation model includes an optimized mBART model and an optimized T5 model; and the translating of the text in the text area by the translation model specifically includes: Detecting separators in the text region, and segmenting text in the text region based on the separators; Inputting the segmented text into the optimized mBART model based on the segments for translation to obtain a first translation result; Inputting the first translation result into the optimized T5 model for optimization to obtain a second translation result; Post-translation processing is performed on the second translation result to obtain the final translation result.

8. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The extracting of basic information of the PDF file specifically includes: Parsing the PDF file by the PDF parsing engine to obtain a page tree structure of the PDF file; According to the page tree structure, objects in each page of the PDF are traversed page by page, metadata, page information, font information and image information of the PDF file are read, and basic information of the PDF file is obtained.

9. The method for translating a PDF text containing complex features according to claim 1, characterized in that: The preprocessing of the image in the complex document includes: noise removal, image enhancement and edge detection.

10. A PDF text translation system containing complex features, characterized in that: include: An initialization module is used to initialize the PDF parsing engine, read the PDF file and extract the basic information of the PDF file; A document judgment module, used to judge whether the PDF file is a complex document and record complex features; A text recognition module, used for preprocessing the image in the complex document and calling a corresponding text detection model to perform text area recognition according to the complex features; A text translation module, used to extract the layout information of the text in the text area, and translate the text in the text area through a translation model to obtain a final translation result; The translation output module is used to typeset the final translation result according to the layout information and customize the output target translation file.

Citation Information

Cited By

  • PDF (Portable Document Format) translation system for PDF content stream analysis based on AI (Artificial Intelligence) identification

    CN121010984A