PDF (Portable Document Format) translation system for PDF content stream analysis based on AI (Artificial Intelligence) identification

By using an AI-based PDF content stream parsing system, the problems of preserving the vector characteristics of complex formulas and ensuring consistency in translation format have been solved, achieving high-precision PDF document translation.

CN121010984AInactive Publication Date: 2025-11-25BEIJING TWEET INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511118757.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies fail to effectively preserve the vector characteristics of complex formulas and the consistency between the translated text and the original format, resulting in low translation accuracy of PDF documents.

Method used

The system employs an AI-based PDF content stream parsing system, which includes a preprocessing module, a content parsing module, an AI visual analysis module, a paragraph recognition module, a formula processing module, a translation engine module, an adaptive typesetting module, and a nested structure processing module. Through the collaborative efforts of these modules, the entire process from PDF preprocessing to final generation is optimized, ensuring the integrity of vector formulas and the consistency of the translated text format.

Benefits of technology

It significantly improves the accuracy and precision of PDF document translation, solves the problems of garbled text and preservation of vector characteristics of complex formulas in traditional PDF translation, and ensures the clarity and consistency of the translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010984A_ABST
    Figure CN121010984A_ABST
Patent Text Reader

Abstract

The invention relates to the field of PDF translation, in particular to a PDF translation system for PDF content stream analysis based on AI recognition. The method comprises the following steps: generating a standardized PDF document through a preprocessing module; generating a structured intermediate representation sample database through a content analysis module; an actual layout director is generated through an AI visual analysis module; extracting a paragraph document analyzer through a paragraph recognition module; extracting a vector formula analyzer through a formula processing module; generating a target language translation and extracting a formula placeholder through a translation engine module; calculating an optimal zoom factor of the paragraph through a self-adaptive typesetting module, and generating a final typesetting sequence; outputting a cross-level text continuity verification result through a nested structure processing module; and generating a format-preserving translation PDF document through a PDF generation module. The method is used for solving the problem of reserving vector characteristics of typesetting messy codes and complex formulas, and the precision of PDF literature translation is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of PDF translation, and in particular to a PDF translation system based on AI recognition of PDF content flow analysis. BACKGROUND

[0002] With the acceleration of globalization and the increasing frequency of cross-language academic exchanges, the demand for PDF document translation is growing exponentially. According to statistics, academic publishing, multinational corporations and international organizations need to handle millions of pages of multilingual PDF files every day, covering technical manuals, scientific papers and legal contracts, etc. However, traditional PDF translation technology faces technical bottlenecks. Violent font scaling destroys the original layout coordination, resulting in a significant reduction in document readability. Formulas lose the vector scaling feature and become blurred and distorted after enlargement, which cannot meet the high-precision requirements of academic literature. Paragraphs are fragmented, and the translated text is semantically broken.

[0003] Patent application publication CN120163169A discloses a PDF text translation method and system containing complex features. The present application belongs to the technical field of text translation and provides a PDF text translation method and system containing complex features. The method includes initializing a PDF parsing engine, reading a PDF file and extracting the basic information of the PDF file, determining whether the PDF file is a complex document, recording complex features, preprocessing images in the complex document, and calling a corresponding text detection model for text region recognition according to the complex features. The layout information of the text in the text region is extracted, the text in the text region is translated by a translation model, and the final translation result is obtained. After typesetting according to the layout information, the target translated text file is output. The present application can intelligently identify complex content in PDF documents and completely preserve the original document format during translation. It also supports multilingual translation to ensure the accuracy of PDF document translation.

[0004] Patent application publication CN115455931A discloses a line break identifier. The present application discloses a line break identifier, which includes a rule and a semantic model combination method for line break identification. The rule-based method is used to identify line breaks. For cases that can be judged by rules, the result is directly returned. When the rule cannot be used for judgment, the semantic model is used to determine the output result. The present application has the following advantages: it improves the accuracy of pdf to word conversion, saves manual labor time when dealing with incorrect line breaks, improves document quality, and ensures the quality of document analysis and translation in the later stage.

[0005] Therefore, the prior art has the following problems:

[0006] Existing technologies do not consider the preservation of vector characteristics of complex formulas, nor do they take into account the differences between the translated text and the original version, resulting in low accuracy of translated PDF documents. Summary of the Invention

[0007] To address this, the present invention provides a PDF translation system based on AI-based PDF content stream parsing, which overcomes the problems of existing technologies failing to consider the preservation of vector characteristics of complex formulas and failing to consider the differences between the translated text and the original format, resulting in low accuracy of the translated PDF documents.

[0008] To achieve the above objectives, this invention provides a PDF translation system based on AI-based PDF content stream parsing, comprising:

[0009] The preprocessing module repairs and generates standardized PDF documents based on the original PDF file;

[0010] The content parsing module, which is connected to the preprocessing module, generates a structured intermediate representation sample database based on the standardized PDF document;

[0011] The AI ​​visual analysis module, which is connected to the content parsing module, generates an actual layout guide based on the structured intermediate representation sample database;

[0012] The paragraph recognition module is connected to the AI ​​visual analysis module and the content parsing module respectively, and extracts the paragraph document analyzer based on the structured intermediate representation sample database corresponding to the actual layout guide;

[0013] The formula processing module, which is connected to the content parsing module, extracts a vector formula analyzer based on the structured intermediate representation sample database;

[0014] A translation engine module, which is connected to the paragraph recognition module and the formula processing module, generates a target language translation based on the paragraph document analyzer; and extracts formula placeholders based on the vector formula analyzer.

[0015] An adaptive typesetting module, which is connected to the translation engine module, calculates the optimal scaling factor of the paragraph based on the bounding box of the original paragraph obtained from the standardized PDF document and the length of the translation, and generates the final typesetting sequence based on the optimal scaling factor of the paragraph.

[0016] The nested structure processing module, which is connected to the content parsing module, outputs cross-level text continuity verification results based on the structured intermediate representation sample database;

[0017] The PDF generation module, connected to the adaptive typesetting module and the nested structure processing module, generates a version-preserving translated PDF document based on the final typesetting sequence and the cross-level text continuity verification results.

[0018] Furthermore, the formula processing module includes: a subscript detection unit that generates subscript identifiers based on the font size variation of adjacent characters in the acquired structured intermediate representation sample database; an offset calculation unit connected to the subscript detection unit that generates fragment offsets based on the acquired reference coordinates; and a vector reconstruction unit connected to the offset calculation unit that generates a vector formula analyzer based on the offsets.

[0019] Furthermore, the adaptive typesetting module includes: a preprocessing calculation unit, which generates an optimal scaling factor for each paragraph through iterative search based on the bounding box of the original paragraph and the length of the translation; a boundary expansion unit, which is connected to the preprocessing calculation unit, and generates a horizontal or vertical expansion instruction based on the available space on the page when the scaling factor is lower than a preset threshold; a global optimization unit, which is connected to the preprocessing calculation unit and the boundary expansion unit, and calculates a global mode scaling factor based on all obtained paragraph scaling factors and performs normalization processing; and a rendering application unit, which is connected to the preprocessing calculation unit, the boundary expansion unit, and the global optimization unit, and generates a final typesetting sequence based on the optimal scaling factor of each paragraph, the horizontal or vertical expansion instruction, and the normalized scaling factor.

[0020] Furthermore, the nested structure processing module includes a call stack construction unit that generates a hierarchical path based on the object reference relationships obtained from the structured intermediate representation sample database; and a continuity verification unit connected to the call stack construction unit that generates the continuity verification result based on the hierarchical path.

[0021] Furthermore, the preprocessing module is used to generate a scanned skip marker based on the proportion of non-text instructions acquired.

[0022] Furthermore, the preprocessing module is used to generate a scanning determination result based on the proportion of the acquired text drawing instructions.

[0023] Furthermore, the paragraph recognition module includes a punctuation rule unit, which generates continuity determination based on the commas and colons at the end of lines obtained by AI recognition; and a position rule unit, which generates soft line break determination based on the proportion of line spacing less than the character height obtained by AI recognition.

[0024] Furthermore, the PDF generation module includes a font subsetation unit, which generates simplified font resources based on the acquired actual characters used; and a directory migration unit, which generates translated navigation bookmarks based on the acquired original directory structure.

[0025] Furthermore, the translation engine module includes generating segmented text based on the acquired semantic delimiters; and generating an optimized translation based on the segmented text.

[0026] Furthermore, the preprocessing module generates a corrected PDF based on the obtained object reference errors.

[0027] Compared with the prior art, the beneficial effects of the present invention are that it provides a PDF translation system based on AI-based PDF content stream parsing. Through the synergistic effect of various modules, it achieves full-process optimization from PDF preprocessing to final generation, solves the problems of garbled text and preservation of vector characteristics of complex formulas in traditional PDF translation, and improves the accuracy of the translation and the consistency of the format, thus significantly improving the accuracy of PDF document translation.

[0028] In particular, this invention, through a preprocessing module based on the original PDF file, utilizes document processing libraries such as pymupdf to load the file and call repair functions to automatically repair structural errors such as corrupted object references and non-standard filters, generating standardized PDF documents. Simultaneously, a scanned version detection mechanism determines whether it is a scanned or native electronic version based on the similarity of page images before and after text drawing instructions, ensuring the structural integrity and standardization of the input file. This provides a reliable foundation for subsequent processing, enhances the system's compatibility with various non-standard PDFs, avoids parsing failures caused by problems with the original file, and improves process adaptability by adopting corresponding processing strategies for different types of PDFs.

[0029] In particular, this invention utilizes a content parsing module based on standardized PDF documents, leveraging PDF content extraction and analysis tools such as the pdfminer.six underlying engine to parse the content page by page and character by character, recording precise coordinates, size, font, color, and other micro-information of each character. Through an intermediate layer builder, such as ILCreater, a structured intermediate representation sample database is constructed. This database fully preserves the PDF's text drawing instructions, resource information, and layout details, providing comprehensive data support for subsequent AI visual analysis, paragraph recognition, and other modules. This avoids analysis errors caused by missing information, ensuring the accuracy and precision of subsequent processing and forming the foundation for the efficient operation of the entire system.

[0030] In particular, this invention renders PDF pages as images using an AI visual analysis module, inputs them into an AI visual model, identifies macro-layout areas such as text blocks, tables, images, and formulas, and outputs an actual layout guide, providing a layout basis for paragraph recognition. The paragraph recognition module combines the actual layout guide with the micro-character flow information from the content parsing module to achieve accurate segmentation of paragraphs, lists, and headings under complex layouts. It no longer relies on fixed rules or text order, significantly improving paragraph recognition accuracy and solving the problem of chaotic paragraph segmentation in complex scenarios such as multi-column layouts and mixed text and images using traditional methods.

[0031] In particular, this invention, through its formula processing module based on a structured intermediate representation sample database, identifies vector formulas using a hybrid strategy of AI layout, special font matching, and character-level heuristic rules. It then processes subscripts, offsets, and fragment merging through a subscript detection unit, an offset calculation unit, and a vector reconstruction unit. During translation, formulas are represented by placeholders and subsequently restored; during typesetting, formulas are treated as a single box. This process ensures that formulas are always clear and scalable vector graphics, solving the problems of traditional tools erroneously translating or blurring formulas, and guaranteeing the professionalism of documents containing a large number of formulas.

[0032] In particular, this invention uses an adaptive typesetting module and an iterative search mechanism to find the minimum scaling ratio for each paragraph that can accommodate all its translated content. A backup mechanism is in place: when the scaling ratio decreases to a preset threshold, the text is considered too small and affects readability; therefore, the text is no longer shrunk further, and the paragraph's container is expanded. If content still overflows within the expanded paragraph's bounding box, the line break restrictions for English text are relaxed. Standardization prevents paragraphs with very little content, such as those with only one word, from receiving a large scaling ratio, while most paragraphs with a lot of content are shrunk to a smaller scaling ratio. Standardization ensures a more uniform and harmonious overall font size in the document, avoiding inconsistent font sizes on the page. By obtaining the optimal scaling ratio for each paragraph, the document is reformatted to generate the final typesetting sequence. This solves the layout chaos caused by the difference in length between the translated and original texts, ensuring the translation is complete and aesthetically pleasing while maintaining the original layout, thus balancing readability and layout consistency.

[0033] In particular, this invention records the xobj ID of a character through a nested structure processing module, returns the corresponding xobj during write-back, generates hierarchical paths through a call stack construction unit, and verifies cross-level text continuity based on the path through a continuity verification unit, handling the inheritance of parent xobj states and resources by child xobjs. This ensures that characters from different xobjs are not incorrectly classified into the same paragraph, guaranteeing the accuracy and consistency of text processing, and maintaining document structure stability even in complex xobj nesting scenarios.

[0034] In particular, this invention, through its PDF generation module, uses atomic typesetting units and cross-level text continuity verification results to generate new drawing instructions from the typeset content. These instructions are then concatenated with the original non-text instructions to replace the original content stream. A font subsetation unit simplifies font resources, and a directory migration unit preserves the original directory navigation function, ultimately generating a format-preserving translated PDF. This module ensures that the translated PDF is highly consistent with the original text in terms of layout, clarity, and navigation, outputting a high-quality translated document. Attached Figure Description

[0035] Figure 1 This is a structural block diagram of a PDF translation system based on AI-recognized PDF content stream parsing, according to an embodiment of the present invention.

[0036] Figure 2 Structural block diagram of the formula processing module in this embodiment of the invention;

[0037] Figure 3 Structural block diagram of the adaptive typesetting module in this embodiment of the invention;

[0038] Figure 4 A structural block diagram of the nested structure processing module in an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0040] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0041] Please see Figure 1 , Figure 2 , Figure 3 as well as Figure 4 The diagram shows the structural block diagrams of the PDF translation system based on AI-recognized PDF content stream parsing according to an embodiment of the present invention; the structural block diagram of the formula processing module according to an embodiment of the present invention; the structural block diagram of the adaptive typesetting module according to an embodiment of the present invention; and the structural block diagram of the nested structure processing module according to an embodiment of the present invention.

[0042] This invention provides a PDF translation system based on AI-based PDF content stream parsing, comprising:

[0043] The preprocessing module repairs and generates standardized PDF documents based on the original PDF file;

[0044] The content parsing module, which is connected to the preprocessing module, generates a structured intermediate representation sample database based on the standardized PDF document;

[0045] The AI ​​visual analysis module, which is connected to the content parsing module, generates an actual layout guide based on the structured intermediate representation sample database;

[0046] The paragraph recognition module is connected to the AI ​​visual analysis module and the content parsing module respectively, and extracts the paragraph document analyzer based on the structured intermediate representation sample database corresponding to the actual layout guide;

[0047] The formula processing module, which is connected to the content parsing module, extracts a vector formula analyzer based on the structured intermediate representation sample database;

[0048] A translation engine module, which is connected to the paragraph recognition module and the formula processing module, generates a target language translation based on the paragraph document analyzer; and extracts formula placeholders based on the vector formula analyzer.

[0049] An adaptive typesetting module, which is connected to the translation engine module, calculates the optimal scaling factor of the paragraph based on the bounding box of the original paragraph obtained from the standardized PDF document and the length of the translation, and generates the final typesetting sequence based on the optimal scaling factor of the paragraph.

[0050] The nested structure processing module, which is connected to the content parsing module, outputs cross-level text continuity verification results based on the structured intermediate representation sample database;

[0051] The PDF generation module, connected to the adaptive typesetting module and the nested structure processing module, generates a version-preserving translated PDF document based on the final typesetting sequence and the cross-level text continuity verification results.

[0052] Through the synergistic effect of various modules, the entire process from PDF preprocessing to final generation has been optimized, solving the problems of garbled text and preservation of vector characteristics of complex formulas in traditional PDF translation. It also improves the accuracy of the translation and the consistency of the format, significantly enhancing the precision of PDF document translation.

[0053] Specifically, the formula processing module includes: a subscript detection unit that generates subscript identifiers based on the font size changes of adjacent characters in the acquired structured intermediate representation sample database; an offset calculation unit connected to the subscript detection unit that generates fragment offsets based on the acquired reference coordinates; and a vector reconstruction unit connected to the offset calculation unit that generates a vector formula analyzer based on the offsets.

[0054] In this embodiment, the subscript detection unit of the formula processing module identifies subscripts and generates subscript identifiers based on the font size changes of adjacent characters in the structured intermediate representation sample database. The subscript identification process includes identifying the current character as a subscript and generating a subscript identifier when the previous character is not a subscript and the current character's font size is less than 0.79 times that of the previous character, or when the previous character is a subscript and the current character's font size is greater than 1.1 times that of the previous character. The offset calculation unit calculates the vertical coordinate offset of special characters such as subscripts based on the character's reference coordinates. The vector reconstruction unit merges fragmented formula segments into complete vector formulas based on these offsets and spacing information, generating a vector formula analyzer. The formula processing module improves the accuracy of formula recognition and preserves its integrity through a hybrid strategy of AI layout, special font matching, and character-level heuristic rules.

[0055] Based on a structured intermediate representation sample database, the formula processing module identifies vector formulas using a hybrid strategy of AI layout, special font matching, and character-level heuristics. Subscript detection, offset calculation, and vector reconstruction units handle subscripts, offsets, and fragment merging. During translation, formulas are represented by placeholders and subsequently restored. In typesetting, formulas are treated as a single, unified box. This process ensures that formulas are always clear and scalable vector graphics, resolving the problems of incorrect or blurred formula translations found in traditional tools, and guaranteeing the professionalism of documents containing numerous formulas.

[0056] The adaptive typesetting module includes: a preprocessing calculation unit, which generates an optimal scaling factor for each paragraph through iterative search based on the bounding box of the original paragraph and the length of the translation; a boundary expansion unit, which is connected to the preprocessing calculation unit and generates horizontal or vertical expansion instructions based on the available space on the page when the scaling factor is lower than a preset threshold; a global optimization unit, which is connected to the preprocessing calculation unit and calculates a global mode scaling factor based on all obtained paragraph scaling factors and performs normalization processing; and a rendering application unit, which is connected to the preprocessing calculation unit, the boundary expansion unit, and the global optimization unit and generates an atomic typesetting unit sequence based on the optimal scaling factor of each paragraph, the horizontal or vertical expansion instructions, and the normalized scaling factor.

[0057] In this embodiment, the preprocessing calculation unit of the adaptive typesetting module finds a minimum scaling ratio for each paragraph that can accommodate all its translated content through an iterative search mechanism. In this embodiment, starting with an initial scaling factor of 1.0 (i.e., no scaling), it checks whether all text content can be placed within the original bounding box of the paragraph. If the content overflows, the scaling ratio is gradually reduced, with a preset reduction of 0.05 or 0.1. Then, the layout is retried until a scaling ratio that can accommodate all content is found, which is determined as the optimal scaling factor.

[0058] In this embodiment, the boundary expansion unit includes a first backup mechanism: when the scaling ratio decreases to a preset threshold, the text is considered too small, affecting readability. In this case, a backup logic is triggered, including ceasing further text shrinking and instead attempting to expand the paragraph's container. The scaling ratio threshold is set to 0.7. The process of generating horizontal or vertical expansion instructions based on available page space includes checking for available space below and to the right of the paragraph. If not blocked by other page elements, the paragraph's bounding box is expanded, and then the layout within the expanded bounding box is re-attempted. The boundary expansion unit also includes a second backup mechanism: if text content overflows within the expanded paragraph's bounding box when the optimal scaling factor is used to shrink the text, the line break restrictions for English text are relaxed, allowing words to break at any position. The system then checks again, starting with the initial scaling factor of 1.0, whether the text will overflow the bounding box.

[0059] In this embodiment, the global optimization unit calculates the mode of the optimal scaling ratios for all paragraphs after the optimal scaling ratios for all paragraphs have been calculated, and then normalizes these optimal scaling ratios. The normalization process includes checking the optimal scaling ratios of all paragraphs; if the optimal scaling ratio of any paragraph is greater than the mode, the optimal scaling ratio of that paragraph is updated to the value corresponding to the mode. This normalization prevents paragraphs with very little content, such as those with only one word, from receiving a large scaling ratio, while most paragraphs with a lot of content are shrunk to a smaller scaling ratio. Normalization also makes the overall font size of the document more uniform and harmonious, avoiding inconsistent font sizes on the page.

[0060] In this embodiment, the rendering application unit obtains the optimal scaling ratio corresponding to the paragraph, reformats the document, and generates the final layout sequence.

[0061] The adaptive typesetting module solves the layout chaos caused by the difference in length between the translation and the original text, ensuring that the translation is complete and beautiful while maintaining the original layout, and taking into account both readability and layout consistency.

[0062] The nested structure processing module includes a call stack construction unit that generates a hierarchical path based on the object reference relationships obtained from the structured intermediate representation sample database; and a continuity verification unit connected to the call stack construction unit that generates the continuity verification result based on the hierarchical path.

[0063] In this embodiment, the call stack construction unit of the nested structure processing module generates a hierarchical path similar to a function call stack based on the reference relationship of xobj in the structured intermediate representation sample database. The continuity verification unit verifies the continuity of text in different xobj based on this hierarchical path, ensuring that characters in different xobj are not incorrectly classified into the same paragraph. At the same time, it processes the inheritance of the parent xobj's state and resources by the child xobj and outputs the cross-level text continuity verification result.

[0064] The nested structure processing module records the xobj ID of each character and returns the corresponding xobj during write-back. A call stack construction unit generates hierarchical paths, and a continuity verification unit verifies cross-level text continuity based on these paths, handling the inheritance of parent xobj states and resources from child xobjs. This ensures that characters from different xobjs are not incorrectly grouped into the same paragraph, guaranteeing the accuracy and consistency of text processing and maintaining document structure stability even in complex xobj nesting scenarios.

[0065] Specifically, the preprocessing module is used to generate a scanned skip mark based on the proportion of non-text instructions acquired.

[0066] In this embodiment, if at least 20% of the pages in a PDF are non-scanned, it is determined to be a native electronic PDF, and a scanned version skip mark is generated based on this.

[0067] Specifically, the preprocessing module is used to generate a scanning determination result based on the proportion of the acquired text drawing instructions.

[0068] In this embodiment, through the scan detection mechanism, if the similarity of the image converted from the page after removing all text drawing instructions is greater than 80% compared with that before removal, it is determined to be a scanned page; if more than 80% of the pages in the PDF are scanned, it is determined to be a scanned PDF.

[0069] Specifically, the paragraph recognition module includes a punctuation rule unit, which generates continuity determination based on the commas and colons at the end of lines obtained by AI recognition; and a position rule unit, which generates soft line break determination based on the proportion of line spacing less than the character height obtained by AI recognition.

[0070] In this embodiment, the paragraph recognition module receives the actual layout guide output by the AI ​​visual analysis module through a document image layout analysis library, such as LayoutParser, and combines it with the micro-character flow information obtained by the content parsing module to perform refined paragraph segmentation. Based on the macro-layout recognized by AI, it no longer simply concatenates characters according to text order or fixed rules, but combines heuristic rules, such as the recognition of table of contents “…”, to accurately identify paragraphs, lists, and headings under complex layouts, generating a paragraph document analyzer.

[0071] Specifically, the PDF generation module includes a font subsetation unit, which generates simplified font resources based on the actual characters used; and a directory migration unit, which generates translated navigation bookmarks based on the original directory structure.

[0072] In this embodiment, the font subset unit of the PDF generation module embeds only the characters required for typesetting, generating simplified font resources and reducing file size; the directory migration unit reads the directory of the original PDF and migrates it to the newly generated PDF, maintaining the consistency of navigation functionality. This module converts the atomic typesetting units generated by the adaptive typesetting module into PDF drawing instructions, concatenates them with the original non-text instructions, replaces the content stream of the corresponding xobj, and finally generates a version-preserving translated PDF document.

[0073] Specifically, the translation engine module includes generating segmented text based on the acquired semantic delimiters; and generating an optimized translation based on the segmented text.

[0074] In this embodiment, the translation engine module submits the text paragraphs extracted by the paragraph recognition module to an external large language model for translation, generating target language translations. For formulas recognized by the formula processing module, placeholders are generated and embedded in the translated text. When parsing the translation output, the placeholders are matched with regular expressions and restored to the original formulas. At the same time, rich text styles are processed, and placeholders are used for texts that are different from the basic styles to ensure that the styles remain consistent after translation.

[0075] Specifically, the preprocessing module generates a corrected PDF based on the obtained object reference errors.

[0076] In this embodiment, the preprocessing module loads the PDF based on the original PDF file using a document processing library, such as the pymupdf library, and calls repair functions, such as fix_null_xref and fix_filter, to automatically repair structural errors such as corrupted object references and non-standard filters, thereby generating a standardized PDF document.

[0077] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0078] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A PDF translation system based on AI-based PDF content stream parsing, characterized in that, include: The preprocessing module repairs and generates standardized PDF documents based on the original PDF file; The content parsing module, which is connected to the preprocessing module, generates a structured intermediate representation sample database based on the standardized PDF document; The AI ​​visual analysis module, which is connected to the content parsing module, generates an actual layout guide based on the structured intermediate representation sample database; The paragraph recognition module is connected to the AI ​​visual analysis module and the content parsing module respectively, and extracts the paragraph document analyzer based on the structured intermediate representation sample database corresponding to the actual layout guide; The formula processing module, which is connected to the content parsing module, extracts a vector formula analyzer based on the structured intermediate representation sample database; A translation engine module, which is connected to the paragraph recognition module and the formula processing module, generates a target language translation based on the paragraph document analyzer; And extract formula placeholders based on the vector formula analyzer; An adaptive typesetting module, which is connected to the translation engine module, calculates the optimal scaling factor of the paragraph based on the bounding box of the original paragraph obtained from the standardized PDF document and the length of the translation, and generates the final typesetting sequence based on the optimal scaling factor of the paragraph. The nested structure processing module, which is connected to the content parsing module, outputs cross-level text continuity verification results based on the structured intermediate representation sample database; The PDF generation module, connected to the adaptive typesetting module and the nested structure processing module, generates a version-preserving translated PDF document based on the final typesetting sequence and the cross-level text continuity verification results.

2. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The formula processing module includes a subscript detection unit, which generates subscript identifiers based on the font size variation of adjacent characters in the acquired structured intermediate representation sample database. And an offset calculation unit, which is connected to the index detection unit, generates fragment offsets based on the acquired reference coordinates; And a vector recombination unit, which is connected to the offset calculation unit, generates a vector formula analyzer based on the offset.

3. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The adaptive typesetting module includes a preprocessing calculation unit that generates the optimal scaling factor for each paragraph through iterative search based on the bounding box of the original paragraph and the length of the translation. And a boundary expansion unit, which is connected to the preprocessing calculation unit, generates a horizontal or vertical expansion instruction based on the available space of the page when the scaling factor is lower than a preset threshold; And a global optimization unit, which is connected to the preprocessing calculation unit and the boundary expansion unit, calculates the global mode scaling factor based on all the obtained paragraph scaling factors and performs normalization processing; And a rendering application unit, which is connected to the preprocessing calculation unit, the boundary expansion unit and the global optimization unit, generates the final layout sequence based on the optimal scaling factor of each paragraph, the horizontal or vertical expansion instruction and the scaling factor of the normalization process.

4. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The nested structure processing module includes a call stack construction unit, which generates hierarchical paths based on the object reference relationships obtained in the structured intermediate representation sample database; And a continuity verification unit, which is connected to the call stack construction unit, generates the continuity verification result based on the hierarchical path.

5. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The preprocessing module is used to generate a scanned skip mark based on the proportion of non-text instructions obtained.

6. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The preprocessing module is used to generate a scanning determination result based on the proportion of the acquired text drawing instructions.

7. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The paragraph recognition module includes a punctuation rule unit, which generates continuity determination based on the commas and colons at the end of lines obtained by AI recognition; And the position rule unit, which generates a soft line break determination based on the proportion of line spacing less than the character height obtained by AI recognition.

8. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The PDF generation module includes a font subsetation unit, which generates simplified font resources based on the actual characters used. And a directory migration unit that generates translated navigation bookmarks based on the acquired original directory structure.

9. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The translation engine module includes generating segmented text based on the acquired semantic delimiters; Optimized translations are generated based on the segmented text.

10. The PDF translation system based on AI-recognized PDF content stream parsing according to claim 1, characterized in that, The preprocessing module generates a corrected PDF based on the obtained object reference errors.

Citation Information

Patent Citations

  • Line feed identification method

    CN115455931A

  • Method and system for translating PDF (Portable Document Format) text containing complex features

    CN120163169A

  • Machine translation method for PDF file

    US20090030671A1

  • Methods and systems for multilingual document translation

    WO2024256866A1