Format conversion method of PDF (Portable Document Format) document, storage medium and computer equipment

By using the combination method of preset document parsing algorithm and pre-trained large model, the content blocks of PDF documents are extracted and converted, and the information loss problem during parsing and conversion in the prior art is solved, efficient and accurate PDF document format conversion is achieved, and the content and layout characteristics of the document are retained.

CN120146000APending Publication Date: 2025-06-13XIAN YOUFANG DIGITAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510039892.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When parsing and converting existing PDF document processing tools, it is difficult to accurately identify and retain complex content such as tables, images, formulas and multi-column text typesetting and layout information, resulting in information loss and incomplete content of converted documents.

Method used

The content blocks of the original PDF document are extracted using a preset document parsing algorithm, the content and layout features of the original PDF document are extracted according to the structure type of the content block, and the pre-trained large model is used for semantic conversion, and the content blocks in the target format are generated, and these content blocks are laid out in the target format document according to the location information of the original content block.

Benefits of technology

It realizes efficient and accurate parsing of PDF document content and converting it into a target format document, retaining the content integrity and layout characteristics of the original document, and improving the readability and secondary utilization efficiency of the converted document.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146000A_ABST
    Figure CN120146000A_ABST
Patent Text Reader

Abstract

The invention discloses a PDF document format conversion method, a storage medium and computer equipment. The format conversion method of the PDF document comprises the following steps: analyzing an original PDF document based on a preset document analysis algorithm to extract a plurality of original content blocks of the original PDF document; based on the structure type of each original content block, extracting document content and layout features included in the corresponding original content block; and performing semantic conversion on the document content and the layout features based on a pre-trained large model to obtain a plurality of target content blocks in a preset target format, and arranging each target content block in the target format document based on the position information of the original content block corresponding to each target content block. By means of the method, the content of the PDF document can be efficiently and accurately analyzed and converted into the corresponding target format document, the target format document can keep the document content and layout characteristics of the original PDF document, and the integrity and readability of the converted document are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method for converting the format of PDF documents, a storage medium, and a computer device. Background Art

[0002] In modern document management and processing workflows, the PDF (Portable Document Format) is widely used in the production, transmission, and storage of various official documents due to its high compatibility and versatility. PDF documents can maintain the original file format, layout, and fonts, unaffected by software or operating system differences, thus ensuring document consistency and readability. However, with the rapid development of artificial intelligence technology and the increasing demand for digital office work, the existing PDF document processing methods are gradually showing their limitations, especially in the parsing and conversion of document content.

[0003] Currently, although there are various PDF content parsing and conversion tools on the market, most of these tools can only perform simple text extraction and format conversion. For complex content contained in PDF documents, such as tables, images, formulas, multi-column text, etc., the processing effects are not ideal. Specifically, when parsing PDF documents, these tools often can only perform basic text copying and pasting operations, and it is difficult to accurately identify and retain the typesetting and layout information in the original document. Therefore, when processing PDF documents containing a large number of tables, pictures, or complex typesetting, these tools are likely to cause information loss, making the integrity and readability of the content of the converted document greatly reduced when reused. In addition, the existing PDF document parsing and conversion technologies also face problems such as insufficient intelligent processing capabilities and poor multi-language adaptability. Most tools lack the ability to intelligently identify and classify document content and cannot perform targeted processing according to the specific characteristics of the document. At the same time, when processing multi-language documents, these tools often cannot automatically adapt the corresponding content output methods, resulting in possible garbled characters or format errors in the document content during the conversion process, further affecting the use effect of the document. Summary of the Invention

[0004] This application mainly provides a method for converting the format of PDF documents, a storage medium, and a computer device, aiming to solve the technical problem of low accuracy in PDF document parsing.

[0005] To solve the above technical problems, the technical solution adopted in this application is: to provide a method for format conversion of PDF documents. The method for format conversion of PDF documents includes: parsing the original PDF document based on a preset document parsing algorithm to extract several original content blocks of the original PDF document; extracting the document content and layout features included in the corresponding original content block based on the structure type of each original content block; performing semantic conversion on the document content and the layout features based on a pre-trained large model to obtain several target content blocks in a preset target format; and laying out each target content block in the target format document based on the position information of the original content block corresponding to each target content block to obtain the target format document.

[0006] In some embodiments, the preset document parsing algorithm includes a first parsing algorithm based on the PyMuPDF library and / or a second parsing algorithm based on the PDFPlumber library; the parsing of the original PDF document based on the preset document parsing algorithm to extract several original content blocks of the original PDF document includes: determining the structure type, the position information, and the element content of several page elements in the original PDF document based on the first parsing algorithm and / or the second parsing algorithm to obtain the original content block corresponding to each page element.

[0007] In some embodiments, the structure type includes at least one of text, image, table, and formula.

[0008] The extracting of the document content and layout features included in the corresponding original content block based on the structure type of each original content block includes: determining that the structure type of the original content block is text, and extracting the text structure information, text style information, and text content corresponding to the original content block based on a preset text processing algorithm; determining that the structure type of the original content block is an image, and extracting the image file and the image content description corresponding to the original content block based on a preset image processing algorithm and generating the acquisition path corresponding to the image file; determining that the structure type of the original content block is a table, and extracting the table structure information and the cell content corresponding to the original content block based on a preset table processing algorithm; determining that the structure type of the original content block is a formula, and extracting the formula content corresponding to the original content block based on a preset formula processing algorithm.

[0009] In some embodiments, extracting the text structure information, text style information, and text content corresponding to the original content block based on a preset text processing algorithm includes: recognizing the layout structure of the original content block based on the structure parsing algorithm provided by the PDFPlumber library to obtain the text structure information; copying the text or performing text character recognition on the element content of the page element corresponding to the original content block to obtain the text content; and parsing the context relationship in the text content based on a preset semantic analysis model to determine the text style information.

[0010] In some embodiments, extracting the image file and image content description corresponding to the original content block based on a preset image processing algorithm and generating the acquisition path corresponding to the image file includes: copying or cropping the image of the original content block to obtain the image file; when the original content block includes text information, copying or performing character recognition on the text information to obtain the image content description; when the original content block does not include text information, selectively inputting the image file into a preset visual understanding model to generate the image content description; saving the image file in a preset storage path, and creating a path index corresponding to the image file based on the preset storage path to generate the acquisition path.

[0011] In some embodiments, extracting the table structure information and cell content corresponding to the original content block based on a preset table processing algorithm includes: recognizing the original content block based on the table parsing algorithm provided by the Camelot library to obtain the table structure information and the cell content; or performing positioning detection on the rectangular area of the original content block and merging the rectangular areas to determine the table structure information, and recognizing each rectangular area to obtain the corresponding cell content.

[0012] In some embodiments, extracting the formula content corresponding to the original content block based on a preset formula processing algorithm includes: performing formula recognition on the original content block based on a preset formula recognition model and converting the corresponding formula recognition result into the formula format in the preset target format.

[0013] In some embodiments, laying out each of the target content blocks in a target format document based on the position information of the original content blocks corresponding to each of the target content blocks to obtain the target format document includes: calculating the layout positions of each of the target content blocks in the target format document based on a preset geometric algorithm provided by the Shapely library and the position information of the corresponding original content blocks; arranging each of the target content blocks at the calculated layout positions in the target format document; and in the target format document, adjusting the typesetting of the target content blocks according to the layout features of the target content blocks to obtain the target format document.

[0014] To solve the above technical problem, another technical solution adopted by this application is: providing a storage medium on which program data is stored, characterized in that when the program data is executed by a processor, the steps of the PDF document format conversion method as described above are implemented.

[0015] To solve the above technical problem, another technical solution adopted by this application is: providing a computer device, which includes a processor and a memory connected to each other, the memory stores a computer program, and when the processor executes the computer program, the steps of the PDF document format conversion method as described above are implemented.

[0016] The beneficial effects of this application are as follows: Different from the prior art, this application discloses a PDF document format conversion method, a storage medium, and a computer device. This application first parses the original PDF document through a preset document parsing algorithm to extract a number of original content blocks, then extracts the document content and layout features according to the structural type of each original content block, such as text, image, table, formula, etc., and then uses a pre-trained large model to perform semantic conversion on these document content and layout features to obtain a number of target content blocks in a preset target format. Finally, according to the position information of the original content blocks corresponding to each target content block, these target content blocks are laid out in the target format document, thereby obtaining the target format document. It can efficiently and accurately parse the content of the PDF document and convert it into the corresponding target format document. The obtained target format document can retain the document content and layout features of the original PDF document, improving the integrity and readability of the converted document, and effectively ensuring the quality and secondary utilization efficiency of the target format document. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0018] Figure 1 It is a schematic flowchart of an embodiment of a method for format conversion of a PDF document provided by this application;

[0019] Figure 2 is Figure 1 a schematic flowchart of an embodiment of step 20 in the embodiment;

[0020] Figure 3 is Figure 1 a schematic flowchart of an embodiment of step 40 in the embodiment;

[0021] Figure 4 It is a schematic structural diagram of an embodiment of a storage medium provided by this application;

[0022] Figure 5 It is a schematic structural diagram of an embodiment of a computer device provided by this application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0024] The terms "first", "second", and "third" in the embodiments of this application are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0025] Referring to "embodiment" in this article means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0026] This application provides a method for converting the format of a PDF document. Refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the method for converting the format of a PDF document provided by this application. The method for converting the format of a PDF document includes:

[0027] Step 10: Parse the original PDF document based on a preset document parsing algorithm to extract several original content blocks of the original PDF document.

[0028] In this embodiment, the original PDF document is an electronic document in PDF format for subsequent conversion. The document includes at least one of elements such as text, images, tables, formulas, etc. The original PDF document can be one page or multiple pages, and it can be specifically created by users in various ways, including but not limited to PDF files generated by scanning paper documents, PDF files directly created using document editing software, and PDF-format electronic books, reports, contracts, etc. downloaded from the Internet. These documents need to maintain their original layout and content integrity during the subsequent conversion process to ensure that the converted documents can meet the usage requirements of users in different scenarios.

[0029] In this embodiment, the preset document parsing algorithm is a preset algorithm used to parse the structure and content of the original PDF document. The algorithm can be a parsing algorithm based on a specific library or a custom-developed parsing algorithm. The input original PDF document is parsed page by page through the preset parsing algorithm to identify and extract several original content blocks in the document. These original content blocks can be text blocks, image blocks, table blocks, etc., and they together constitute the overall content and structure of the original PDF document. In this way, it can be ensured that the content extracted from the original PDF document includes not only text information but also elements such as images, tables, and formulas, providing accurate data support for subsequent semantic conversion and the layout of the target format document.

[0030] Optionally, in some embodiments, the preset document parsing algorithm includes a first parsing algorithm based on the PyMuPDF library and / or a second parsing algorithm based on the PDFPlumber library.

[0031] In this alternative embodiment, the PyMuPDF library is an open-source Python library that provides powerful PDF document processing capabilities, including but not limited to page rendering, text extraction, image processing, and other functions. The use of the PyMuPDF library can significantly improve the efficiency and accuracy of document parsing. Especially when processing PDF documents containing complex elements such as tables, text, formulas, and multi-column text, it can provide more stable and reliable parsing results. The first parsing algorithm corresponding to this library, namely the parsing algorithm based on the PyMuPDF library, can perform a structured analysis of the page elements in the PDF document, identify different types of elements such as text, images, and tables, and extract their position information and content.

[0032] In this alternative embodiment, the PDFPlumber library is also an open-source Python library that can provide functions such as extracting text from PDF pages, extracting table data, obtaining image information, analyzing page layout, and searching for specific text. The second parsing algorithm corresponding to this library refers to the parsing algorithm based on the PDFPlumber library, which can perform more detailed analysis and extraction of the text in the PDF document, including attribute information such as the style, font size, and color of the text, as well as the layout and structure information of the text.

[0033] In this alternative embodiment, PDFPlumber focuses more on the extraction of text and tables in PDF files and the analysis of page layout. It provides functions such as extracting text, table, and image position information and can search for text in PDF. The functions of PyMuPDF are more comprehensive. It can not only extract text, images, and tables but also perform operations such as merging, splitting, encrypting, and decrypting PDF files. The preset document parsing algorithm based on the first parsing algorithm or the second parsing algorithm can more precisely process complex content in PDF documents, such as embedded images, tables, and formulas. By combining the use of PyMuPDF and / or PDFPlumber libraries, a comprehensive parsing of PDF documents can be achieved, ensuring that the content blocks extracted from the original document are both comprehensive and accurate, laying a solid foundation for subsequent format conversion.

[0034] Optionally, the original PDF document is parsed based on the preset document parsing algorithm to extract several original content blocks of the original PDF document, including: based on the first parsing algorithm and / or the second parsing algorithm, determining the structure type, position information, and element content of several page elements in the original PDF document to obtain the original content blocks corresponding to each page element.

[0035] In this alternative embodiment, the original PDF document may include various page elements, such as text blocks, image blocks, table blocks, etc. Different page elements have corresponding structure types, position information, and element content. Among them, the structure type refers to the type of page element, such as text, image, table, or formula, etc.; the position information refers to the specific position coordinates of the page element in the document; and the element content is the specific information contained in the page element, such as text content, image data, table row and column information, and cell content text, etc. By determining the detailed information of these page elements, each original content block in the original PDF document can be accurately extracted, providing a data basis for subsequent format conversion and content extraction.

[0036] Step 20: Based on the structure types of the respective original content blocks, extract the document content and layout features included in the corresponding original content blocks.

[0037] In this embodiment, the document content in the original content block refers to the specific content of text, image, table, or formula, such as text characters, pixel information of the image, and the corresponding text, image, or formula information contained in the cell. The document content is the essential information of the PDF document and is the key to ensuring the accuracy of the converted document content. The layout features include visual presentation attributes such as the arrangement layout of the document content and font paragraph styles. These features are crucial for maintaining the original appearance and reading experience of the document. By extracting this information, it can be ensured that the converted document is as close as possible to the original PDF document in terms of both content and visual effects.

[0038] In this embodiment, there are various ways to extract the document content included in the original content block. For example, it can be recognized through a specific recognition model, extraction rules can be set for extraction, or it can also be processed by combining model recognition methods and logical rules to extract the document content. For example, the text content can be extracted through a preset text recognition model, the image content can be analyzed and extracted through image processing technology, the table content can be recognized and structured using a table parsing algorithm, and the formula content can be parsed and converted using formula recognition technology. These extraction processes ensure the accuracy and integrity of the document content, providing high-quality input data for subsequent format conversion. The extraction of layout features involves analyzing the visual layout of the document content, including font size, color, paragraph spacing, line spacing, arrangement method, etc., as well as the relative position and size of images and tables. The extraction of these layout features helps to maintain the visual style and layout of the original document during the conversion process, making the target format document visually consistent with the original PDF document.

[0039] Optionally, refer to Figure 2, based on the structural types of the original content blocks, extract the document content and layout features included in the corresponding original content blocks, which can be executed according to the following steps:

[0040] Step 21: Determine that the structural type of the original content block is text, and extract the text structure information, text style information, and text content of the corresponding original content block based on a preset text processing algorithm.

[0041] Step 22: Determine that the structural type of the original content block is an image, and extract the image file and image content description of the corresponding original content block based on a preset image processing algorithm, and generate the acquisition path corresponding to the image file.

[0042] Step 23: Determine that the structural type of the original content block is a table, and extract the table structure information and cell content of the corresponding original content block based on a preset table processing algorithm.

[0043] Step 24: Determine that the structural type of the original content block is a formula, and extract the formula content of the corresponding original content block based on a preset formula processing algorithm.

[0044] In this optional embodiment, the extraction methods of the document content and layout features in different cases where the structural types of the original content blocks are text, image, table, and formula are specified, so as to perform precise processing on different types of original content blocks, and can automatically adapt the corresponding extraction methods according to the structural type to ensure the accuracy and reliability of the content.

[0045] In this optional embodiment, text structure information, text style information, and text content can be extracted from the original content blocks of the text type. Among them, the text structure information refers to the organizational structure of the text, such as single-column, double-column, triple-column, block, etc.; the text style information includes the style attributes of text, paragraphs, lists, etc., such as the font, color, font size, glyph, and font effects of the text, the alignment method, indentation, line spacing, etc. of the paragraphs; the text content is the actual text information included in the text block. By extracting these detailed text-related information, it can be ensured that the converted document is consistent with the original PDF document in terms of text content and format.

[0046] In this alternative embodiment, an image file, an image content description, and an image acquisition path can be extracted from the original content block of the image type. Among them, the image file refers to the digital file of the image, which can be common picture formats such as JPEG, PNG, etc. It can be an image obtained by copying or a captured image. According to the image file, attribute information such as the size, color depth, and resolution of the corresponding image can also be obtained. When necessary, these attributes can be described by the image content description; the image content description is the explanatory text of the picture content, which can be the text appended to the picture, the text appended below the picture, the surrounding text, the watermark text, the overset text, or the underset text, etc., as long as these texts are the text content specifically used to describe the corresponding image file in the original PDF document; the image acquisition path refers to the location information of the image file in the file system and the network end, such as the local path, the network URL (uniform resource locator), etc. Through these information, it can be ensured that the image content is accurately presented in the converted document and can be conveniently managed and used accordingly.

[0047] In this alternative embodiment, table structure information and cell content can be extracted from the original content block of the table type. Among them, the table structure information refers to the layout and format of the table, including the number of rows, the number of columns, the merging situation of cells, and the styles of the table border and shading, etc.; the cell content is the data contained in each cell of the table, such as text, images, formulas, etc. And the characteristics such as the font, color, and alignment of the text and formulas, as well as the characteristics such as the size, color depth, and color mode of the images, can also be extracted together to maintain the visual style of the original document. By extracting these detailed table-related information, it can be ensured that the layout and data of the table in the converted document are consistent with the original PDF document.

[0048] In this alternative embodiment, formula content can be extracted from the original content block of the formula type. Among them, the formula content refers to the specific expression form of mathematical formulas, chemical equations, or other professional formulas, including symbols, operators, variables, and constants in the formula, etc. The extraction of these information is crucial for maintaining the accuracy of scientific and technical content in the document. In addition, the visual presentation characteristics of the formula, such as font style, size, color, etc., can also be extracted to ensure that the converted document is visually consistent with the original PDF document.

[0049] In this alternative embodiment, the corresponding preset text processing algorithm, preset image processing algorithm, preset table processing algorithm, and preset formula processing algorithm are algorithms used for optimizing and processing different types of original content blocks. The preset text processing algorithm can process the text in the text block, including identifying and extracting the structural and styling information of the text, as well as the text content itself. The preset image processing algorithm focuses on processing the image block, can extract the image file and generate a description of the image content, and at the same time provides the image acquisition path. The preset table processing algorithm is responsible for parsing the table block, extracting the structural information of the table and the cell content, ensuring the accuracy and integrity of the table data. The preset formula processing algorithm is specifically used to process the formula block, extract the formula content and retain its visual presentation features.

[0050] In this alternative embodiment, these algorithms can be obtained by training a learning model or developed based on specific rules. Together, they form a comprehensive processing system that can handle the conversion requirements of various complex documents. The different forms of document content and layout features extracted by them, namely the text structure information, text styling information, and text content of the original content block of the above text type, the image file and image content description of the original content block of the image type, the table structure information and cell content of the original content block of the table type, the formula content of the original content block of the formula type, etc. The effective extraction of these document content and layout features can ensure that the converted document is consistent with the original PDF document in terms of specific content and layout, maintain the original structure and semantics, meet the actual needs of complex document parsing and conversion, and maintain content integrity and layout accuracy.

[0051] Optionally, in some embodiments, based on the preset text processing algorithm, extracting the text structure information, text styling information, and text content of the corresponding original content block specifically includes: based on the structure parsing algorithm provided by the PDFPlumber library, identifying the layout structure of the original content block to obtain the text structure information; performing text copying or text character recognition on the element content of the page element corresponding to the original content block to obtain the text content; based on the preset semantic analysis model, parsing the context relationship in the text content to determine the text styling information.

[0052] In this alternative embodiment, the structure parsing algorithm provided by the PDFPlumber library parses the original content blocks of the text type through the powerful parsing ability of the library to extract text position information. It can accurately analyze the text layout in the PDF document, including the hierarchical structures of elements such as paragraphs, lists, and headings. Through this algorithm, the text content in the document can be accurately reconstructed according to its layout in the original PDF document, ensuring that the converted document is consistent with the original document in format and structure. In addition, this algorithm can also identify the style information of the text, such as font, color, size, etc. These information are crucial for maintaining the visual style and reading experience of the document. By combining the structure parsing algorithm of the PDFPlumber library, in-depth analysis and extraction of complex text content in the PDF document can be achieved, ensuring the correct logical order of the content and providing solid data support for subsequent format conversion.

[0053] In this alternative embodiment, the acquisition of text content is achieved through the text copying function of the PDFPlumber library or by using OCR (Optical Character Recognition) technology for text character recognition. The text copying function can directly extract text from the PDF document, while the OCR technology can recognize the text in images or scanned documents. These two methods each have their advantages. The text copying function is suitable for PDF documents with clear text and good format, while the OCR technology is more effective when dealing with image-based text or low-quality scanned documents. Through these technologies, it can be ensured that the text content extracted from the original PDF document is both accurate and complete.

[0054] In this alternative embodiment, the preset semantic analysis model is a model used to analyze the context relationship of text content. It can determine the semantic structure relationship of the text content according to the vocabulary, phrases, and sentence structures in the text content, and can avoid being affected by factors such as headers, footers, binding lines, line number markings, or incorrect copying order during the text copying or recognition process of PDF files. It can effectively parse the context relationship in the text content, thereby ensuring the accuracy of the text content and being able to effectively analyze the text content, making the division of corresponding paragraphs and headings more accurate. The preset semantic analysis model can specifically be a dedicated natural language processing model (Natural Language Processing, NLP), such as a multi-modal large language model (Large Language Model, LLM) or other types of text analysis tools. By applying this preset speech analysis model, the accuracy of text content extraction and the depth of semantic understanding can be further improved, ensuring that the converted document is consistent with the original PDF document in terms of the logic and coherence of the content.

[0055] Optionally, in some embodiments, an image file and an image content description corresponding to the original content block are extracted based on a preset image processing algorithm, and an acquisition path corresponding to the image file is generated, specifically including: performing image copying or image cropping on the original content block to obtain the image file; when the original content block includes text information, copying the text information or performing character recognition to obtain the image content description; when the original content block does not include text information, selectively inputting the image file into a preset visual understanding model to generate the image content description; saving the image file in a preset storage path, and creating a path index corresponding to the image file based on the preset storage path to generate the acquisition path.

[0056] In this optional embodiment, for the original content block of the image type, the image copying or image cropping operation is the basic step for extracting the image file. When the corresponding image in the original PDF document can be copied, the image file can be directly extracted from the original PDF document through image copying, while image cropping is used to process the image embedded in the document that cannot be directly copied, and the image file is extracted through the cropping operation. These image files can be in vector or bitmap format and are saved as needed. When extracting the image file, if the original content block contains text information, such as figure captions or explanatory text, this information will also be copied or recognized through OCR technology to form the image content description. When the original content block does not contain text information, a preset visual understanding model can be used to analyze the image content, thereby generating a text description of the image content. The preset visual understanding model can specifically be an image recognition model based on deep learning, such as a model fine-tuned from a large language model that emphasizes visual functions like GPT-4V, or an all-round large language model like GPT-4o, or a visual model trained based on a convolutional neural network. Once the obtained image file is extracted, it will be saved in a preset storage path, which can be a path in the local file management system or a cloud storage path. The acquisition path corresponding to the image file is generated by creating a path index based on this storage path, ensuring that these images can be conveniently accessed and referenced in the converted document. For example, in a Markdown format file, the image can be referenced in the format of "!(image content description)[acquisition path]" through this acquisition path. Through these steps, the extraction of image content not only ensures the image quality but also ensures the accessibility and manageability of the image content.

[0057] Optionally, in some embodiments, table structure information and cell content corresponding to the original content block are extracted based on a preset table processing algorithm, specifically including: identifying the original content block based on the table parsing algorithm provided by the Camelot library to obtain table structure information and cell content; or performing positioning detection on the rectangular area of the original content block and merging the rectangular areas to determine the table structure information, and identifying each rectangular area to obtain the corresponding cell content.

[0058] In this optional embodiment, the Camelot library is a tool library specifically used to parse table content in PDF documents. By analyzing cell boundaries, text alignment, and spatial distribution, it can accurately identify table structures, especially performing well when dealing with borderless tables or complex merged cells. Through the table parsing algorithm provided by the Camelot library, the table content in the original PDF document can be accurately positioned and parsed, thereby extracting the structure information of the table, including layout features such as the number of rows, columns, and merged cells, as well as cell content, including text, numbers, or other data. The extraction of this information is crucial for maintaining the accuracy and integrity of table data in the document. In addition, the Camelot library also supports further processing of the table, such as data cleaning, format conversion, etc., to meet different data processing requirements. By using the Camelot library, structured table data can be effectively extracted from complex PDF documents, providing convenience for subsequent data analysis and processing.

[0059] In this optional embodiment, in addition to using the Camelot library to parse complex tables, for simple tables, the method of rectangular area detection can be adopted. First, perform positioning detection on the rectangular area of the original content block to determine the table structure information, and the boundaries of the table can be identified to determine the number of rows and columns. Then, by merging adjacent rectangular areas, the situation of merged cells can be identified. For each rectangular area, the cell content, including text, numbers, or other data, can be further identified. This method is especially suitable for documents with simple and regular table structures and can quickly and accurately extract table data. Through these steps, it can be ensured that the table content is accurately presented in the converted document and the visual style and data integrity of the original document are maintained.

[0060] Optionally, in some embodiments, formula content corresponding to the original content block is extracted based on a preset formula processing algorithm, specifically including: performing formula recognition on the original content block based on a preset formula recognition model and converting the corresponding formula recognition result into the formula format in a preset target format.

[0061] In this alternative embodiment, the preset formula recognition model is a model specifically designed to recognize and convert mathematical formulas, chemical equations, or other professional formulas. This model can be established based on deep learning techniques such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), or it can be a model obtained by fine-tuning large language models such as GPT-4V or GPT-4o. By using this preset formula recognition model to identify elements such as symbols, operators, variables, and constants in the formula, the model can understand the structure and semantics of the formula, and thus accurately convert the recognized formula into the formula format in the preset target format, such as LaTeX or MathML. This conversion is crucial for maintaining the accuracy of scientific and technical content in documents, especially in academic papers, technical reports, and educational materials. In addition, the model can also preserve the visual presentation features of the formula, such as font style, size, color, etc., to ensure that the converted document is visually consistent with the original PDF document. Through this preset formula processing algorithm, documents containing complex formulas can be effectively processed, ensuring accuracy and readability during the conversion process.

[0062] Step 30: Perform semantic conversion on the document content and layout features based on the pre-trained large model to obtain a number of target content blocks in the preset target format.

[0063] In this embodiment, the pre-trained large model is a machine learning model trained with a large amount of data or a processing model obtained by fine-tuning a large language model. It can understand and process natural language text, as well as recognize and convert various document formats. In this embodiment, the pre-trained large model is used to perform semantic conversion on the extracted document content and layout features to generate a number of target content blocks that conform to the preset target format. The preset target format can be formats such as Markdown format, HTML format, or Word document format. These document formats can all present text, pictures, tables, and formulas. For example, in the Markdown format, the title levels can be represented by symbols such as "#" and "##", pictures can be referenced by "!" and "[]", tables can be created using "|" and "-", and mathematical formulas can be represented by specific Latex syntax. And in Markdown files, partial HTML format can be used, and the style adjustment of elements such as titles, tables, and pictures can be carried out using the cascading style sheet configuration interfaces provided by relevant editor software. Finally, target content blocks consistent with the content and layout of the original PDF document can be obtained.

[0064] In this embodiment, the pre-trained large model can convert elements such as the extracted text, images, tables, and formulas into a preset target format by understanding the semantics of the document content, while maintaining the original structure and semantics. For example, for text content, the model can identify elements such as paragraphs, lists, and headings and format them according to the requirements of the target format; for images and tables, the model can generate corresponding reference codes or table codes; for formulas, the model can convert the recognized formula content into LaTeX or MathML codes. In this way, the pre-trained large model can not only handle the semantic conversion of the document content, but also handle the layout of the document, ensuring that the converted document is visually consistent with the original PDF document and meeting the user's requirements for document format conversion.

[0065] In this embodiment, the pre-trained large language model preferably adopts an Agent (intelligent agent) module integrated with the GPT-4o model, which interacts with the GPT-4o model through an API (application programming interface). The model receives the pre-processed image and text information and generates target content blocks in a preset target format according to predefined prompts. Among them, the predefined prompts are used to guide the model to generate target content blocks in a target format that meets the requirements according to the document content and layout features. The predefined prompts specifically include three main contents. The first is to prompt the GPT-4o model to be used for the parsing of standard text content, the second is to prompt the GPT-4o model to be used for the parsing of area content with annotations, and the third is to prompt the GPT-4o model to be used to define the role and output style of the model. For example, when the preset target format is Markdown format, the corresponding prompt can be: "Please convert the following content into Markdown format, ensuring that the heading levels, lists, image references, and table codes are correct. At the same time, for mathematical formulas, please use LaTeX syntax for conversion." Through such predefined prompts, the model can clearly know the specific requirements of the conversion task and thus generate accurate target content blocks. In addition, the use of the GPT-4o model can also achieve content planning and processing in a multi-language environment, ensuring accurate conversion of the content.

[0066] Step 40: Based on the position information of the original content blocks corresponding to each target content block, layout each target content block in the target format document to obtain the target format document.

[0067] In this embodiment, when laying out each target content block in the target format document according to the position information of the original content block, it is first necessary to determine the position of the target content block in the new document according to the position information of the content block in the original PDF document. This step is crucial because it ensures that the logical order and structural layout of the document content are consistent with the original document. By accurately positioning each content block, the readability of the document and the reading experience of the user can be guaranteed. During the layout process, the system will consider the relative positions and sizes of the original content blocks, as well as their layout relationships on the page. For example, if the title in the original document is located at the top of the page, then in the target format document, the corresponding title content block will also be placed at the top. Similarly, if the pictures and tables in the original document are arranged in a certain order, then in the target format document, these content blocks will also be laid out in the same order. After the layout is completed, the final target format document can be obtained. This document not only visually matches the original PDF document, but also matches the original document in terms of content logic and coherence, meeting the requirements for conversion from PDF to other formats, improving the integrity and readability of the converted document, and effectively ensuring the quality and secondary utilization efficiency of the target format document.

[0068] Optionally, refer to Figure 3 , and based on the position information of the original content blocks corresponding to each target content block, lay out each target content block in the target format document to obtain the target format document. The following steps can be performed:

[0069] Step 41: Calculate the layout positions of each target content block in the target format document based on the preset geometric algorithms provided by the Shape ly library and the position information of the corresponding original content blocks.

[0070] Step 42: Arrange each target content block at the calculated layout positions in the target format document.

[0071] Step 43: In the target format document, adjust the typesetting of the target content blocks according to the layout characteristics of the target content blocks to obtain the target format document.

[0072] In this alternative embodiment, the Shapely library is a Python library for processing and analyzing two-dimensional geometric objects. It provides rich geometric algorithms that can be used to calculate properties such as the positional relationships, boundaries, and areas of geometric objects. In this embodiment, the Shapely library is used to assist in determining the specific layout positions of target content blocks in a target format document. Through the geometric algorithms provided by the Shapely library, the positions of each target content block relative to the entire document layout can be accurately calculated, including their relative positions and dimensions. These calculation results will guide the placement of target content blocks in the target format document to ensure that their layout in the new document is consistent with that in the original PDF document.

[0073] In this alternative embodiment, after calculating and determining the layout positions of each target content block in the target format document, the target content blocks can be arranged in the target format document according to the calculated layout positions. After that, the typesetting of the target content blocks can be adjusted according to their layout characteristics to ensure the overall aesthetics and readability of the document. For example, for text content blocks, the font size, line spacing, and paragraph spacing can be adjusted to adapt to the layout design of the target format document. For images and tables, their positions and sizes on the page can be adjusted to ensure that they are visually coordinated with the surrounding text. In addition, for formula content blocks, their positions on the page can be adjusted to ensure that they are clearly readable and maintain an appropriate spacing from the relevant text content. Through these meticulous typesetting adjustments, the content and arrangement layout of the target format document are made consistent with the original PDF document, which can effectively improve the integrity and readability of the converted document and effectively guarantee the quality and secondary utilization efficiency of the target format document.

[0074] In this embodiment, during the process of processing the original PDF documents of each page, multi-threading can also be used for corresponding processing, so that the content parsing and conversion processing of multiple pages can be performed simultaneously, ensuring efficient operation even when faced with large-scale documents. At the same time, to further meet different format conversion requirements, flexible configuration options can be provided. The parsing and conversion rules can be customized through a configuration file, and the output format and content processing method suitable for specific needs can be selected. For example, users can specify the corresponding formula format, table format, or the embedding form of pictures in the preset target format. In addition, an automatic cleaning mechanism can be introduced into this format conversion method. This automatic cleaning mechanism can be executed automatically or with parameters set manually. For example, it can be automatically cleaned in response to the completion of parsing and conversion, or parameters such as frequency, time interval, or time point can be set for scheduled or frequency-based cleaning to ensure the cleanliness of the output and improve the resource utilization efficiency of the system.

[0075] Refer to Figure 4 , Figure 4It is a schematic structural diagram of an embodiment of the storage medium provided by this application.

[0076] The storage medium 50 stores program data 51. When the program data 51 is executed by a processor, it implements the Figures 1 to 3 format conversion method of the PDF document as described.

[0077] The program data 51 is stored in a storage medium 50 and includes several instructions for causing a network device (which can be a network device such as a router, a personal computer, a server, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of this application.

[0078] Optionally, the storage medium 50 can be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk, or an optical disc that can store the program data 51.

[0079] Refer to Figure 5 , Figure 5 It is a schematic structural diagram of an embodiment of the computer device provided by this application.

[0080] The computer device 60 includes a processor 62 and a memory 61 that are interconnected. The memory 61 stores a computer program. When the processor 62 executes the computer program, it implements the Figures 1 to 3 format conversion method of the PDF document as described. Among them, the memory 61 can include the storage medium 50 or can be other separately developed memories.

[0081] Different from the prior art, this application discloses a format conversion method, a storage medium, and a computer device for a PDF document. First, the original PDF document is parsed through a preset document parsing algorithm to extract several original content blocks. Then, according to the structural types of each original content block, such as text, image, table, formula, etc., its document content and layout features are extracted. Then, a pre-trained large model is used to perform semantic conversion on these document contents and layout features to obtain several target content blocks in a preset target format. Finally, according to the position information of the original content block corresponding to each target content block, these target content blocks are laid out in the target format document, thereby obtaining the target format document. It can efficiently and accurately parse the content of the PDF document and convert it into a corresponding target format document. The obtained target format document can retain the document content and layout features of the original PDF document, improving the integrity and readability of the converted document, and effectively ensuring the quality and secondary utilization efficiency of the target format document.

[0082] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the storage medium and the computer device, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0083] This application can be used in many general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0084] In several implementation manners provided in this application, it should be understood that the disclosed methods, storage media, and computer devices can be implemented in other ways. For example, the above-described implementation manner of the storage medium is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this implementation manner.

[0086] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0087] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of this application's specification and drawings, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.

Claims

1. A method for converting the format of a PDF document, characterized in that: include: Parsing the original PDF document based on a preset document parsing algorithm to extract several original content blocks of the original PDF document; Based on the structural type of each of the original content blocks, extracting the document content and layout features included in the corresponding original content block; Performing semantic conversion on the document content and the layout features based on the pre-trained large model to obtain a plurality of target content blocks in a preset target format; Based on the position information of the original content blocks corresponding to the target content blocks, the target content blocks are laid out in the target format document to obtain the target format document.

2. The method for converting the format of a PDF document according to claim 1, wherein: The preset document parsing algorithm includes a first parsing algorithm based on the PyMuPDF library and / or a second parsing algorithm based on the PDFPlumber library; The parsing of the original PDF document based on a preset document parsing algorithm to extract a plurality of original content blocks of the original PDF document includes: Based on the first parsing algorithm and / or the second parsing algorithm, the structural type, the position information, and the element content of a plurality of page elements in the original PDF document are determined to obtain the original content block corresponding to each of the page elements.

3. The method for converting the format of a PDF document according to claim 1, wherein: The structure type includes at least one of text, image, table and formula; The step of extracting the document content and layout features included in the corresponding original content block based on the structural type of each original content block includes: Determining that the structure type of the original content block is text, and extracting text structure information, text style information, and text content corresponding to the original content block based on a preset text processing algorithm; Determining that the structure type of the original content block is an image, extracting an image file and an image content description corresponding to the original content block based on a preset image processing algorithm, and generating an acquisition path corresponding to the image file; Determining that the structure type of the original content block is a table, and extracting table structure information and cell content corresponding to the original content block based on a preset table processing algorithm; It is determined that the structural type of the original content block is a formula, and the formula content corresponding to the original content block is extracted based on a preset formula processing algorithm.

4. The method for converting the format of a PDF document according to claim 3, wherein: The extracting of text structure information, text style information and text content corresponding to the original content block based on a preset text processing algorithm includes: Based on the structure parsing algorithm provided by the PDFPlumber library, the typesetting structure of the original content block is identified to obtain the text structure information; Performing text copying or text character recognition on the element content of the page element corresponding to the original content block to obtain the text content; Based on a preset semantic analysis model, the contextual relationship in the text content is parsed to determine the text style information.

5. The method for converting the format of a PDF document according to claim 3, characterized in that: The extracting the image file and the image content description corresponding to the original content block based on a preset image processing algorithm, and generating an acquisition path corresponding to the image file, includes: Copying or cropping the original content block to obtain the image file; When the corresponding original content block includes text information, copying or character recognition is performed on the text information to obtain the image content description; When the corresponding original content block does not include text information, selectively inputting the image file into a preset visual understanding model to generate the image content description; The image file is saved in a preset storage path, and a path index corresponding to the image file is created based on the preset storage path to generate the acquisition path.

6. The method for converting the format of a PDF document according to claim 3, characterized in that: The extracting the table structure information and cell content corresponding to the original content block based on a preset table processing algorithm includes: Identify the original content block based on the table parsing algorithm provided by the Camelot library to obtain the table structure information and the cell content; or Positioning and detecting the rectangular areas of the original content blocks are performed, and the rectangular areas are merged to determine the table structure information, and each of the rectangular areas is identified to obtain the corresponding cell content.

7. The method for converting the format of a PDF document according to claim 3, characterized in that: The extracting the formula content corresponding to the original content block based on a preset formula processing algorithm includes: Formula recognition is performed on the original content block based on a preset formula recognition model, and the corresponding formula recognition result is converted into a formula format in the preset target format.

8. The method for converting the format of a PDF document according to claim 1, characterized in that: The step of laying out each target content block in a target format document based on the position information of the original content block corresponding to each target content block to obtain the target format document comprises: Calculate the layout position of each target content block in the target format document based on a preset geometric algorithm provided by the Shapely library and the position information corresponding to the original content block; Arranging each of the target content blocks into the target format document according to the calculated layout position; In the target format document, the layout of the target content block is adjusted according to the layout features of the target content block to obtain the target format document.

9. A storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the steps of the method for format conversion of a PDF document as claimed in any one of claims 1 to 8 are implemented.

10. A computer device, characterized in that: The method comprises a processor and a memory connected to each other, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for format conversion of a PDF document as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • PDF text extraction method and system based on large language model

    CN120599643A

  • Multi-modal data preprocessing and fusion technology and system based on artificial intelligence

    CN120833614A

  • Document conversion method, electronic equipment, storage medium, chip and product

    CN121881984A