PDF (Portable Document Format) image-text separation method and device and medium

By reading the Script information of PDF and separating the operator information, and processing graphics and images in combination with Java's image processing package, the problem of damage to graphics and images during graphics and text separation in the prior art is solved, and efficient graphics and text separation and restoration effects are achieved.

CN120107417APending Publication Date: 2025-06-06SICHUAN LAN-BRIDGE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236766.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, when converting complex graphic and text mixed PDF files into word documents, it is difficult to effectively separate graphics, images and text, resulting in the impact of the integrity of graphics and images.

Method used

By reading the Script information of PDF, obtaining operator information, separating text description operators and graphic image description operators, and processing graphics and images using Java's image processing package, combining coordinate information of images and text, and finally outputting the processed image and text into a word document.

Benefits of technology

The separation of graphics, images and text is achieved, ensuring the integrity of graphics and images to the greatest extent and improving the restoration degree.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107417A_ABST
    Figure CN120107417A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of PDF (Portable Document Format) text processing, in particular relates to a PDF image-text separation method, a PDF image-text separation device and a medium, and aims to complete the processing of complex image-text mixed PDF, abandon the traditional OCR (Optical Character Recognition) technology, and perform processing through the characteristic that the characters, the images and the pictures of the PDF are layered so as to realize the purpose of image-text separation. By adopting the method and the device, the graph image and the character are processed separately in the processing process, so that the integrity of the graph and the image can be ensured to the greatest extent, and the reduction degree is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of PDF text processing, and in particular, relates to a PDF image-text separation method, device and medium. Background Art

[0002] In PDF files, there are often scenes of complex mixed graphics and text, such as PDF files converted from PPT-related files. When converting such PDF files to word documents, it is necessary to separate graphics, images and text, so as to restore the original graphics and images while also being able to edit the text; in the prior art, PDF document processing usually uses OCR technology for image recognition, and after recognizing the text, it actually cuts out the text on the original layer, but only compensates and repairs the background when backfilling, which will affect the original image and even damage the graphics and images. Summary of the invention

[0003] The purpose of the present invention is to provide a PDF image-text separation method to solve the technical problems existing in the prior art.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A PDF image-text separation method comprises the following steps: Step S1: Read the Script information of the PDF and obtain all operator information, including: text description operators and graphic image description operators; Step S2: First traversal: traverse all operators, process identifiers and record index information of all text description operators, recorded as TextIndex; Step S3: Second traversal: traverse all operators, process and identify and record all graphic image operators, and process them according to their path description using the Java image processing package, record the coordinate position information of all graphics and images, and at the same time record the index information of their operators, recorded as ImageIndex; Step S4: traverse ImageIndex, and complete the merging of ImageIndex and TextImageIndex according to the coordinate information of the graphics and image operators and other key information of PDF, which is recorded as MergeImageIndex; Step S5: The third traversal: traverse all operators, the operators in MergeImageIndex, process them using the GeneralPath class under the image processing package according to their parameter information and draw them onto BufferedImage, and output BufferedImage as a complete picture to the word document; Step S6: The operators in TextIndex are merged according to their parameter information, and then the text is output to the word document using the text box.

[0005] In one implementation, in step S4, it is determined whether there is an intersection between the coordinates of ImageIndex and TextImageIndex. If there is an intersection, the two can be merged.

[0006] In one embodiment, in step S6, operators located in TextIndex with the same Y coordinate are merged.

[0007] In order to achieve the purpose of the present invention, the present invention further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the PDF image and text separation method as described above.

[0008] In order to achieve the purpose of the present invention, the present invention also provides a device for separating PDF images and texts, characterized in that it includes: a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the device for separating PDF images and texts performs the PDF image and text separation method as described above.

[0009] Compared with the prior art, the present invention has the following beneficial effects: The present invention abandons the traditional OCR technology, but uses the characteristics of PDF itself that text, graphics and pictures are layered to process, thereby achieving the purpose of separating text and graphics. The present invention separates graphics and images from text during the processing process, which can ensure the integrity of graphics and images to the greatest extent, thereby ensuring the restoration degree. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a schematic diagram of the process of Example 1 of the present invention. DETAILED DESCRIPTION

[0011] In order to enable those skilled in the art to have a clearer understanding and knowledge of the present invention, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described below are only used to explain the present invention, and are convenient for understanding. The technical solutions provided by the present invention are not limited to the technical solutions provided by the following embodiments, and the technical solutions provided by the embodiments should not limit the protection scope of the present invention.

[0012] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present invention in a schematic manner. Therefore, the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the form, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0013] Example 1 like Figure 1 As shown, this embodiment provides a PDF image and text separation method, which aims to complete the processing of a PDF with complex mixed images and texts, and specifically includes the following steps: 1. Step S1: Read the PDF script information and obtain all operator information According to the official PDF documentation, operators are divided into text description operators (TOP), such as Tc, Tw, Tz, TL, Tf, Tr, Ts, Tj and other operations, and graphic image description operators (GOP), such as m, l, s, S, h, re, f, F, f*, c, v, y and other operators.

[0014] 2. Step S2: First traversal: traverse all operators, only process, identify and record the index information of all text description operators (TOP).

[0015] The processing in this step refers to the processing of text by org.apache.pdfbox.text.PDFTextStripper#processTextPosition in the PDFBOX open source framework. By rewriting the showFontGlyph method, the glyph information and its coordinate information currently described in the PDF Script can be obtained. Only the index information of all text description operators (TOP) is processed, identified and recorded, recorded as TextIndex. For some special character information (scientific notation characters, mathematical characters, abnormal width and height characters, icons, etc.), their index information is recorded separately, recorded as TextImageIndex.

[0016] 3. Step S3: Second traversal: traverse all operators, process and mark and record all graphic image operators (GOP).

[0017] The above processing refers to the org.apache.pdfbox.contentstream.PDFStreamEngine#processOperator(org.apache.pdfbox.contentstream.operator.Operator, java.util.List<org.apache.pdfbox.cos.COSBase> ) For the identification and processing of operators in the PDF Script, we can get which operators are related to graphics according to the operators (such as l, c, etc.), identify and record all graphics and image operators, and process them according to their path description using the Java image processing package. This processing refers to the java.awt.Graphics2D class in the Java library, in which the draw method can draw related graphics on the drawing board by specifying points and path information (java.awt.geom.GeneralPath), and java.awt.geom.GeneralPath is obtained based on the description information of the graphics in the above operator (including the starting point coordinates, end point coordinates, fill color, etc.); record the coordinate position information of all graphics and images, and at the same time record the index information of their operators, recorded as ImageIndex.

[0018] 4. Step S4: traverse ImageIndex, and complete the merging of ImageIndex and TextImageIndex according to the coordinate information of the graphics and image operators and other key information of PDF.

[0019] In this step, other key information that comes with PDF includes: fill color (Fill), brush path color (Stroke) and area declaration information (operators such as gs). The relevant operator information can be found in the official PDF documentation. This technology refers to (PDF Reference 1.7); the merging of ImageIndex and TextImageIndex is based on judgment, and the method is as follows: judge whether there is an intersection between the coordinates of ImageIndex and TextImageIndex. If there is an intersection, the two can be merged.

[0020] 5. Step S5: Third traversal: All operators are traversed for the third time. The operators in MergeImageIndex are processed using the GeneralPath class under the (java.awt) package according to their parameter information and drawn onto the BufferedImage.

[0021] In this step, the parameter information includes the coordinates of the point, fill color, brush path color, etc. (defined and introduced in the official PDF document). The GeneralPath class under the image processing package is used to process and draw on BufferedImage, and the BufferedImage is output to the word document as a complete picture. Among them, java.awt.geom.GeneralPath is the official SDK package of Java. By sequentially passing in the information of points and paths, java.awt.geom.GeneralPath can be constructed inside it, and then the current java.awt.geom.GeneralPath can be drawn on the drawing board through the draw method of BufferedImage. Output the BufferedImage as a complete picture to the word document.

[0022] 6. Step S6: The operators in TextIndex are merged according to their parameter information, and then the text is output to the word document using the text box.

[0023] In this step, the parameter information includes the coordinates of the glyph, the width and height of the glyph, and the coordinate transformation matrix related information. This information comes from the PDF Script file, and the meanings of these parameters are defined in the PDF official document; the above merging is based on judgment, and the method is as follows: the operators located in the TextIndex have the same Y coordinate and are merged.

[0024] Through the above method, the present invention applies the feature that the text, graphics and pictures of PDF are layered, and processes the graphic images and text separately during the processing, thereby achieving the purpose of separating the text and pictures and ensuring the integrity of the graphics and images to the greatest extent.

[0025] Example 2 This embodiment provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the PDF image and text separation method provided in Embodiment 1. A person skilled in the art can understand that all or part of the steps of implementing the method provided in Embodiment 1 can be completed by hardware related to the computer program, and the above-mentioned computer program can be stored in a computer-readable storage medium, and when the program is executed, the steps of the method provided in Embodiment 1 are executed; and the above-mentioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.

[0026] Example 3 This embodiment provides a device for separating PDF images and texts, including: a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the device for separating PDF images and texts performs the PDF image and text separation method provided in Example 1.

[0027] Specifically, the memory includes: ROM, RAM, disk, USB flash drive, memory card or CD and other media that can store program codes.

[0028] Preferably, the processor can be a general-purpose processor, including a central processing unit, a network processor, etc.; it can also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.

[0029] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A PDF image and text separation method, characterized in that: The following steps are involved: Step S1: Read the Script information of the PDF and obtain all operator information, including: text description operators and graphic image description operators; Step S2: First traversal: traverse all operators, process identifiers and record index information of all text description operators, recorded as TextIndex; Step S3: Second traversal: traverse all operators, process and identify and record all graphic image operators, and process them according to their path description using the Java image processing package, record the coordinate position information of all graphics and images, and at the same time record the index information of their operators, recorded as ImageIndex; Step S4: traverse ImageIndex, and complete the merging of ImageIndex and TextImageIndex according to the coordinate information of the graphics and image operators and other key information of PDF, which is recorded as MergeImageIndex; Step S5: The third traversal: traverse all operators, the operators in MergeImageIndex, process them using the GeneralPath class under the image processing package according to their parameter information and draw them onto BufferedImage, and output BufferedImage as a complete picture to the word document; Step S6: The operators in TextIndex are merged according to their parameter information, and then the text is output to the word document using the text box.

2. The PDF image-text separation method according to claim 1, characterized in that: In step S4, it is determined whether there is an intersection between the ImageIndex and TextImageIndex coordinates. If there is an intersection, the two can be merged.

3. The PDF image-text separation method according to claim 2, characterized in that: In step S6, operators located in TextIndex with the same Y coordinate are merged.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the PDF image and text separation method according to any one of claims 1 to 3.

5. A device for separating PDF images and texts, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the PDF image and text separation device performs the PDF image and text separation method according to any one of claims 1 to 3.