Differential visualization method and device after content preprocessing based on PDF document and storage medium
By extracting and mapping text and image content in PDF documents, and processing and difference marking, the problem of difficult to intuitively compare content differences after preprocessing of PDF documents is solved, and a more accurate corpus quality evaluation is achieved.
Patent Information
- Application Number
- CN202510473467.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
After preprocessing of PDF documents, it is difficult to intuitively compare the content differences before and after processing, which affects the evaluation of corpus quality.
By extracting the text content and picture content in the PDF document, the mapping file record location information is constructed, the text file is split into paragraph units, and it is processed with the image file input processing operator, the target document is integrated to generate, and the original and target documents are output in the same display interface, marking the difference part.
This realizes the visual comparison of the differences between PDF document content before and after preprocessing, and improves the accuracy of corpus quality evaluation.
Smart Images

Figure CN120012711A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of document processing, and in particular to a method, device and storage medium for visualizing differences after content preprocessing based on a PDF document. Background Art
[0002] During the model training process, the quality of the original data will directly affect the effect and stability of the model training. PDF documents, as a common format that carries a large amount of information, are often used as an important source of data collection and model training. However, PDF documents have a fixed layout and may contain unstructured data such as pictures and tables. These characteristics make it complicated to directly use PDF documents for model training. Therefore, in order to improve the effect and stability of model training, PDF documents usually need to be preprocessed, including content cleaning, deduplication, desensitization and other operations.
[0003] However, the current PDF document preprocessing technology extracts PDF documents into TXT format or image files and preprocesses the PDF documents. The preprocessed PDF documents lose the precise mapping of the location information with the original PDF documents. Therefore, it is difficult for users to intuitively compare the content differences before and after processing. In the corpus quality assessment stage, it is impossible to evaluate whether the preprocessing results meet the expected requirements. Summary of the invention
[0004] The main purpose of this application is to provide a method, device and storage medium for visualizing the differences after content preprocessing of a PDF document, aiming to solve the technical problem that it is difficult to intuitively compare the content differences before and after the PDF document is preprocessed.
[0005] To achieve the above-mentioned purpose, an embodiment of the present application provides a method for visualizing the difference of content preprocessing based on a PDF document, and the method for visualizing the difference of content preprocessing based on a PDF document comprises: Extracting text content and image content from a PDF document to obtain a text file and an image file, and recording original document information of the PDF document and location information of the text content and the image content in the text file and the image file; Based on the original document information and the location information, construct a mapping file of the PDF document, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content; Based on the mapping file, the text file is split into paragraphs to obtain paragraph units; Input the paragraph unit and the image file into corresponding processing operators for processing respectively, and integrate to obtain a target document according to the processing results of the processing operators; The PDF document and the target document are outputted in a same display interface, and the difference between the PDF document and the target document is marked in the display interface.
[0006] In one embodiment, the step of extracting text content and image content from a PDF document, obtaining a text file and an image file, and recording original document information of the PDF document and position information of the text content and the image content in the text file and the image file includes: Utilizing a PDF parsing library, extracting the text content and the image content of the PDF document; Determine original document information of the PDF document according to the document structure and semantic features of the PDF document, wherein the original document information includes chapters, paragraphs, page ranges, PDF paragraph coordinates, and image page numbers, image coordinates, and sizes; Matching the text content with the chapters and paragraphs in the PDF document, and obtaining a text line number, an offset, and a text index position corresponding to the text content in the text file according to the chapters and paragraphs; When saving the extracted picture content as the picture file, recording the file path of the picture file; The text line number, offset and text index position of the text content in the text file and the file path of the picture file are integrated to obtain the position information of the text content and the picture content in the text file and the picture file.
[0007] In one embodiment, the step of constructing a mapping file of the PDF document based on the original document information and the location information, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content comprises: Creating a mapping entry according to the original document information, wherein the mapping entry includes chapters, paragraphs, and pictures in the PDF document; Matching the chapters, paragraphs and pictures in the PDF document with the text content and the picture content; Determine, according to the original document information and the position information, a specific position of the mapping entry in the PDF document, and position information of the mapping entry in the text file and the image file; The mapping entries are integrated according to a preset data structure to obtain a mapping file of the PDF document.
[0008] In one embodiment, the step of splitting the text file into paragraphs based on the mapping file to obtain paragraph units includes: Based on the text line number, offset and text index position of the paragraph in the mapping file, the text content corresponding to each paragraph is extracted from the text file to obtain a plurality of paragraph units.
[0009] In one embodiment, the step of inputting the paragraph unit and the image file into corresponding processing operators for processing respectively, and integrating the target document according to the processing results of the processing operators comprises: Distributing the paragraph unit and the picture file to the corresponding processing operator for processing to obtain a processing result; The processing results are integrated according to the structure in the mapping file to obtain a target document.
[0010] In one embodiment, the step of outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface includes: Determining a page matching relationship between the PDF document and corresponding content in the target document according to the mapping relationship in the mapping file; Based on the page matching relationship, corresponding pages of the PDF document and the target document are displayed side by side in the same display interface; Different colors, marks or highlighting methods are used to mark the different parts between the PDF document and the target document.
[0011] In one embodiment, the step of outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface, further includes: When the selection instruction stays on a mark, a processing record of the text content or the picture content corresponding to the mark is displayed based on the process file.
[0012] In one embodiment, when the selection instruction stays on the mark, the step of displaying the processing record of the text content or the picture content corresponding to the mark based on the process file further includes: In the process of processing the paragraph unit and the image file, a process file is created to record the processing record of the processing operator processing the paragraph unit and the image file, wherein the processing record includes the processing start time, end time, processing type, processing status, location information, and the text content and the image content before and after processing; When the selection instruction stays on a mark, the processing record corresponding to the mark is retrieved and extracted from the process file, and displayed in the display interface in the form of a sidebar, a pop-up window or a floating frame.
[0013] An embodiment of the present application further provides a device for visualizing differences after content preprocessing based on a PDF document, the device for visualizing differences after content preprocessing based on a PDF document comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method for visualizing differences after content preprocessing based on a PDF document as described above.
[0014] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by a processor, the steps of the method for visualizing the differences after content preprocessing based on the PDF document as described above are implemented.
[0015] The embodiment of the present application discloses a method for visualizing differences after content preprocessing based on a PDF document. The method comprises extracting text content and image content from a PDF document to obtain a text file and an image file, and recording original document information of the PDF document and position information of the text content and the image content in the text file and the image file; constructing a mapping file of the PDF document based on the original document information and the position information, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content; splitting the text file into paragraphs based on the mapping file to obtain paragraph units; inputting the paragraph units and the image file into corresponding processing operators for processing, and integrating them to obtain a target document according to the processing results of the processing operators; outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface. This application establishes a positional mapping relationship between text content and image content and the original PDF document, and synchronously saves the processing records when processing the extracted content using the processing operator. This enables the document content before and after processing to be synchronously displayed according to the mapping relationship when performing a visual comparison display of the differences, so that users can intuitively compare the content differences before and after processing. In the corpus quality assessment stage, it can effectively evaluate whether the preprocessing results meet the expected effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of a first embodiment of a method for visualizing differences after content preprocessing based on a PDF document according to an embodiment of the present application; Figure 2 This is a schematic diagram of the mapping relationship of the first embodiment of the method for visualizing the difference after content preprocessing based on PDF documents involved in the embodiment of the present application; Figure 3This is a flow chart of a second embodiment of a method for visualizing differences after content preprocessing based on a PDF document according to an embodiment of the present application; Figure 4 This is a flow chart of a third embodiment of a method for visualizing differences after content preprocessing based on a PDF document according to an embodiment of the present application; Figure 5 This is a flow chart of a fourth embodiment of a method for visualizing differences after content preprocessing based on a PDF document according to an embodiment of the present application; Figure 6 This is a schematic diagram of the structure of a device for visualizing differences after content preprocessing based on a PDF document according to an embodiment of the present application.
[0017] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0018] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] During the model training process, the quality of the original data will directly affect the effect and stability of the model training. PDF documents, as a common format that carries a large amount of information, are often used as an important source of data collection and model training. However, PDF documents have a fixed layout and may contain unstructured data such as pictures and tables. These characteristics make it complicated to directly use PDF documents for model training. Therefore, in order to improve the effect and stability of model training, PDF documents usually need to be preprocessed, including content cleaning, deduplication, desensitization and other operations.
[0020] However, the current PDF document preprocessing technology extracts PDF documents into TXT format or image files and preprocesses the PDF documents. The preprocessed PDF documents lose the precise mapping of the location information with the original PDF documents. Therefore, it is difficult for users to intuitively compare the content differences before and after processing, and thus it is impossible to evaluate whether the preprocessing results meet the expected requirements.
[0021] In order to solve the above-mentioned defects existing in the related art, an embodiment of the present application proposes a method for visualizing differences after content preprocessing based on a PDF document. The method obtains a text file and a picture file by extracting text content and image content in a PDF document, and records the original document information of the PDF document, as well as the position information of the text content and the picture content in the text file and the picture file; based on the original document information and the position information, a mapping file of the PDF document is constructed, and the mapping file includes a mapping relationship between the PDF document and the text content and the picture content; based on the mapping file, the text file is split into paragraphs to obtain paragraph units; the paragraph units and the picture file are respectively input into corresponding processing operators for processing, and a target document is obtained by integration according to the processing results of the processing operators; the PDF document and the target document are output in the same display interface, and the difference parts between the PDF document and the target document are marked in the display interface.
[0022] This application establishes a positional mapping relationship between text content and image content and the original PDF document, and synchronously saves the processing records when processing the extracted content using the processing operator. This enables the document content before and after processing to be synchronously displayed according to the mapping relationship when performing a visual comparison display of the differences, so that users can intuitively compare the content differences before and after processing. In the corpus quality assessment stage, it can effectively evaluate whether the preprocessing results meet the expected effects.
[0023] It should be noted that the execution subject of this embodiment can be a content preprocessing difference visualization system based on PDF documents, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a content preprocessing difference visualization device based on PDF documents that can realize the above functions. The following takes the content preprocessing difference visualization system based on PDF documents as an example to illustrate this embodiment and the following embodiments.
[0024] The first embodiment of the present application proposes a method for visualizing the difference of content preprocessing based on PDF documents. Figure 1 The method comprises steps S10 to S50: Step S10: extracting text content and image content from the PDF document to obtain a text file and an image file, and recording original document information of the PDF document and location information of the text content and the image content in the text file and the image file.
[0025] In this embodiment, the PDF parsing library is used to extract the text content and image content in the PDF document, and the text content and image content are saved as text files and image files. Among them, the text content is saved in plain text format, and the image content is saved as a corresponding image file. In the process of extracting the PDF document, according to the document structure of the PDF document, the original document information of each part of the content in the PDF document is synchronously recorded in the PDF document. The original document information includes chapters, paragraphs, page ranges (starting page numbers and ending page numbers), PDF paragraph coordinates, and image page numbers, image coordinates and sizes.
[0026] In addition, the location information of the text content and the image content in the extracted text file and the image file will also be recorded. Specifically, when the text content and the image content are extracted using the PDF parsing library and the text file and the image file are obtained, the text line number, offset, and text index position of the text content corresponding to the document structure such as chapters and paragraphs in the original document information will be recorded in the text file, and after the image content in the original document information is extracted, the file path where the corresponding image file is saved will be recorded.
[0027] It should be noted that when the text content is extracted and saved in a plain text format, each text paragraph or text line will have a corresponding text line number in the generated text file, and the text line number is used to indicate the position of the text content in the text file.
[0028] Step S20: constructing a mapping file of the PDF document based on the original document information and the location information, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content.
[0029] In this embodiment, mapping entries are created according to the document structure in the original document information, such as chapters, paragraphs or pictures, etc. The mapping entries will serve as the basic unit of the mapping file and are used to establish the mapping relationship between the PDF document and the extracted text files and picture files.
[0030] The document structures such as chapters, paragraphs and images in the PDF document are matched one by one with the text content and image files in the extracted text file. For example, a chapter in the PDF document exists in a certain page range in the original PDF document, and the text content of the chapter obtained through extraction corresponds to a specific line number range in the text file, and a picture in the chapter corresponds to an image file with a specific file path. In this way, it is ensured that the mapping file can accurately reflect the location of the content in the PDF document in the extracted file.
[0031] In the mapping file, the original document information of each mapping entry in the PDF document is recorded, including chapter number, paragraph number, page range, PDF paragraph coordinates, image page number, image coordinates and size, etc. At the same time, the location information of these contents in the extracted text file and image file is also recorded, such as text line number, offset, text index position and image file path, etc. Finally, all mapping entries are integrated according to the preset data structure to form a complete mapping file.
[0032] It should be noted that the data structure can be a list, array, JSON, XML or a custom data structure, which can be quickly queried and accessed so that the content in the PDF document and the content of the corresponding text file and image file can be quickly located in the subsequent steps. The mapping file at least includes the correspondence between the original position information of the text content in the PDF document and the text line number, offset and text index position of the text content in the text file, and the correspondence between the original position information of the image content in the PDF document and the file path and file name of the extracted image file.
[0033] For example, a json file named metadata-mapping_file number_version is constructed, for example, named metadata-mapping.json, which includes the mapping relationship between the PDF document and the extracted text files and image files. Figure 2 As shown: First, determine the basic information of the PDF document and the text file that stores the extracted content, such as the PDF file number, the temporary text file number, the total number of pages and chapters of the PDF document, etc., to distinguish and manage different PDF documents, as well as the mapping relationship between the text file and image file pre-extracted from the PDF document.
[0034] Next, based on the document structure of the PDF document, the basic information of each document structure and the text content information corresponding to the extracted text files and image files are recorded. For example, the document structure in a PDF document includes chapters, paragraphs, etc. Taking the chapter as an example, when recording the chapter information of each chapter, the chapter information further includes the chapter number, PDF chapter information and text chapter information, paragraphs and pictures in the chapter. Among them, the chapter number is used to identify each chapter in the PDF document; the PDF chapter information is the original position of the chapter in the PDF document, including the starting page number and the ending page number of the chapter; the text chapter information indicates the position information of the chapter in the extracted text file, including the starting line number, the ending line number, the text index starting position and the text index ending position of the chapter, etc.
[0035] Furthermore, there are multiple paragraphs in a chapter, so the paragraph numbers, PDF paragraphs, and text paragraphs corresponding to the paragraphs in the chapter are also recorded. The paragraph number is used to identify each paragraph in the chapter; the PDF paragraph indicates the original position of the paragraph in the PDF document, including the page position and PDF coordinates of the paragraph in the PDF document; the text paragraph indicates the position information of the paragraph in the extracted text file, including the starting line number, ending line, offset, text index starting position, and text index ending position of the paragraph in the text file.
[0036] In addition, if there are pictures in the chapter, the PDF picture information and the file path of the extracted picture file will be recorded. The PDF picture information indicates the original position of the picture in the PDF document, including the page number, picture coordinates and size, so that the processing of the picture can be accurately displayed when visualizing the difference.
[0037] Finally, the above information is integrated into a structured mapping file, which contains the mapping relationship between the original content of the PDF document and the extracted text file and image file, so that in the subsequent difference analysis, the original PDF document and the processed target document can be quickly located according to the mapping file.
[0038] For example, in a certain PDF document, the PDF file number is A, the extracted temporary text file number is A1, the total number of pages of the PDF document is 100, and the total number of chapters is 3.
[0039] For a chapter numbered 1, the following information is recorded: PDF chapter information includes the starting page number (1) and ending page number (30) of the chapter in the PDF document; text chapter information includes the starting line (1), ending line (500), text index starting position (0), and text index ending position (10000) of the chapter in the text file.
[0040] For a paragraph with paragraph number 1, the following information is recorded: the PDF paragraph includes the page number of the paragraph in the PDF document (1), and the PDF coordinates of the paragraph in the PDF document (PDF coordinates Y u :700, PDF coordinate Y d :680), a text paragraph includes the paragraph's start line (1), end line (3), text index start position (0), and text index end position (150) in the text file.
[0041] In addition, for the images in the chapter, the following information is recorded: PDF image information includes the page number of the image in the PDF document (1), image coordinates (coordinate x: 100, coordinate y: 650) and image size (width: 100, height: 100); image file information includes the stored image file path (image file path: / A / dataset / images / 123.png).
[0042] It should be noted that the image coordinates only record the coordinate value of a certain corner of the image. This coordinate value is used as a reference point of the PDF page, and combined with the image size, the specific position of the image can be determined. The reference point can be selected as the upper left corner, lower left corner, upper right corner, lower right corner or center point of the image, etc., and this embodiment does not impose any restrictions on this. Different reference point selections only affect the subsequent calculation method of the image position, but no matter which point is selected as a reference, combined with the corresponding image size information, the position range of the image in the PDF page can be uniquely determined.
[0043] Step S30: Based on the mapping file, the text file is split into paragraphs to obtain paragraph units.
[0044] In this embodiment, based on the text line number, offset and text index position of the paragraph in the mapping file, the text content corresponding to each paragraph is extracted from the text file to obtain multiple paragraph units. Specifically, according to the information about the text content part in the mapping file, the start line, end line, offset relative to the start line and end line of each text paragraph and the text index position are found, and then the corresponding text content is intercepted from the text file and divided into different paragraph units.
[0045] The line number refers to the number of each line in the text file. The offset is the starting position of the actual content of the text paragraph starting from the specified line. Through the offset, you can accurately locate the character in the line from which the paragraph actually starts, and the character in the line at which the paragraph ends.
[0046] In addition, the text index position is a description of the paragraph position from a global perspective. It regards the text file as a continuous character sequence and then numbers each character from the beginning of the text file. The text index position includes the text index start position and the text index end position. The text index start position in a text paragraph is the number of the starting character of the paragraph, and the text index end position is the number of the ending character of the paragraph.
[0047] In a PDF document, the text may have complex layout, a paragraph may not start at the beginning of the line, or when extracted as a text file, it may contain additional formatting information, such as paragraph indentation, special symbols, etc. Therefore, it is impossible to accurately determine the start and end positions of a paragraph by relying solely on the line number. The offset can accurately determine the number of characters in the line where the paragraph starts and ends, or combine the text index position to further determine the paragraph position, so as to accurately intercept the required paragraph content.
[0048] For example, a paragraph starts at line number 10, with an offset of 5, and ends at line number 20, with an offset of 8. The actual content of the paragraph starts at the 5th character of line 10 and ends at the 8th character of line 20.
[0049] For example, the starting line number of a paragraph is 1, the ending line number is 3, the text index starting position is 0, and the text index ending position is 150. In this case, since the text index starting position is 0, it means that the actual content of the paragraph starts from the beginning of the first line and extends to the position of the 150th character.
[0050] It should be noted that the text file is regarded as a continuous character sequence here, and each character is numbered from the beginning of the file. Therefore, the end position of the text index is determined based on the global character sequence of the text file, not limited to a certain line, but counted cumulatively from the beginning of the file, not from the third line.
[0051] It should be noted that, since paragraphs are the smallest units with relatively complete semantics and logical structures in a document, segmenting a document by paragraphs can better ensure the integrity and coherence of information. Therefore, in this embodiment, paragraphs are preferentially used as the basic units of segmentation, and the entire PDF document is split into multiple relatively independent and complete parts according to the starting line number, ending line number, offset, text index position, and related semantics and logical boundaries of the paragraphs. However, this embodiment does not limit paragraphs to be the only basic unit of segmentation, and chapters, page numbers, etc. can also be used as the basis for segmentation.
[0052] Step S40: input the paragraph unit and the image file into corresponding processing operators for processing respectively, and integrate them to obtain a target document according to the processing results of the processing operators.
[0053] In an optional embodiment, step S40 further includes steps S41-S42: Step S41: Distribute the paragraph unit and the picture file to the corresponding processing operator for processing to obtain a processing result.
[0054] Step S42: Integrate the processing results according to the structure in the mapping file to obtain a target document.
[0055] In this embodiment, the processing operator includes multiple text processing functions, which are intended to operate on text content or image content to achieve the purpose of modifying or converting documents. These processing operators can be algorithms such as text cleaning, deduplication, desensitization, format conversion, content enhancement, etc., which process text content according to predetermined rules, while the processing of image content may include image compression, resolution adjustment, format conversion, etc.
[0056] For example, for each paragraph unit, the processing operator may first perform text cleaning to remove extra spaces, line breaks or special characters, etc.; then perform text formatting to ensure that the font, size, color, etc. of the text meet the requirements; for image files, the processing operator may also perform format conversion, compress the image to reduce the file size, etc., or perform image enhancement to improve the image quality.
[0057] Therefore, when the PDF document is extracted to obtain text files and image files, and the text files are further split into multiple paragraph units, the paragraph units and image files can be distributed to the corresponding processing operators. These processing operators will perform parallel processing on the input paragraph units and image files according to preset functions to obtain processing results.
[0058] After completing the processing of paragraph units and image files, the processing results are integrated into the target document according to the document structure and logical relationship of the original PDF document, ensuring that the processed text content and image content in the target document still maintain the original logical order and layout.
[0059] Exemplarily, according to the mapping relationship in the mapping file, the processed paragraph units are arranged in the original order of chapters and paragraphs, and the processed images are accurately placed in the corresponding positions of the target document.
[0060] It should be noted that parallel processing can take advantage of multi-core processors or distributed computing environments to improve the processing speed of PDF documents when processing large amounts of data. Parallel processing can be achieved using a parallel computing framework, such as using multi-threaded or multi-process programming techniques to assign each processing unit as an independent task to different threads or processes for processing. Alternatively, Java's ExecutorService, Python's multiprocessing library, etc. can be used in a multi-threaded environment.
[0061] During the processing, each processing operator call will follow the corresponding relationship defined in the mapping file to ensure that the processed content can accurately correspond to the location in the original PDF document. After the processing is completed, the integrated target document also maintains the same structure as the original PDF document, but the content has been modified and optimized accordingly according to the definition of the processing operator.
[0062] Exemplarily, the content of metadata-mapping.json is parsed, and according to the start and end line numbers of the text paragraph, as well as the offset and text index position, the text paragraph is used as the basic processing unit (i.e., paragraph unit), the paragraph unit and the image file are loaded into the operator center module, the corresponding processing operator is called for processing, and then the processing results are integrated according to the document structure of the original PDF document to obtain the target document.
[0063] Step S50: outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface.
[0064] In this embodiment, the PDF document and the target document are loaded in the same display page, and the page number position where the PDF document is triggered is activated according to the mapping file, so that the PDF document and the target document are displayed synchronously. Then, using a visualization tool, such as a difference comparison software or a custom-developed comparison interface, the difference between the PDF document and the target document is intuitively displayed by highlighting, red frame marking or other special markings, sidebar annotation, and parallel display, so that the user can intuitively compare the context and the processed content. For example, the PDF document and the target document are visually compared and displayed by parallel display. The original PDF document is displayed on the left, and the processed content of the text file is displayed on the right.
[0065] In addition, the top of the display page can also be shrunk to display or paged to display the processing actions of the difference part.
[0066] Please refer to Figure 3 In the second embodiment of the present application, the method for visualizing the difference after content preprocessing based on the PDF document is as follows: Step S10 includes steps S11 to S15: Step S11: using a PDF parsing library to extract the text content and the image content of the PDF document.
[0067] In this embodiment, a PDF parsing library, such as PyPDF2, PDFMiner or Poppler, is used to extract the text content and image content in the PDF document. These libraries can read the structure of the PDF file, parse the text and image elements in the PDF document, and convert them into an operable data format. The text content is usually extracted as a plain text format, while the image content is saved as an image file, such as a JPEG or PNG format.
[0068] Step S12: determining original document information of the PDF document according to the document structure and semantic features of the PDF document, wherein the original document information includes chapters, paragraphs, page ranges, PDF paragraph coordinates, and image page numbers, image coordinates, and sizes.
[0069] In this embodiment, the document structure includes elements such as chapters and paragraphs, which are based on the layout and content logic of the PDF document, for example, different font sizes, indents, line breaks, etc. are used to distinguish chapters and paragraphs. Semantic features refer to the information logic and hierarchy conveyed by the PDF document, for example, semantic elements such as titles and subtitles are used to assist in determining the document structure of the PDF document.
[0070] Specifically, in the process of parsing the PDF document using the PDF parsing library, the start and end positions of chapters and paragraphs are determined according to the typesetting format of Odifen, for example, by observing different fonts, font sizes, indentation styles, etc., and then the chapters and paragraphs are determined. At the same time, according to the page number marks and page layouts of the chapters and paragraphs, the page number ranges and PDF paragraph coordinates of the chapters and paragraphs are obtained.
[0071] In addition, for the images in the PDF document, the page number where the image is located and the coordinates and size on the page are determined according to the position and layout of the image on the page.
[0072] Step S13: Match the text content with the chapters and paragraphs in the PDF document, and obtain the text line number, offset and text index position corresponding to the text content in the text file according to the chapters and paragraphs.
[0073] After the text content is extracted from the PDF document, since the layout and structure of the text content in the text file may change during the extraction process, it is necessary to match the text content with the chapters and paragraphs of the original PDF document.
[0074] According to the chapter and paragraph information in the original PDF document, the extracted text content is re-associated according to the original structure of the document. By analyzing the text content line by line and combining the judgment of the starting and ending positions of chapters and paragraphs, the text content is divided into corresponding chapters and paragraphs.
[0075] At the same time, when storing the text content as a text file, the text content is numbered and its position is recorded, thereby obtaining line numbers of the text. The offset refers to the position information of the text content within a line relative to the starting position of that line, and the text index position is the starting character sequence number and the ending character sequence number corresponding to this part calculated from the beginning of the text file.
[0076] In an alternative embodiment, a line counter is initialized and set to 1. When storing the extracted text content into a text file, the text content is processed line by line. For each line of text content, it is stored in the text file, and the value of the current line counter is used as the line number of that line.
[0077] Whenever a line of text content is processed, the line counter is incremented by 1. By processing each line of text content in sequence, the line numbers corresponding to each line of text content in the text file can be obtained. At the same time, when storing the text lines, the line numbers are associated with the text of that line. For example, a dictionary or a list can be used to store the association between the text and the line numbers for convenient subsequent lookup.
[0078] Furthermore, when processing each line of text content, for the text content within each line, the offset is calculated based on the starting position of the text content of that line. The offset of the first character of each line is initialized to 0. When traversing each character of that line, the offset increases with the position of the character. For example, for a line "The weather today is very nice", when processing the character "天" (which might be a Chinese character here, assuming it's the 4th character in the line), its offset is 4 because "天" is the 4th character in this line. The offset is updated as the characters are processed, and when storing the text content, the offset is associated with the chapter or paragraph. For example, using a nested data structure, when storing the text line information of a chapter or paragraph, the offsets of the text content at the start line and the end line of the chapter or paragraph are stored simultaneously.
[0079] In addition, the text index position is the character sequence number calculated from the beginning of the entire text file. Initially, the starting value of the text index position is set to 0. When storing each line of text content into the text file, for each character stored, the index position of that character is set to the current text index position and then stored. At the same time, the text index position is updated by increasing it by the length of that character.
[0080] Finally, according to the starting position information and the ending position information of the PDF chapters and PDF paragraphs in the PDF document, such as the starting page number and the ending page number of the PDF chapter, and the starting and ending positions of the PDF paragraph, when extracting the text content of the PDF document, it is determined which text chapter and text paragraph the currently extracted text line belongs to in the text file, thereby determining the range information of the text chapter and text paragraph.
[0081] Then, based on the range information of the text chapters and text paragraphs, as well as the text line number, offset and text index position of the text content, the text content is accurately divided into corresponding text chapters and text paragraphs, so that the position information such as the starting line number, ending line number, offset and text index position corresponding to each text chapter and text paragraph in the text file can be obtained.
[0082] Step S14: when saving the extracted picture content as the picture file, recording the file path of the picture file.
[0083] In this embodiment, after extracting the image content from the PDF document, it needs to be saved as an independent image file. During the saving process, the image is stored in a specific location, which is represented by a file path. The file path of the image file is recorded so that these images can be quickly found and operated in the subsequent operator processing process.
[0084] Specifically, the extracted image content is stored in a corresponding file format (such as JPEG, PNG, etc.), and a unique file path is generated for each image file. The file path can be based on a custom directory structure and file naming rules.
[0085] Step S15: Integrate the text line number, offset and text index position of the text content in the text file and the file path of the picture file to obtain the position information of the text content and the picture content in the text file and the picture file.
[0086] In this embodiment, the text line number, offset and text index position of the text content in the text file, and the file path of the picture file and other information are integrated to obtain the position information of the text content and the picture content in the text file and the picture file.
[0087] Please refer to Figure 4 In the third embodiment of the present application, the method for visualizing the difference after content preprocessing based on the PDF document is proposed. Step S20 includes steps S21 to S24: Step S21: creating mapping entries according to the original document information, wherein the mapping entries include chapters, paragraphs and pictures in the PDF document.
[0088] In this embodiment, the original document information includes the document structure of the PDF document, such as chapters, paragraphs, page ranges, image information, etc. The mapping entry is created to manage the document structure in the PDF document. The mapping entry can be regarded as a structured description of elements such as chapters, paragraphs, and images in the PDF document, and the mapping entry will serve as a basic element for constructing a mapping file.
[0089] When creating a mapping entry, the PDF document will be traversed to determine each chapter, paragraph, and image in the PDF document. For example, for a chapter, the chapter number, chapter name, start page number, and end page number will be extracted; for a paragraph, the paragraph number, chapter, page number, and paragraph coordinate range will be extracted; for an image, the image number, page number, image coordinates, and size will be extracted. These chapter, paragraph, and image information will serve as the content of the mapping entry, aiming to record and store the elements in the PDF document in a structured manner, providing a basis for mapping and positioning operations with text files in subsequent steps.
[0090] Step S22: Matching the chapters, paragraphs and pictures in the PDF document with the text content and the picture content.
[0091] In this embodiment, since the positions and structural relationships of the text content and picture content in the chapters, paragraphs and pictures in the PDF document may be different from those in the PDF document after extraction, it is necessary to correspond the chapters, paragraphs and pictures in the PDF document to the text content and picture content in the text file and picture file.
[0092] Specifically, it is necessary to match chapters and paragraphs with the extracted text content to ensure that each PDF chapter and PDF paragraph in the PDF document can find the corresponding text chapter and text paragraph in the text file. This correspondence relationship requires that the chapters and paragraphs are consistent in text content, and that the positions of the chapters and paragraphs in the PDF document and the text document correspond.
[0093] In addition, it is necessary to match the images with the extracted image content to ensure that the images in the PDF document can find the corresponding images stored in the image file.
[0094] Step S23: determining the specific position of the mapping entry in the PDF document and the position information of the mapping entry in the text file and the image file according to the original document information and the position information.
[0095] Based on the original document information and the position information, the specific position of each PDF chapter and PDF paragraph in the PDF document, as well as the position information of each text chapter and text paragraph in the text file can be determined.
[0096] Specifically, in the process of extracting the text content and image content of a PDF document, since the position information of each chapter, paragraph and image of the PDF document is determined in advance, it is possible to determine, based on the original document information of the PDF document, which text chapter and text paragraph the currently extracted text line belongs to in the text file, and which PDF document the currently extracted image corresponds to in the image file.
[0097] It should be noted that in the process of converting the content of a PDF document into a text file, although the presentation of the PDF document changes, the actual document structure remains the same. Therefore, you can refer to the framework of the PDF document such as chapters and paragraphs, combined with the progress of text content extraction and the characteristics of the text line, to determine the range to which the text line belongs in the text file.
[0098] For example, a chapter starts from page 1 and ends at page 30, and the text currently being extracted is from page 15. It can be determined that the text content currently extracted also belongs to the chapter in the text file.
[0099] Step S24: Integrate the mapping entries according to a preset data structure to obtain a mapping file of the PDF document.
[0100] In this embodiment, the preset data structure includes a list, an array, JSON, XML or a custom data structure, so that the content in the PDF document and the content of the corresponding text file and image file can be quickly located in subsequent steps.
[0101] Taking JSON as an example, the mapping entries are stored hierarchically. In the JSON structure, chapters are the top-level elements, which contain chapter information related to the chapters. The chapters contain multiple paragraphs and pictures, and each paragraph and picture further contains paragraph information and picture information. These chapter information, paragraph information and picture information also include corresponding location information.
[0102] In the process of integrating mapping entries, the mapping entries and their corresponding location information are organized according to a specific structure. For example, the page range, starting position and ending position of each chapter in the chapter information, as well as the paragraph information contained in each chapter, are first determined. Then, corresponding attributes are added to each chapter, such as chapter number and chapter name, and the paragraph information and image information in the chapter are stored as sub-elements under the chapter information.
[0103] In addition, chapter information, paragraph information and image information include two parts respectively: one is the information about the PDF document, and the other is the information about the text file and image file extracted from the PDF document. Specifically, chapter information includes PDF chapter information and text chapter information, paragraph information includes PDF paragraph information and text paragraph information, and image information is divided into PDF image information and image file information.
[0104] Please refer to Figure 5 In the fourth embodiment of the present application, the method for visualizing the difference after content preprocessing based on the PDF document, step S50 includes steps S51 to S53: Step S51: determining a page matching relationship between the PDF document and corresponding content in the target document according to the mapping relationship in the mapping file.
[0105] In this embodiment, the mapping file provides the correspondence information between the original PDF document and the target document, including the location information of elements such as chapters, paragraphs and pictures. In the process of operator processing of text content and picture content, the content may be moved, scaled, rotated and so on. Therefore, when displaying the comparison content of the PDF document and the target document before and after processing, the page matching relationship between the PDF document and the corresponding content in the target document can be determined through the mapping relationship in the mapping file. In this way, the displayed pages of the PDF document and the target document can belong to the same text content.
[0106] Step S52: Based on the page matching relationship, corresponding pages of the PDF document and the target document are displayed side by side in the same display interface.
[0107] In this embodiment, the side-by-side display means that the document contents of the two versions are displayed on the interface at the same time according to their positional relationship, so that the user can compare and analyze intuitively. The PDF document and the target document are loaded through a PDF loading tool, such as PyMuPDF. After loading, the two document contents are displayed side by side on the same interface using a visualization tool. The side-by-side display method allows the user to view the original state and the processed state of the PDF document at the same time, which is convenient for intuitive comparison and analysis.
[0108] Commonly used visualization tools include Diff2Html, which can convert the difference results into HTML format and support side-by-side display.
[0109] For example, a dual-view or multi-view layout is used to display the original PDF document and the target document side by side or overlapping. In this way, the user can clearly see the position difference of each content in the two versions, as well as their relative relationship in the document structure, which improves the readability of the difference visualization.
[0110] Step S53: using different colors, marks or highlighting methods to mark the difference between the PDF document and the target document.
[0111] On the basis of the side-by-side display, the text content and image content before and after processing are further highlighted in the display interface. Highlighting is an effective visual prompt method that can highlight the differences in the document content, allowing users to quickly identify which content has changed.
[0112] For the difference marks of text content, you can use the highlighting function provided by tools such as Diff2Html. For example, when the text content is modified or deleted, you can use different colors or backgrounds to highlight the differences.
[0113] For marking the difference of the image content, image comparison and difference detection technology can be used, such as using OpenCV and Matplotlib libraries to identify the differences between images and mark them on the images.
[0114] Through highlighting, users can more intuitively understand the specific changes in document content, improving the difference visualization effect and user experience.
[0115] In the display interface, the position of each text content and image content before and after processing is highlighted by highlighting or special marking. Highlighting is a common visual prompt method, which can be achieved by changing the background color, font style, border style, etc. of the text or image. Special marking can be achieved by using graphic elements such as icons, labels, lines, etc. to mark the position of the content.
[0116] For example, different highlight colors can be set for the positions before and after processing, so that users can quickly identify changes in content. Alternatively, an arrow or icon can be added next to the content before and after processing to point to the positional relationship before and after processing.
[0117] In the fifth embodiment of the present application, the method for visualizing differences after content preprocessing based on PDF documents, step S50 further includes step S54: Step S54: When the selection instruction stays on a mark, a processing record of the text content or the picture content corresponding to the mark is displayed based on the process file.
[0118] In the display interface, when the user places the mouse or other selection command on a highlighted mark, the difference information and processing record of the text content or image content corresponding to the current mark are displayed based on the process file generated by the processing operator when processing the text content and image content.
[0119] In this way, users can not only see the differences in document content, but also further understand how these differences are caused, including the specific processing operations and execution processes.
[0120] In an optional implementation, step S54 further includes steps S541-S542: Step S541: In the process of processing the paragraph unit and the image file, a process file is created to record the processing record of the processing operator processing the paragraph unit and the image file, and the processing record includes the start time, end time, processing type, processing status, location information, and the text content and the image content before and after processing.
[0121] Step S542: When the selection instruction stays on a mark, the processing record corresponding to the mark is retrieved and extracted from the process file, and displayed in the display interface in the form of a sidebar, pop-up window or floating box.
[0122] In this embodiment, in order to better manage the processing of PDF documents, it is necessary to record the execution of each processing operator in detail. Creating a process file is an effective process management method that can help users track the processing progress and results of each processing unit.
[0123] By recording the start and end time of the processing, you can understand the execution time of each processing operator, thereby evaluating the processing efficiency and performance. The processing type and processing status respectively indicate the specific operation performed by the processing operator and the success or failure of the execution process. The location information records the specific coordinates of the text content and image content in the PDF document, ensuring that the processing results can be accurately displayed in the subsequent difference visualization process. Recording the text content and image content before and after processing can be used to compare and analyze the document content before and after processing when displaying on the same display page.
[0124] When the user places the selection instruction on these marks, the processing record corresponding to the mark is retrieved and extracted from the process file. The extracted processing record is then displayed in the display interface in the form of a pop-up window or a floating box, so as to intuitively present the processing record to the user without affecting the user's browsing of the PDF document and the target document.
[0125] An embodiment of the present application provides a device for visualizing differences after content preprocessing based on a PDF document. The device for visualizing differences after content preprocessing based on a PDF document includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for visualizing differences after content preprocessing based on a PDF document in the first embodiment.
[0126] Reference below Figure 6 , which shows a schematic diagram of a structure of a device for visualizing differences after content preprocessing based on a PDF document suitable for implementing an embodiment of the present application. The device for visualizing differences after content preprocessing based on a PDF document in an embodiment of the present application may include various hardware and software components for implementing a method for visualizing differences after content preprocessing based on a PDF document. Figure 6 The device for visualizing the difference after content preprocessing based on the PDF document shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0127] like Figure 6 As shown, the device for visualizing the difference after content preprocessing based on the PDF document may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the device for visualizing the difference after content preprocessing based on the PDF document are also stored. The processing device 1001, the read-only memory 1002 and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the PDF document-based content preprocessing difference visualization device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a PDF document-based content preprocessing difference visualization device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0128] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0129] The device for visualizing the difference of content after preprocessing based on PDF documents provided by the present application adopts the method for visualizing the difference of content after preprocessing based on PDF documents in the above embodiment, which can solve the technical problem that it is difficult to visually compare the difference of content before and after the PDF document is preprocessed. Compared with the prior art, the beneficial effects of the device for visualizing the difference of content after preprocessing based on PDF documents provided by the present application are the same as the beneficial effects of the method for visualizing the difference of content after preprocessing based on PDF documents provided by the above embodiment, and other technical features of the device for visualizing the difference of content after preprocessing based on PDF documents are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0130] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0131] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0132] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the method for visualizing differences after content preprocessing based on a PDF document in the above embodiment.
[0133] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.
[0134] The computer-readable storage medium may be included in the device for visualizing the difference after content preprocessing of a PDF document; or may exist independently without being assembled into the device for visualizing the difference after content preprocessing of a PDF document.
[0135] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a content preprocessing difference visualization device based on a PDF document, the content preprocessing difference visualization device based on a PDF document: extracts text content and image content in a PDF document to obtain a text file and an image file, and records original document information of the PDF document and position information of the text content and the image content in the text file and the image file; constructs a mapping file of the PDF document based on the original document information and the position information, the mapping file including a mapping relationship between the PDF document and the text content and the image content; splits the text file into paragraphs based on the mapping file to obtain paragraph units; inputs the paragraph units and the image file into corresponding processing operators for processing, and integrates them to obtain a target document according to the processing results of the processing operators; outputs the PDF document and the target document in the same display interface, and marks the difference between the PDF document and the target document in the display interface.
[0136] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0137] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0138] The modules involved in the embodiments described in the present application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0139] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned method for visualizing the difference after content preprocessing based on PDF documents, and can solve the technical problem that it is difficult to visually compare the content difference before and after the PDF document is preprocessed. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the method for visualizing the difference after content preprocessing based on PDF documents provided by the above-mentioned embodiments, and will not be described in detail here.
[0140] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for visualizing differences after content preprocessing based on a PDF document.
[0141] The computer program product provided by the present application can solve the technical problem that it is difficult to visually compare the content differences before and after the PDF document is pre-processed. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as the beneficial effects of the method for visualizing the differences after the content pre-processing of the PDF document provided by the above embodiment, and will not be described in detail here.
[0142] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.
[0143] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0144] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0145] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for visualizing differences after content preprocessing of PDF documents, characterized in that: The method for visualizing the difference after content preprocessing based on the PDF document includes: Extracting text content and image content from a PDF document to obtain a text file and an image file, and recording original document information of the PDF document and location information of the text content and the image content in the text file and the image file; Based on the original document information and the location information, construct a mapping file of the PDF document, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content; Based on the mapping file, the text file is split into paragraphs to obtain paragraph units; Input the paragraph unit and the image file into corresponding processing operators for processing respectively, and integrate to obtain a target document according to the processing results of the processing operators; The PDF document and the target document are outputted in a same display interface, and the difference between the PDF document and the target document is marked in the display interface.
2. The method for visualizing the difference after content preprocessing based on PDF documents according to claim 1, characterized in that: The steps of extracting text content and image content from a PDF document, obtaining a text file and an image file, and recording original document information of the PDF document and position information of the text content and the image content in the text file and the image file include: Utilizing a PDF parsing library, extracting the text content and the image content of the PDF document; Determine original document information of the PDF document according to the document structure and semantic features of the PDF document, wherein the original document information includes chapters, paragraphs, page ranges, PDF paragraph coordinates, and image page numbers, image coordinates, and sizes; Matching the text content with the chapters and paragraphs in the PDF document, and obtaining a text line number, an offset, and a text index position corresponding to the text content in the text file according to the chapters and paragraphs; When saving the extracted picture content as the picture file, recording the file path of the picture file; The text line number, offset and text index position of the text content in the text file and the file path of the picture file are integrated to obtain the position information of the text content and the picture content in the text file and the picture file.
3. The method for visualizing the difference after content preprocessing based on PDF documents according to claim 1, characterized in that: The step of constructing a mapping file of the PDF document based on the original document information and the position information, wherein the mapping file includes a mapping relationship between the PDF document and the text content and the image content comprises: Creating a mapping entry according to the original document information, wherein the mapping entry includes chapters, paragraphs, and pictures in the PDF document; Matching the chapters, paragraphs and pictures in the PDF document with the text content and the picture content; Determine, according to the original document information and the position information, a specific position of the mapping entry in the PDF document, and position information of the mapping entry in the text file and the image file; The mapping entries are integrated according to a preset data structure to obtain a mapping file of the PDF document.
4. The method for visualizing the difference of content preprocessing based on PDF documents according to claim 1, characterized in that: The step of splitting the text file into paragraphs based on the mapping file to obtain paragraph units includes: Based on the text line number, offset and text index position of the paragraph in the mapping file, the text content corresponding to each paragraph is extracted from the text file to obtain a plurality of paragraph units.
5. The method for visualizing the difference after content preprocessing based on PDF documents according to claim 1, characterized in that: The step of inputting the paragraph unit and the image file into corresponding processing operators for processing respectively, and integrating to obtain a target document according to the processing results of the processing operators comprises: Distributing the paragraph unit and the picture file to the corresponding processing operator for processing to obtain a processing result; The processing results are integrated according to the structure in the mapping file to obtain a target document.
6. The method for visualizing the difference of content preprocessing based on PDF documents according to claim 1, characterized in that: The step of outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface includes: Determining a page matching relationship between the PDF document and corresponding content in the target document according to the mapping relationship in the mapping file; Based on the page matching relationship, corresponding pages of the PDF document and the target document are displayed side by side in the same display interface; Different colors, marks or highlighting methods are used to mark the different parts between the PDF document and the target document.
7. The method for visualizing the difference of content preprocessing based on PDF documents according to claim 1, characterized in that: The step of outputting the PDF document and the target document in the same display interface, and marking the difference between the PDF document and the target document in the display interface, further includes: When the selection instruction stays on a mark, a processing record of the text content or the picture content corresponding to the mark is displayed based on the process file.
8. The method for visualizing the difference of content preprocessing based on PDF documents according to claim 7, characterized in that: The step of displaying the processing record of the text content or the picture content corresponding to the mark based on the process file when the selection instruction stays on the mark also includes: In the process of processing the paragraph unit and the image file, a process file is created to record the processing record of the processing operator processing the paragraph unit and the image file, wherein the processing record includes the processing start time, end time, processing type, processing status, location information, and the text content and the image content before and after processing; When the selection instruction stays on a mark, the processing record corresponding to the mark is retrieved and extracted from the process file, and displayed in the display interface in the form of a sidebar, a pop-up window or a floating frame.
9. A device for visualizing differences after content preprocessing of PDF documents, characterized in that: The device for visualizing the difference after content preprocessing based on a PDF document comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for visualizing the difference after content preprocessing based on a PDF document as claimed in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the method for visualizing differences after content preprocessing based on a PDF document are implemented as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Document processing method, electronic equipment and storage medium
CN114118011A
PDF archive resource efficient extraction method and text recognition system
CN118762368A
PDF document difference comparison method and related device
CN119720992A
Document comparison tracing method and device and computer storage medium
CN119808752A
A graphical user interface for document review on a computer (Daguan Data Intelligent Document Review System)
CN305828066S