A PDF editing and review method and system on PC

By establishing a text-level feature index and format inheritance method in PDF documents, identifying and handling multi-user modification conflicts, the problem of document format confusion under multi-user collaboration is solved, the efficiency and consistency of document editing are improved, and the fairness of modifications and document quality are ensured.

CN120124600BActive Publication Date: 2025-10-03LUDONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510226765.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-10-03
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing PDF editing tools lack the ability to intelligently handle modification attribution and conflicts in a multi-user collaborative environment, resulting in confusing document formats and incorrect content, affecting the efficiency of document reading and use.

Method used

By collecting and parsing text data in PDF documents, establishing a text-level feature index, setting the format inheritance method, identifying and adjusting format conflicts, identifying the ownership of multi-source modifications, and assigning modification conflict priorities based on authority and impact, the review change tracking results are finally generated.

Benefits of technology

It improves the efficiency and consistency of document editing, reduces the need for manual formatting, optimizes the collaborative review process, ensures the fairness and accuracy of revisions, and improves document quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124600B_ABST
    Figure CN120124600B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of document processing technology, specifically to a PC-side PDF editing and review method and system, comprising the following steps: collecting text data in a PC-side PDF document, parsing a text block structure, analyzing the corresponding positional relationship between the text block and the preceding and following texts, and establishing a text hierarchical feature index. In the present invention, document editing is made more intuitive and efficient through detailed analysis of the position and structural relationship of the text blocks, which is particularly important when processing batch documents. By analyzing the text hierarchy and format inheritance method, the consistency and professionalism of the document format are ensured, the automation level of document processing is improved, and the collaborative review process is optimized by identifying and resolving modification conflicts between multiple reviewers, thereby improving the review efficiency and the quality of the document. The intelligent sorting and processing of modification conflicts not only reduces the burden on the reviewers, but also ensures more fair and accurate modification adoption, thereby improving the overall process and output quality of the document review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document processing, and in particular to a method and system for editing and reviewing PDFs on a PC. Background Art

[0002] The document processing technology field aims to efficiently manage and manipulate digital documents, encompassing a wide range of technologies, from basic text editing to complex content management and automated workflows. Core aspects of this technology area include document creation, editing, storage, retrieval, and sharing, taking into account document security, access control, and integration with information systems. Document processing technology also focuses on improving processing efficiency and user experience, supporting multiple formats and devices, and ensuring data consistency and integrity.

[0003] The PC-based PDF editing and review method refers to a specific technology for efficiently editing and reviewing PDF files. The technical matters covered by this patent subject matter include text editing, annotation addition, content modification, and layout adjustment. Specific technical implementation methods include the use of graphical user interface tools to operate directly on PDF documents, as well as the included tracking changes and commenting system, which enables multiple reviewers to collaborate and record input. This is achieved through a software application designed to run on a personal computer to support users' direct manipulation and modification of PDF files.

[0004] Existing technologies are insufficient for handling large-scale document editing and multi-user collaboration. This is particularly true during the review process, where changes by multiple users can lead to formatting confusion and content errors, impacting the readability and further use of the document. While existing PDF editing tools support basic annotation and change tracking, they are insufficient for efficient and precise handling of complex document structures and format inheritance. Existing technologies also lack support for intelligent conflict resolution and modification attribution. Consequently, during the review process in a collaborative environment, reviewers must spend significant time manually resolving conflicts and errors, reducing work efficiency and compromising document accuracy and consistency due to incorrect handling. Summary of the Invention

[0005] In order to solve the deficiencies in the existing technology in handling large-scale document editing and multi-user collaboration, especially in the review process, modifications by multiple users lead to confusion in the document format and errors in the content, affecting the reading and further use of the document. Although the existing PDF editing tools support basic annotations and modification tracking, when dealing with complex document structures and format inheritance issues, the functions of the tools are not sufficient to meet the needs of efficiency and precision. The existing technology lacks support for intelligent processing of modification attribution and conflicts, resulting in the review process of documents in a multi-person collaborative environment. Reviewers need to spend a lot of time manually resolving conflicts and errors, which not only reduces work efficiency, but also damages the accuracy and consistency of documents due to error handling. The embodiment of the present invention provides a PC-side PDF editing and review method and system. The technical solution is as follows:

[0006] On the one hand, a method for editing and reviewing PDF on a PC is provided, the method comprising:

[0007] S1: Collect text data from PDF documents on the PC, parse the text block structure, analyze the corresponding position relationship between the text block and the surrounding text, and establish a text hierarchical feature index based on the title level and numbering sequence;

[0008] S2: Using the text-level feature index, setting the format inheritance mode, extracting format parameters of text blocks at the same level, identifying the inheritance range of paragraph indentation and line spacing, filtering format inheritance conflicts, and obtaining format inheritance adjustment results;

[0009] S3: Extracting the PDF review modification records based on the format inheritance adjustment results, parsing the review comments, text modifications, and format adjustments, recording the start and end positions of the modified text blocks, evaluating the degree of match between the modified content and the original text blocks, identifying reviewer modification conflict information, and obtaining multi-source modification attribution relationships;

[0010] S4: Using the multi-source modification attribution relationship, filter the multi-source modification records of the same text block, classify and adjust them according to the modification type, and refer to the reviewer authority, modification scope, and modification impact to obtain the modification conflict priority allocation result;

[0011] S5: Based on the modification conflict priority allocation result, the merged modification blocks and unresolved conflicts are recorded, the format consistency of the modified text blocks is evaluated, and the review change tracking result is obtained.

[0012] As a further solution of the present invention, the text level feature index includes hierarchical title codes, position indexes between text blocks, and hierarchically associated text serial numbers; the format inheritance adjustment results include unified font size parameters, color gradient ranges, and standardized paragraph indentation and line spacing settings; the multi-source modification attribution relationship includes the reviewer to whom each modification is attributed, the type identifier of the modification, and the text block position interval; the modification conflict priority allocation result includes the processing priority of each modification type, the reviewer authority level of the modification, and the classification strategy for conflict resolution; the review change tracking result includes the merged modified text block record, the list of unresolved conflicts, and the format consistency evaluation index.

[0013] As a further solution of the present invention, the steps of obtaining the text-level feature index are specifically as follows:

[0014] S101: Collect text data from a PDF document on a PC, extract location information, font features, and paragraph structure of text blocks, identify text block areas with reference to page numbers, coordinate ranges, and text formats, and obtain spatial distribution of text blocks;

[0015] S102: Based on the spatial distribution of the text blocks, analyzing the horizontal spacing, vertical spacing, and paragraph alignment features of adjacent text blocks, recording the arrangement of the text blocks, identifying the hierarchical relationship of the text blocks, determining the logical attribution between the text blocks based on the degree of cohesion of the text content, typesetting rules, and page layout information, matching the order of the text blocks, and obtaining the corresponding relationship between the text block positions;

[0016] S103: using the text block position correspondence, parsing the title level, numbering sequence and hierarchical relationship of the text block, matching the title hierarchy according to the number of words, font size and format of the title text, and generating a text hierarchical feature index.

[0017] As a further solution of the present invention, the steps of obtaining the format inheritance adjustment result are specifically as follows:

[0018] S201: Using the text-level feature index, obtaining format parameters of text blocks at the same level, extracting font size, color gradient, paragraph indentation, and line spacing data of the text blocks, analyzing the distribution range of the format parameters in text blocks at the same level, identifying the change trend of the format parameters, and obtaining the distribution of the format parameters of the text at the same level;

[0019] S202: Based on the format parameter distribution of the text at the same level, a format inheritance mode is set, the format inheritance range between text blocks is identified, the inheritance mode of font size, color gradient, paragraph indentation, and line spacing is analyzed, the deviation range of format parameters is identified, abnormalities in format inheritance of text blocks are screened, and a format inheritance conflict determination result is obtained;

[0020] S203: Using the format inheritance conflict determination result, screening the format mismatch issues of text blocks at the same level, analyzing the priority of format adjustment, adjusting format parameters according to the inheritance relationship, correcting the conflicts of font size, color gradient, paragraph indentation and line spacing, and generating a format inheritance adjustment result.

[0021] As a further solution of the present invention, the steps of obtaining the multi-source modified ownership relationship are specifically as follows:

[0022] S301: Using the format inheritance adjustment result, extracting the review and modification records of the text data in the PDF document on the PC, parsing the review comments, text modifications, deletion operations, and format adjustment contents of the text blocks, recording the text block numbers and modification timestamps associated with the modified contents, and obtaining a review and modification dataset;

[0023] S302: Based on the manuscript review and modification dataset, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of the modified characters, and the text structure changes, evaluate the matching degree of the modified content, determine the scope of the modified text block, and obtain the modification content matching evaluation result;

[0024] S303: Using the modified content matching evaluation result, number the text blocks to which the modified text block belongs, analyze the original source of the modified text, combine text level features and format parameters, match the attribution relationship, identify independent modified texts and merge modified contents, and obtain modified attribution text block data;

[0025] S304: Based on the modification attribution text block data, detect the modifications made by multiple reviewers to the same text block, identify the modification time, content differences and format adjustment deviations, screen modification conflict points, analyze modification priorities, and obtain multi-source modification attribution relationships.

[0026] As a further solution of the present invention, the type of text modification, the length of the modified characters and the change in text structure are analyzed using the formula:

[0027]

[0028] Among them, I m represents the text modification intensity index, C new Represents the number of characters in the modified text, C orig Represents the number of characters in the original text, Δs i represents the change value of the structural modification at the i-th location, and n represents the total number of modified locations.

[0029] As a further solution of the present invention, the step of obtaining the result of modifying the conflict priority allocation is specifically as follows:

[0030] S401: Filtering the multi-source modification records of the same text block based on the multi-source modification attribution relationship, integrating the modified content of the same text block based on the modified text block number, modification type, and timestamp, analyzing the differences in the text modifications, identifying the overlapping areas of the modification ranges, and obtaining a multi-source modification set for the text block;

[0031] S402: Based on the multi-source modification set of the text block, adjusting the text block modification record by modification type classification, identifying the coverage of deletion, replacement, and format adjustment operations, calculating a coverage index, evaluating the impact of the modified text, matching the degree of change of the text before and after the modification, screening key modification content, and obtaining the modification type classification adjustment result;

[0032] S403: Using the modification type classification adjustment result, the priority weight of the modification operation is evaluated according to the reviewer authority, modification scope and modification impact, the processing order of the conflicting modifications is determined, the modification application scope is adjusted, and the modification conflict priority allocation result is generated.

[0033] As a further solution of the present invention, the formula for calculating the coverage index is:

[0034]

[0035] Among them, CA represents the coverage index, M o Represents the feature value after the text at position o is modified, T o represents the feature value before the text at position o is modified, W o represents the text weight coefficient at position o, and M represents the number of text segments.

[0036] As a further solution of the present invention, the steps for obtaining the review change tracking results are specifically as follows:

[0037] S501: Using the modification conflict priority allocation result, record the merged modification blocks and unresolved conflicts, extract the modification text block number, modification type and applicable scope, filter the unresolved conflict records, analyze the unresolved reasons for the modification conflicts, and obtain the merged modification and conflict record data;

[0038] S502: Based on the merged modification and conflict record data, and according to the format inheritance mode, extract the format parameters of the modified text block, evaluate the degree of matching of font size, color gradient, paragraph indentation, and line spacing, analyze the consistency of the format adjustment with the inheritance range, determine whether the format of the modified text complies with the format inheritance mode, and obtain a text format consistency evaluation result;

[0039] S503: Based on the modified text format consistency assessment result, combined with the status of merged modifications and unresolved conflicts, the format adjustment status is recorded, the change scope of the modified text block is matched, and the review change tracking result is generated.

[0040] On the other hand, the PC-side PDF editing and reviewing system is used to execute the above-mentioned PC-side PDF editing and reviewing method, and the system includes:

[0041] The feature recognition module collects text data from PDF documents on the PC, analyzes the positional relationship between text blocks and surrounding text, identifies heading levels and numbering sequences, and establishes a text-level feature index.

[0042] The format inheritance module uses the text-level feature index to set inheritance rules for format parameters of text blocks at the same level in the PDF document on the PC, analyzes font size, color gradient, and paragraph indentation, identifies conflicts in format inheritance, verifies format consistency, and generates format inheritance adjustment results;

[0043] The review and modification module uses the format inheritance adjustment results to extract the review and modification records of the text data in the PDF document on the PC side, identifies text modification, deletion and format adjustment information, analyzes modification conflicts between different reviewers, and generates multi-source modification attribution relationships;

[0044] The modification conflict classification module classifies and organizes the multi-source modification records based on the multi-source modification attribution relationship, assigns priorities according to the modification type and reviewer authority, optimizes the overall consistency of the document, and obtains the modification conflict priority assignment result;

[0045] The change identification module records and merges the conflicting modification blocks according to the modification conflict priority allocation result, evaluates the format consistency of the modified text, and obtains the review change tracking result.

[0046] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0047] By collecting and structurally analyzing text data in documents, the efficiency of text organization and retrieval is improved. Detailed analysis of the location and structural relationships of text blocks makes document editing more intuitive and efficient, which is particularly important when processing batches of documents. By analyzing text hierarchies and format inheritance methods, the consistency and professionalism of document formats are ensured, the need for manual formatting adjustments is reduced, and the level of automation in document processing is improved. By identifying and resolving revision conflicts between multiple reviewers, the collaborative review process is optimized, improving review efficiency and document quality. Intelligent sorting and handling of revision conflicts not only reduces the burden on reviewers, but also ensures more fair and accurate adoption of revisions, improving the overall process and output quality of document review. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0049] Figure 2 This is a detailed flow chart of S1 of the present invention;

[0050] Figure 3 This is a detailed flow chart of S2 of the present invention;

[0051] Figure 4 This is a detailed flow chart of S3 of the present invention;

[0052] Figure 5 This is a detailed flow chart of S4 of the present invention;

[0053] Figure 6 This is a detailed flow chart of S5 of the present invention;

[0054] Figure 7 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0055] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0056] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0057] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0058] See also Figure 1 The embodiment of the present invention provides a method for editing and reviewing PDF on a PC. The processing flow of the method may include the following steps:

[0059] S1: Collect text data from PDF documents on the PC, parse the text block structure, analyze the corresponding position relationship between the text block and the surrounding text, and establish a text hierarchical feature index based on the title level and numbering sequence;

[0060] S2: Use the text-level feature index to set the format inheritance method, extract the format parameters of text blocks at the same level, identify the inheritance range of font size, color gradient, paragraph indentation, and line spacing, filter format inheritance conflicts, analyze the format mismatch of text blocks at the same level, and obtain the format inheritance adjustment results;

[0061] S3: Extract PDF review modification records based on format inheritance adjustment results, analyze review comments, text modifications, deletion operations, and format adjustments, record the start and end positions of modified text blocks, evaluate the degree of match between the modified content and the original text blocks, analyze the modified attribution text blocks, identify multi-reviewer modification conflicts, and obtain multi-source modification attribution relationships;

[0062] S4: Using the multi-source modification attribution relationship, filter the multi-source modification records of the same text block, classify and adjust them according to the modification type, and refer to the reviewer's authority, modification scope, and modification impact to obtain the modification conflict priority allocation result;

[0063] S5: Based on the result of the modification conflict priority assignment, record the merged modification blocks and unresolved conflicts, evaluate the format consistency of the modified text blocks according to the format inheritance method, and obtain the review change tracking results;

[0064] The text level feature index includes hierarchical title codes, position indexes between text blocks, and hierarchically associated text serial numbers. The format inheritance adjustment results include unified font size parameters, color gradient ranges, and standardized paragraph indentation and line spacing settings. The multi-source modification attribution relationship includes the reviewer of each modification, the type identifier of the modification, and the text block position interval. The modification conflict priority allocation results include the processing priority of each modification type, the reviewer authority level of the modification, and the classification strategy for conflict resolution. The review change tracking results include the merged modified text block records, the list of unresolved conflicts, and the evaluation indicators of format consistency.

[0065] See also Figure 2 , the steps for obtaining the text-level feature index are as follows:

[0066] S101: Collect text data from a PDF document on a PC, extract location information, font features, and paragraph structure of text blocks, identify text block areas with reference to page numbers, coordinate ranges, and text formats, and obtain spatial distribution of text blocks;

[0067] Extracting text data from a PDF document involves multiple steps. It requires obtaining the binary data of the PDF file and parsing the text content. Using a text parsing library such as pdfp lumber or PyMuPDF, the text blocks of each page are obtained and classified according to coordinate information. For example, for the standard A4 paper size (210mm×297mm), pixel coordinate mapping can be used (such as 595×842 pixels at a 72dpi resolution). For different font information, the font name, font size, bold, italic and other features are extracted through the OCR engine or the built-in font parsing library. The font size can be determined by the glyph boundary (boundary i ngbox) calculation. For example, if the character height of a certain paragraph of text is 12 pixels, the font size can be inferred to be 12pt. All text blocks need to be clustered according to page number and coordinate range for subsequent spatial distribution analysis. When extracting text blocks, it is necessary to consider whether there is a multi-column layout. For a double-column format, the separator between the two columns can be determined by detecting the X-coordinate distribution of the text blocks (for example, if the X-coordinate range of one column of text blocks is 50-250px and the X-coordinate range of the other column is 300-550px). Finally, for non-pure text data, such as tables and embedded text in images, table detection algorithms (such as OpenCV edge detection or Tabu La's table recognition method) can be combined to obtain the text content and the spatial distribution of text blocks.

[0068] S102: Based on the spatial distribution of text blocks, analyze the horizontal spacing, vertical spacing, and paragraph alignment features of adjacent text blocks, record the arrangement of the text blocks, identify the hierarchical relationship of the text blocks, determine the logical attribution between the text blocks based on the degree of text content cohesion, typesetting rules, and page layout information, match the order of the text blocks, and obtain the corresponding relationship between the text block positions;

[0069] It is necessary to further calculate the horizontal and vertical spacing of adjacent text blocks to determine the logical relationship of the text blocks. If the distance between two text blocks on the X axis is less than the set paragraph spacing threshold (such as 10px) and the alignment is consistent, they can be determined to be parts of the same paragraph. For text blocks with obvious indentation (such as the first line indentation exceeds 20px), it can be inferred that it is the first line of a paragraph. For the judgment of paragraph spacing, the minimum line spacing corresponding to different font sizes can be set. The line spacing of text with a 12pt font size is about 14-16px. If the spacing between two text blocks exceeds this value, 1.5 times (such as 24px) can be inferred as different paragraphs. When identifying the hierarchical relationship of text blocks, it is necessary to combine the text format. For example, the title has a larger font size (such as 14pt or above) and uses bold or center alignment format. The hierarchical relationship of the text can be established by matching features. For example, if there are four title text blocks on a PDF page, and their font sizes are 16pt, 14pt, 12pt, and 10pt respectively, a four-level title structure can be constructed, and the previous and next text blocks can be matched according to the order of text content connection to obtain the corresponding relationship between the text block positions.

[0070] S103: Using the position correspondence of the text blocks, parsing the title level, numbering sequence, and hierarchical relationship of the text blocks, matching the title hierarchy according to the number of words, font size, and format of the title text, and generating a text hierarchical feature index;

[0071] When parsing the title level of a text block, it is necessary to analyze it based on its font size, numbering format, alignment and other information. For titles numbered with Arabic numerals (such as "1.1"), their hierarchical relationship can be matched according to the numbering format. Common rules include using a larger font size (such as 16pt) for the first-level title, a slightly smaller font size (such as 14pt) for the second-level title, and a decreasing font size (such as 12pt) for the third-level title. For automatic matching of title levels, regular expressions can be used to extract title numbers. When extracting the hierarchical number, the hierarchical structure is determined in combination with font features. By constructing a text hierarchical feature index, According to the correspondence between the title and the text, a structured index similar to a directory is formed. A document is set to contain chapters "1. Data Processing", "1.1 Data Collection", and "1.1.1 Sensor Data Acquisition". It can be parsed into a tree hierarchy and stored as structured data (such as JSON format). At the same time, the hierarchical division of the text content can be based on the alignment and indentation of the subsequent text blocks of the title. If the indentation of the first line of the text block exceeds the set threshold (such as 20px), it can be inferred that it is the beginning of a new paragraph, otherwise it is the continuation of the current paragraph. The text structure is parsed and a text hierarchical feature index is generated.

[0072] See also Figure 3 , the specific steps for obtaining the format inheritance adjustment results are as follows:

[0073] S201: Using a text-level feature index, obtain format parameters of text blocks at the same level, extract font size, color gradient, paragraph indentation, and line spacing data of the text blocks, analyze the distribution range of format parameters in text blocks at the same level, identify the change trend of format parameters, and obtain the distribution of format parameters of text blocks at the same level;

[0074] When obtaining the formatting parameters of text blocks at the same level, it is necessary to extract the text block's font size, color gradient, paragraph indentation, and line spacing data, and analyze the distribution range of the formatting parameters within the text blocks at the same level. All text blocks at the same level are traversed to obtain the font information of each text block, including font size, font type, bold or italic style, etc. By extracting the text coordinate information, the vertical and horizontal spacing between text blocks is calculated to identify the line spacing and paragraph indentation. The color gradient calculation requires detecting the text color information and analyzing its distribution in RGB or CMYK color mode. If the font sizes of text blocks at the same level are mainly concentrated between 12pt and 14pt, and a text block has a font size of 16pt, it is necessary to record the abnormality of the formatting parameters and analyze the changing trend of the formatting parameters. If the formatting of text blocks on a certain page gradually increases or the indentation changes significantly, it means that the hierarchy adjustment or typesetting within the document is inconsistent, and further inspection is required to ensure compliance with the overall formatting specifications. Based on the statistical results of the formatting parameters of all text blocks, the formatting parameter distribution of text at the same level is generated.

[0075] S202: Based on the format parameter distribution of text at the same level, a format inheritance mode is set, the format inheritance range between text blocks is identified, the inheritance mode of font size, color gradient, paragraph indentation, and line spacing is analyzed, the deviation range of format parameters is identified, abnormalities in format inheritance of text blocks are screened, and a format inheritance conflict determination result is obtained;

[0076] When setting the format inheritance method, you need to define the inheritance range of different format parameters based on the distribution of format parameters of text at the same level. You also need to analyze the inheritance mode of font size, color gradient, paragraph indentation, and line spacing to set the reference range for format inheritance. If the font sizes of text blocks at the same level are mainly between 12pt and 14pt, you can set this range as the default inheritance range. If the font size of a text block is 10pt or 16pt, this is a format deviation and requires further inspection to ensure compliance with the inheritance rules. Color gradient inheritance requires analysis of text color changes. For example, if the color of text blocks at the same level is mainly dark gray (RGB50, 50, 50) and a text block is pure black (RGB0, 0, 0), this will lead to inconsistent layout and require inspection to ensure compliance with the inheritance mode. Paragraph indentation and line spacing analysis is mainly used to identify text alignment. For example, if the indentation of all text blocks is 20px, but the indentation of one text block is 40px, this is a format anomaly. You need to record this deviation and filter all format parameter inheritance anomalies. By comparing the format inheritance relationships of all text blocks, you can obtain the format inheritance conflict determination results.

[0077] S203: Using the format inheritance conflict determination result, screen for format mismatches among text blocks at the same level, analyze the priority of format adjustment, adjust format parameters based on the inheritance relationship, correct conflicts in font size, color gradient, paragraph indentation, and line spacing, and generate a format inheritance adjustment result;

[0078] When filtering out format mismatches between text blocks at the same level, you need to analyze the priority of format adjustments and adjust format parameters based on inheritance relationships to correct conflicts in font size, color gradient, paragraph indentation, and line spacing. Set adjustment priorities based on the impact range of format conflicts. For example, font size adjustment has a higher priority than color gradient, which has a higher priority than indentation adjustment, while line spacing adjustment has the lowest priority. If the font size of a text block exceeds the inheritance range, you need to adjust its font size to the default range, for example, adjust 16pt to 14pt to match the text block at the same level. The color gradient can be adjusted by interpolation to gradually transition it to be consistent with the surrounding text. If the color of a text block is pure black and the surrounding text is dark gray, its color can be adjusted to an intermediate level (such as RGB25, 25, 25). Indentation adjustment mainly involves the correction of alignment. If the indentation of a text block is obviously too large, its indentation value needs to be adjusted to the standard range, such as 40px to 20px. By adjusting all format parameters and ensuring that the format of all text blocks complies with the inheritance rules, the document's layout consistency can be improved and the format inheritance adjustment results can be generated.

[0079] See also Figure 4 The specific steps for obtaining the multi-source modification ownership relationship are as follows:

[0080] S301: Using the format inheritance adjustment result, extract the review modification records of the text data in the PDF document on the PC, parse the review comments, text modification, deletion operations and format adjustment content of the text block, record the text block number and modification timestamp associated with the modification content, and obtain the review modification dataset;

[0081] Extracting the review and modification records of PC-side PDF documents involves multiple steps. It is necessary to parse the annotation layer in the PDF file to obtain the metadata of all review modifications. For example, using PyMuPDF to parse PDF annotations can extract the text content, modification type (such as addition, deletion, replacement), reviewer name, timestamp and other information of each modification record. At the same time, the number of the text block is parsed to match the correspondence between the modified content and the original text. If a text block is numbered T123, and its modification content involves deleting the original text "data processing" and replacing it with "data optimization", then the record Modified content ("data optimization"), original text ("data processing"), modification time (such as 2025-02-2010:30:15), modification type (replacement), and modifier ID. For records of deletion operations, the text before deletion can be extracted and the deletion status can be marked in the review modification dataset. For format adjustment content, such as font size changes, color changes, etc., the parameters of the original format and the modified format can be compared. For example, if the font size is adjusted from 12pt to 14pt, the format adjustment content can be recorded and all related information of the modification can be stored to form the review modification dataset.

[0082] S302: Based on the manuscript review and modification dataset, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of the modified characters, and the text structure changes, evaluate the matching degree of the modified content, determine the scope of the modified text block, and obtain the modification content matching evaluation result;

[0083] Analyze the type of text modification, the length of modified characters, and the changes in text structure using the formula:

[0084]

[0085] Among them, I m represents the text modification intensity index, C new Represents the number of characters in the modified text, C orig Represents the number of characters in the original text, Δs i represents the change value of the structural modification at position i, and n represents the total number of modified positions;

[0086] Parameter meaning:

[0087] C new is the number of characters in the modified text, obtained by counting the total number of characters in the modified text;

[0088] Corig is the number of characters in the original text, obtained by counting the total number of characters in the original text;

[0089] Δs i is the modification value of the text structure at position i, which represents the difference between the text structure before and after the modification. It is calculated by comparing the text structure features before and after the modification (such as the number of paragraphs, sentence length, etc.);

[0090] n is the total number of modification positions in the text, which is determined by detecting the number of all modification points in the text;

[0091] Calculation derivation process:

[0092] Assume that the original text contains 1000 characters, the modified text contains 1050 characters, there are 5 structural modifications, and the structural modification values ​​of each (for example, the sentence changes from 10 words to 12 words) are 2, 1, 0, 3, and 4 respectively;

[0093] Calculate the change in the number of characters:

[0094] |C new -C orig |=|1050-1000|=50;

[0095] Calculate the sum of squares of structural changes:

[0096]

[0097] Calculate the square root of the sum of the squares of the structural changes:

[0098]

[0099] Substitute the values ​​into the formula to calculate the text modification intensity index:

[0100]

[0101] The results show that the text modification intensity index is 0.05548, which reflects the degree of change from the original text to the modified text. This index reveals the overall intensity of text modification, helps to evaluate the depth and scope of text modification, and is an important metric for evaluating the similarity and structural changes between text modifications and the original text.

[0102] S303: Using the modified content matching evaluation results, number the text blocks to which the modified text block belongs, analyze the original source of the modified text, combine text level features and format parameters, match the attribution relationship, identify independent modified texts and merge modified content, and obtain the modified attribution text block data;

[0103] It is necessary to number the ownership relationship of the modified text blocks and match their original sources, and extract the associated text blocks of the modified text. If the modified text belongs to T101, the modification record of T101 is updated in the dataset. At the same time, combined with the text level characteristics and format parameters, the format inheritance of the modification is analyzed. For example, if the font size of a modified text block is 14pt, and the font size of its affiliated text block is 12pt, the record format deviation ΔF = 2pt is used for subsequent adjustment. For merged modification content, if multiple modified text blocks are adjacent and there is continuity in the modification record, they are merged into one modification block. For example, the modification contents of T201 and T202 are "optimize model parameters" and "improve calculation accuracy" respectively. If the modification time interval is less than the set threshold (such as 5 minutes), they are merged into one modification record "optimize model parameters and improve calculation accuracy". For independent modification text, if a modification content fails to match the affiliated text block, it is marked as an independent modification and the modification affiliated text block data is obtained.

[0104] S304: Based on the modification attribution text block data, detect the modifications made by multiple reviewers to the same text block, identify modification time, content differences, and format adjustment deviations, screen modification conflicts, analyze modification priorities, and obtain multi-source modification attribution relationships;

[0105] When multiple reviewers make revisions, it is necessary to detect the differences in modifications to the same text block, extract all modification records, and group them according to the text block number. For text block T305, extract the modification content of all modification records A, B, and C, and calculate the modification time interval. For example, if A modified it at 10:30, B modified it at 10:32, and C modified it at 10:35, then record the time sequence, analyze the content differences, and calculate the text similarity after modification by different reviewers. For example, if A modifies it to "improve calculation efficiency", B modifies it to "optimize calculation method", and C modifies it to "adjust calculation formula", then the AB similarity is calculated to be 0.75, BC similarity is 0.6, and AC similarity is 0.55. If the modification difference is greater than the set threshold (such as 0.7), it is marked as a conflict and sorted according to modification priority. If A is the main editor and its modification priority is higher than B and C, then A's modification is adopted first, and the multi-source modification attribution relationship is stored.

[0106] See also Figure 5 , the specific steps for obtaining the conflict priority allocation result are as follows:

[0107] S401: Filtering multi-source modification records of the same text block based on the multi-source modification attribution relationship, integrating the modified content of the same text block based on the modified text block number, modification type, and timestamp, analyzing the differences in the text modifications, identifying the overlapping areas of the modification ranges, and obtaining a multi-source modification set for the text block;

[0108] When filtering multi-source modification records for the same text block, it is necessary to integrate the modifications of different reviewers and classify them according to the modified text block number, modification type, and timestamp. All relevant modification records are filtered by the text block number (such as T501) and classified according to the modification type. If the text block T501 has three modification records A, B, and C, involving deletion (A), replacement (B), and format adjustment (C) respectively, the modification contents are stored in a classified manner. Next, by analyzing the differences in the text modifications, the similarity of the modification contents is calculated using the formula:

[0109]

[0110] If the modified content of A is "Optimize calculation method" and the modified content of B is "Improve calculation efficiency", the number of characters in the intersection of the two is 6 and the total number of characters is 12, then calculate the similarity:

[0111]

[0112] If the similarity is lower than the set threshold (such as 0.6), it is determined that the content difference is large, and the overlapping areas of the modification range are further analyzed. If the modification content of A involves the text "computational model optimization" and the modification of B involves "computational model adjustment", the proportion of the overlapping areas of the two is calculated. If the overlap rate of the "computational model" part is 80%, the part is identified as the overlapping area, forming a multi-source modification set of the text block.

[0113] S402: Based on the multi-source modification set of the text block, the text block modification record is adjusted according to the modification type classification, the coverage of the deletion, replacement, and format adjustment operations is identified, the coverage index is calculated, the impact range of the modified text is evaluated, the change degree of the text before and after the modification is matched, the key modification content is screened, and the modification type classification adjustment result is obtained;

[0114] The formula for calculating the coverage index is:

[0115]

[0116] Among them, CA represents the coverage index, M o Represents the feature value after the text at position o is modified, T o represents the feature value before the text at position o is modified, W o represents the weight coefficient of the text at position o, and M represents the number of text segments;

[0117] Parameter meaning:

[0118] Text feature value M o and T o Calculation:

[0119] Character edit distance: Calculates the edit distance between the corresponding segments before and after the text modification, measures character-level changes using Levenshtein distance, and normalizes it to the 0.1 range to make texts of different sizes comparable.

[0120] Based on word vector similarity: word embedding (Word2Vec, BERT) is used to calculate the cosine similarity between sentences, and the similarity value is taken as the text feature value;

[0121] Based on syntactic structure analysis: use dependency syntactic analysis to calculate the degree of sentence structure change, combined with the syntactic tree transformation rate assignment;

[0122] Weight coefficient W o Calculation:

[0123] Word frequency weight: Calculate the importance of each text segment based on TFIDF. If the text modification involves high-weight words, a high W is assigned. o value;

[0124] Sentence position weight: Important paragraphs such as the title, first sentence, and ending of the text have high weights. The position factor is normalized to the range of 0.1 and used as a correction factor for the weight coefficient;

[0125] Modification type weight: Deletion, replacement, formatting, and other operations have different impacts. To ensure that the contribution of each type of modification matches, a basic weight value is set for each modification type.

[0126] Specific calculation example:

[0127] The data of a certain text before and after modification are as follows: the number of text segments M = 5;

[0128] Eigenvalue T before modification o : T1=0.8, T2=0.6, T3=0.7, T4=0.5, T5=0.9;

[0129] Modified eigenvalue M o : M1=0.3, M2=0.9, M3=0.6, M4=0.4, M5=0.8;

[0130] Weight coefficient W o : W1=0.5, W2=0.7, W3=0.8, W4=0.6, W5=0.9;

[0131] Calculate|M o -T o |:

[0132] |M1-T1|=|0.3-0.8|=0.5;

[0133] |M2-T2|=|0.9-0.6|=0.3;

[0134] |M3-T3|=|0.6-0.7|=0.1;

[0135] |M4-T4|=|0.4-0.5|=0.1;

[0136] |M5-T5|=|0.8-0.9|=0.1;

[0137] Calculate the molecular part:

[0138]

[0139] Calculate the denominator:

[0140]

[0141] Substitute into the formula for calculation:

[0142]

[0143] The results show that the coverage index of text modification is 0.321, indicating that there are certain differences between the text before and after modification, but the overall degree of change is low. This value can be used as an adjustment basis in the subsequent text classification adjustment process.

[0144] S403: Using the modification type classification adjustment results, the priority weight of the modification operation is evaluated based on the reviewer's authority, the modification scope, and the degree of modification impact, and the order in which conflicting modifications are handled is determined. The scope of modification application is adjusted to generate a modification conflict priority allocation result;

[0145] It is necessary to evaluate the priority weight of the modification operation to determine the order of handling modification conflicts. The weight is assigned according to the reviewer's authority. The editor's modification weight W = 1.0, the ordinary editor's modification weight W = 0.8, and the external reviewer's modification weight W = 0.5. For multiple modifications to the same text block, the weighted modification impact S is calculated. W as follows:

[0146]

[0147] Among them, I i represents the impact range of the i-th modification, and n is the number of modifications;

[0148] Assume that the modification impact ranges of the three reviewers are 20, 30, and 50 characters respectively, and the corresponding weights are 1.0, 0.8, and 0.5 respectively. The calculation is as follows:

[0149] S W =(1.0×20)+(0.8×30)+(0.5×50)=64;

[0150] If a modified S W If the value is higher than 20% of the modification, the modification is retained first and the scope of the modification is analyzed. If one modification involves the entire text, while another modification only affects a single word, the modification with the larger coverage is preferred. The scope of application of the modification is adjusted based on the scope of modification impact, permissions, timestamps and other factors, and the priority allocation result of the modification conflict is generated.

[0151] See also Figure 6 The specific steps for obtaining the review change tracking results are as follows:

[0152] S501: Using the modification conflict priority allocation result, record the merged modification blocks and unresolved conflicts, extract the modification text block number, modification type and applicable scope, filter the unresolved conflict records, analyze the unresolved reasons for the modification conflicts, and obtain the merged modification and conflict record data;

[0153] Filter all text blocks involved in the modification and organize them according to the text block number, modification type and applicable scope. For the merged modification blocks, it is necessary to extract the modification source, applicable scope and adopted modification version. A certain text block T701 has undergone three rounds of modification. The first round deleted part of the content, the second round replaced part of the text, and the third round adjusted the format. The text is mainly based on the content of the second round of modification, and the format adjustment is retained, while the deletion operation is not adopted. Therefore, it is necessary to mark T701 as adopting the second round of modification content in the merge record, and attach format adjustment information. For unresolved conflict records, it is necessary to identify the reasons for their unresolved nature. If a certain text block T702 has multiple reviewers, If there are conflicts in the formatting, such as one reviewer adjusting the indentation and another reviewer changing the font size, and the two are incompatible, the modification conflict cannot be automatically merged, and the text block needs to be additionally marked as "formatting conflict pending". Further analysis of the scope of the conflict is required. For example, if a text block only involves local phrase modification, and another text block involves adjustment of the entire paragraph, the former has a smaller conflict and can wait for manual intervention, while the latter affects the layout of the entire document and needs to be handled first. All merged modification blocks and unresolved conflicts need to be recorded for subsequent format matching and modification tracking to obtain merged modification and conflict record data.

[0154] S502: Based on the merged modification and conflict record data and the format inheritance mode, the format parameters of the modified text block are extracted, the matching degree of font size, color gradient, paragraph indentation, and line spacing is evaluated, the consistency of the format adjustment with the inheritance range is analyzed, and whether the format of the modified text conforms to the format inheritance mode is determined, and the text format consistency evaluation result is obtained;

[0155] Extract the format parameters of all modified text blocks, including font size, color gradient, paragraph indentation and line spacing, and compare them with the format inheritance method. If a text block T801 is modified and the original font size changes from 12pt to 14pt, and the default format of the text block at the same level is 12pt, it is necessary to determine whether the modification complies with the format inheritance method. If 14pt is within the inheritance range, the modification complies with the format inheritance. Otherwise, it is necessary to adjust back to the default format or redefine the inheritance range. The color gradient evaluation requires detecting the color change of the modified text. If the original text color is dark gray and changes to black after modification, it is necessary to determine whether the color gradient change is within the acceptable range. If the change is If the text is too large, the overall layout of the document will be inconsistent and it needs to be adjusted back to the original color. The matching of indentation parameters is mainly for the first line indent and the overall alignment. If a text block originally used the left-aligned format, but was changed to center-aligned after modification, it is necessary to check whether the format meets the alignment requirements of the current text level. If not, it needs to be adjusted. The evaluation of line spacing parameters is mainly used to check changes in text density. If the original text block uses 1.5 times the line spacing and is changed to single line spacing after modification, it affects readability and needs to be checked for compliance with the format inheritance standard. By comparing the format parameters of the text blocks before and after modification, it is determined that some modifications meet the inheritance rules and some need to be corrected to generate the text format consistency assessment results.

[0156] S503: Modify the text format consistency assessment results, combine the status of merged modifications and unresolved conflicts, record the format adjustment status, match the change scope of the modified text block, and generate the review change tracking results;

[0157] It is necessary to combine the status of merged modifications and unresolved conflicts, record the format adjustment status, and match the change range of the modified text block. For all merged modified text blocks, store the modified content adopted and record its format adjustment status. If a text block T901 has undergone multiple rounds of modifications and uses the content of the second round of modifications, and the format is adjusted to 1.2 times the line spacing and 14pt font size, it is necessary to save the modification record and mark the format adjustment as "confirmed". Secondly, for unresolved conflict text blocks, it is necessary to mark their current status. For example, if T902 is still in an unresolved format conflict state, mark it as "pending" and record its status. Specific problem points, such as "indent mismatch" or "color gradient changes too large", need to match the change scope of the modified text block to determine the impact of the modification on the entire document. For example, if a modification only involves local phrase adjustment, the impact is small, but if the modification involves changes to the entire paragraph and affects the format of the surrounding text, it is necessary to record the scope of the modification and evaluate whether additional adjustments to the format of the text block are needed to ensure overall consistency. The status, format adjustment status and unresolved conflict records of all modified text blocks must be integrated into the review change tracking data to facilitate subsequent format correction and document typesetting, and generate review change tracking results.

[0158] like Figure 7 As shown, a PC-side PDF editing and review system includes:

[0159] The feature recognition module collects text data from PDF documents on the PC, analyzes the positional relationship between text blocks and surrounding text, identifies heading levels and numbering sequences, and establishes a text-level feature index.

[0160] The format inheritance module uses the text-level feature index to set inheritance rules for the format parameters of text blocks at the same level in the PC-side PDF document. It analyzes font size, color gradient, and paragraph indentation, identifies conflicts in format inheritance, verifies format consistency, and generates format inheritance adjustment results.

[0161] The review and modification module uses the format inheritance adjustment results to extract the review and modification records of the text data in the PDF document on the PC side, identify text modification, deletion and format adjustment information, analyze the modification conflicts between different reviewers, and generate multi-source modification attribution relationships;

[0162] The modification conflict classification module classifies and organizes multi-source modification records based on the multi-source modification ownership relationship, assigns priorities based on modification type and reviewer permissions, optimizes the overall consistency of the document, and obtains the modification conflict priority allocation results;

[0163] The change identification module assigns the result according to the modification conflict priority, records and merges the conflicting modification blocks, evaluates the format consistency of the modified text, and obtains the review change tracking result.

[0164] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A PDF editing and review method on a PC, characterized in that: The following steps are involved: S1: Collect text data from PDF documents on the PC, parse the text block structure, analyze the corresponding position relationship between the text block and the surrounding text, and establish a text hierarchical feature index based on the title level and numbering sequence; S2: Using the text-level feature index, setting the format inheritance mode, extracting format parameters of text blocks at the same level, identifying the inheritance range of paragraph indentation and line spacing, filtering format inheritance conflicts, and obtaining format inheritance adjustment results; S3: Extracting the PDF review modification records based on the format inheritance adjustment results, parsing the review comments, text modifications, and format adjustments, recording the start and end positions of the modified text blocks, evaluating the degree of match between the modified content and the original text blocks, identifying reviewer modification conflict information, and obtaining multi-source modification attribution relationships; S4: Using the multi-source modification attribution relationship, filter the multi-source modification records of the same text block, classify and adjust them according to the modification type, and refer to the reviewer authority, modification scope, and modification impact to obtain the modification conflict priority allocation result; S5: Based on the modification conflict priority allocation result, record the merged modification blocks and unresolved conflicts, evaluate the format consistency of the modified text blocks, and obtain the review change tracking result; The steps for obtaining the multi-source modified ownership relationship are specifically as follows: S301: Using the format inheritance adjustment result, extracting the review and modification records of the text data in the PDF document on the PC, parsing the review comments, text modifications, deletion operations, and format adjustment contents of the text blocks, recording the text block numbers and modification timestamps associated with the modified contents, and obtaining a review and modification dataset; S302: Based on the manuscript review and modification dataset, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of the modified characters, and the text structure changes, evaluate the matching degree of the modified content, determine the scope of the modified text block, and obtain the modification content matching evaluation result; S303: Using the modified content matching evaluation result, number the text blocks to which the modified text block belongs, analyze the original source of the modified text, combine text level features and format parameters, match the attribution relationship, identify independent modified texts and merge modified contents, and obtain modified attribution text block data; S304: Based on the modification attribution text block data, detecting the modifications made by multiple reviewers to the same text block, identifying modification time, content differences, and format adjustment deviations, screening modification conflict points, analyzing modification priorities, and obtaining multi-source modification attribution relationships; Analyze the type of text modification, the length of modified characters, and the changes in text structure using the formula: ; in, Represents the text modification intensity index, Represents the number of characters in the modified text. Represents the number of characters of the original text, Represents the change value of the structural modification at position i, Represents the total number of modified positions.

2. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, wherein: The text level feature index includes hierarchical title codes, position indexes between text blocks, and hierarchically associated text serial numbers; the format inheritance adjustment results include unified font size parameters, color gradient ranges, and standardized paragraph indentation and line spacing settings; the multi-source modification attribution relationship includes the reviewer to whom each modification is attributed, the type identifier of the modification, and the text block position interval; the modification conflict priority allocation result includes the processing priority of each modification type, the reviewer authority level of the modification, and the classification strategy for conflict resolution; the review change tracking result includes the merged modified text block record, a list of unresolved conflicts, and format consistency evaluation indicators.

3. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, wherein: The steps for obtaining the text-level feature index are specifically as follows: S101: Collect text data from a PDF document on a PC, extract location information, font features, and paragraph structure of text blocks, identify text block areas with reference to page numbers, coordinate ranges, and text formats, and obtain spatial distribution of text blocks; S102: Based on the spatial distribution of the text blocks, analyzing the horizontal spacing, vertical spacing, and paragraph alignment features of adjacent text blocks, recording the arrangement of the text blocks, identifying the hierarchical relationship of the text blocks, determining the logical attribution between the text blocks based on the degree of cohesion of the text content, typesetting rules, and page layout information, matching the order of the text blocks, and obtaining the corresponding relationship between the text block positions; S103: using the text block position correspondence, parsing the title level, numbering sequence and hierarchical relationship of the text block, matching the title hierarchy according to the number of words, font size and format of the title text, and generating a text hierarchical feature index.

4. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, wherein: The steps for obtaining the format inheritance adjustment result are specifically as follows: S201: Using the text-level feature index, obtaining format parameters of text blocks at the same level, extracting font size, color gradient, paragraph indentation, and line spacing data of the text blocks, analyzing the distribution range of the format parameters in text blocks at the same level, identifying the change trend of the format parameters, and obtaining the distribution of the format parameters of the text at the same level; S202: Based on the format parameter distribution of the text at the same level, a format inheritance mode is set, the format inheritance range between text blocks is identified, the inheritance mode of font size, color gradient, paragraph indentation, and line spacing is analyzed, the deviation range of format parameters is identified, abnormalities in format inheritance of text blocks are screened, and a format inheritance conflict determination result is obtained; S203: Using the format inheritance conflict determination result, screening the format mismatch issues of text blocks at the same level, analyzing the priority of format adjustment, adjusting format parameters according to the inheritance relationship, correcting the conflicts of font size, color gradient, paragraph indentation and line spacing, and generating a format inheritance adjustment result.

5. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, wherein: The steps for obtaining the result of modifying the conflict priority allocation are specifically as follows: S401: Filtering the multi-source modification records of the same text block based on the multi-source modification attribution relationship, integrating the modified content of the same text block based on the modified text block number, modification type, and timestamp, analyzing the differences in the text modifications, identifying the overlapping areas of the modification ranges, and obtaining a multi-source modification set for the text block; S402: Based on the multi-source modification set of the text block, adjusting the text block modification record by modification type classification, identifying the coverage of deletion, replacement, and format adjustment operations, calculating a coverage index, evaluating the impact of the modified text, matching the degree of change of the text before and after the modification, screening key modification content, and obtaining the modification type classification adjustment result; S403: Using the modification type classification adjustment result, the priority weight of the modification operation is evaluated according to the reviewer authority, modification scope and modification impact, the processing order of the conflicting modifications is determined, the modification application scope is adjusted, and the modification conflict priority allocation result is generated.

6. The method for editing and reviewing PDF manuscripts on a PC according to claim 5, wherein: The formula for calculating the coverage index is: ; in, represents the coverage index, Representative The feature value after the text is modified, Representative The feature value before the text is modified, Representative The text weight coefficient, Represents the number of text segments.

7. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, wherein: The specific steps for obtaining the review change tracking results are as follows: S501: Using the modification conflict priority allocation result, record the merged modification blocks and unresolved conflicts, extract the modification text block number, modification type and applicable scope, filter the unresolved conflict records, analyze the unresolved reasons for the modification conflicts, and obtain the merged modification and conflict record data; S502: Based on the merged modification and conflict record data, and according to the format inheritance mode, extract the format parameters of the modified text block, evaluate the degree of matching of font size, color gradient, paragraph indentation, and line spacing, analyze the consistency of the format adjustment with the inheritance range, determine whether the format of the modified text complies with the format inheritance mode, and obtain a text format consistency evaluation result; S503: Based on the modified text format consistency assessment result, combined with the status of merged modifications and unresolved conflicts, the format adjustment status is recorded, the change scope of the modified text block is matched, and the review change tracking result is generated.

8. A PC-based PDF editing and review system, characterized by: The method for editing and reviewing PDF manuscripts on a PC according to any one of claims 1 to 7, wherein the system comprises: The feature recognition module collects text data from PDF documents on the PC, analyzes the positional relationship between text blocks and surrounding text, identifies heading levels and numbering sequences, and establishes a text-level feature index. The format inheritance module uses the text-level feature index to set inheritance rules for format parameters of text blocks at the same level in the PDF document on the PC, analyzes font size, color gradient, and paragraph indentation, identifies conflicts in format inheritance, verifies format consistency, and generates format inheritance adjustment results; The review and modification module uses the format inheritance adjustment results to extract the review and modification records of the text data in the PDF document on the PC side, identifies text modification, deletion and format adjustment information, analyzes modification conflicts between different reviewers, and generates multi-source modification attribution relationships; The modification conflict classification module classifies and organizes the multi-source modification records based on the multi-source modification attribution relationship, assigns priorities according to the modification type and reviewer authority, optimizes the overall consistency of the document, and obtains the modification conflict priority assignment result; The change identification module records and merges the conflicting modification blocks according to the modification conflict priority allocation result, evaluates the format consistency of the modified text, and obtains the review change tracking result.

Citation Information

Patent Citations

  • Method for editing PDF (Portable Document Format) document and terminal equipment

    CN107977346A

  • Parallel editing method for splitting Word document into multiple rich texts

    CN117252164A