PDF editing and manuscript reviewing method and system of PC terminal

By establishing text-level feature index and format inheritance methods in PDF editing tools, identifying and adjusting format inheritance conflicts, and optimizing multi-source modification conflicts, the problems of inconsistent document formats and wrong content in the existing technology are solved, and an efficient and accurate document editing and review process is achieved.

CN120124600AActive Publication Date: 2025-06-10LUDONG UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510226765.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-10
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

When existing PDF editing tools handle large-scale document editing and multi-user collaboration, it is difficult to maintain the consistency of document formats and the accuracy of content, resulting in confusion in formats and content errors during the review process, affecting the reading and further use of documents.

Method used

By collecting and parsing text data in PDF documents, establishing text-level feature indexes, setting format inheritance methods, identifying and adjusting format inheritance conflicts, evaluating the degree of matching of modified content, identifying and optimizing multi-source modification conflicts, and generating review changes tracking results.

Benefits of technology

It improves the efficiency and accuracy of document editing, ensures the consistency and professionalism of document formats, optimizes the multi-person collaborative review process, reduces the need for manual format adjustment, and improves the automation level of document processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124600A_ABST
    Figure CN120124600A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of document processing, in particular to a PDF (Portable Document Format) editing and manuscript reviewing method and system for a PC (Personal Computer) end, which comprises the following steps of: collecting text data in a PDF document of the PC end, analyzing a text block structure, analyzing a corresponding position relationship between a text block and front and back texts, and establishing a text level feature index. According to the method, document editing is more intuitive and efficient through detailed analysis of the positions and the structural relations of the text blocks and is particularly important when batch documents are processed, the consistency and the speciality of document formats are ensured through analysis of the text hierarchy and the format inheritance mode, the automation level of document processing is improved, and the document editing efficiency is improved. By identifying and solving the modification conflicts among multiple reviewers, optimizing the collaborative review process, improving the review efficiency and the document quality, and intelligently sorting and processing the modification conflicts, the burden of the reviewers is reduced, more fair and accurate modification and adoption are ensured, and the overall process and the output quality of document review are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document processing, and particularly to a method and system for PDF editing and review on a PC side. Background Art

[0002] The technical field of document processing aims to efficiently manage and operate digital documents, covering a wide range of technologies from basic text editing to complex content management and automated workflows. At the core of this technical field, it includes the creation, editing, storage, retrieval, and sharing of documents, taking into account document security, access control, and the ability to integrate with information systems. Document processing technology also focuses on improving processing efficiency and user experience, supporting multiple formats and devices, and ensuring data consistency and integrity.

[0003] Among them, the method for PDF editing and review on the PC side refers to a specific technology for efficiently editing and reviewing PDF files. The technical matters involved in this patent theme include text editing, annotation addition, content modification, and layout adjustment. Specific technical implementation methods include using graphical user interface tools to directly operate on PDF documents, as well as including a track changes and commenting system that enables multiple reviewers to collaborate and record inputs. The means is implemented through a software application designed to run on a personal computer to support users' direct operation and modification of PDF files.

[0004] The prior art has deficiencies in dealing with large-scale document editing and multi-user collaboration. Especially during the review process, the modifications made by multiple users lead to document format chaos and content errors, affecting the reading and further use of the document. Although existing PDF editing tools support basic annotation and change tracking, when dealing with complex document structures and format inheritance issues, the functions of the tools are insufficient to meet the requirements of efficiency and precision. The prior art lacks support for intelligent processing of modification attribution and conflicts, resulting in the need for reviewers to spend a lot of time manually resolving conflicts and errors during the document review process in a multi-user collaboration environment, not only reducing work efficiency but also damaging the accuracy and consistency of the document due to error handling. Summary of the Invention

[0005] To address the deficiencies in the prior art in handling large-scale document editing and multi-user collaboration, especially in the review process where modifications by multiple users lead to document format chaos and content errors, affecting the reading and further use of the document. Although existing PDF editing tools support basic annotation and change tracking, when dealing with complex document structures and format inheritance issues, the functions of the tools are insufficient to meet the requirements of efficiency and precision. The prior art lacks support for intelligent handling of modification attribution and conflicts, resulting in the need for reviewers to spend a large amount of time manually resolving conflicts and errors during the document review process in a multi-user collaboration environment, not only reducing work efficiency but also compromising the accuracy and consistency of the document due to error handling. Embodiments of the present invention provide a PDF editing and review method and system for a PC. The technical solution is as follows:

[0006] On the one hand, a PDF editing and review method for a PC is provided, and the method includes:

[0007] S1: Collect text data in the PDF document on the PC, parse the text block structure, analyze the corresponding positional relationship between the text block and the surrounding text, and establish a text hierarchy feature index with reference to the title level and numbering order;

[0008] S2: Using the text hierarchy feature index, set the format inheritance method, extract the format parameters of text blocks at the same level, identify the inheritance range of paragraph indentation and line spacing, and filter format inheritance conflicts to obtain a format inheritance adjustment result;

[0009] S3: Through the format inheritance adjustment result, extract the PDF review modification records, parse the review annotations, text modifications, and format adjustments, record the start and end positions of the modified text blocks, evaluate the matching degree of the modified content with the original text blocks, and identify the reviewer modification conflict information to obtain a multi-source modification attribution relationship;

[0010] S4: Using the multi-source modification attribution relationship, filter the multi-source modification records of the same text block, classify and adjust according to the modification type, and refer to the reviewer's permissions, modification range, and modification impact degree to obtain a modification conflict priority allocation result;

[0011] S5: According to the modification conflict priority allocation result, record the merged modification blocks and unresolved conflicts, and evaluate the format consistency of the modified text blocks to obtain a review change tracking result.

[0012] As a further solution of the present invention, the text hierarchical feature index includes hierarchical title coding, position index between text blocks, and text sequence numbers associated with levels. The format inheritance adjustment result includes unified font size parameters, color gradient ranges, standardized paragraph indentation and line spacing settings. The multi-source modification attribution relationship includes the reviewer to whom each modification belongs, the type identifier of the modification, and the position interval of the text block. The modification conflict priority assignment result includes the processing priority of each modification type, the permission level of the reviewer for the modification, and the classification strategy for conflict resolution. The reviewer change tracking result includes the merged modified text block records, the list of unresolved conflicts, and the evaluation indicators for format consistency.

[0013] As a further solution of the present invention, the specific steps for obtaining the text hierarchical feature index are as follows:

[0014] S101: Collect the text data in the PDF documents on the PC side, extract the position information, font features, and paragraph structure of the text blocks, identify the text block areas with reference to the page numbers, coordinate ranges, and text formats, and obtain the spatial distribution of the text blocks;

[0015] S102: Based on the spatial distribution of the text blocks, analyze the horizontal spacing, vertical spacing, and paragraph alignment features of adjacent text blocks, record the arrangement of the text blocks, identify the superior-subordinate relationship of the text blocks, judge the logical attribution between the text blocks according to the degree of text content connection, typesetting rules, and page layout information, match the front-back order of the text blocks, and obtain the position correspondence relationship of the text blocks;

[0016] S103: Use the position correspondence relationship of the text blocks to analyze the title level, numbering order, and hierarchical relationship of the text blocks, match the title hierarchy according to the title text length, font size, and format, and generate the text hierarchical feature index.

[0017] As a further solution of the present invention, the specific steps for obtaining the format inheritance adjustment result are as follows:

[0018] S201: Use the text hierarchical feature index to obtain the format parameters of the same-level text blocks, extract the font size, color gradient, paragraph indentation, and line spacing data of the text blocks, analyze the distribution range of the format parameters in the same-level text blocks, identify the change trend of the format parameters, and obtain the distribution of the same-level text format parameters;

[0019] S202: Based on the distribution of the same-level text format parameters, set the format inheritance method, identify the format inheritance range between text blocks, analyze the inheritance mode of the font size, color gradient, paragraph indentation, and line spacing, identify the deviation interval of the format parameters, screen the abnormal situations of the format inheritance of the text blocks, and obtain the format inheritance conflict determination result;

[0020] S203: Adopt the format inheritance conflict determination result, screen out the problems of inconsistent text block formats at the same level, analyze the priority of format adjustment, adjust the format parameters according to the inheritance relationship, correct the conflicts in font size, color gradient, paragraph indentation and line spacing, and generate the format inheritance adjustment result.

[0021] As a further solution of the present invention, the step of obtaining the attribution relationship of multi-source modifications is specifically as follows:

[0022] S301: Utilize the format inheritance adjustment result to extract the review modification records of the text data in the PDF document on the PC side, parse the review annotations, text modifications, deletion operations and format adjustment contents of the text blocks, record the text block numbers and modification timestamps associated with the modification contents, and obtain the review modification data set;

[0023] S302: Based on the review modification data set, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of modified characters and the change of text structure, evaluate the matching degree of the modification content, judge the attribution range of the modified text block, and obtain the modification content matching evaluation result;

[0024] S303: Adopt the modification content matching evaluation result to number the attribution text blocks of the modified text blocks, analyze the original sources of the modified text, combine the text hierarchy features and format parameters, match the attribution relationship, identify independent modified texts and merged modification contents, and obtain the modification attribution text block data;

[0025] S304: Based on the modification attribution text block data, detect the modification situations of multiple reviewers on the same text block, identify the differences in modification time, content and format adjustment, screen out the modification conflict points, analyze the modification priority, and obtain the multi-source modification attribution relationship.

[0026] As a further solution of the present invention, analyze the type of the text modification, the length of modified characters and the change of text structure, and adopt the formula:

[0027]

[0028] where I m represents the text modification intensity index, C new represents the number of characters in the modified text, C orig represents the number of characters in the original text, Δs i represents the change value of the i-th structural modification, and n represents the total number of modification positions.

[0029] As a further solution of the present invention, the step of obtaining the modification conflict priority allocation result is specifically as follows:

[0030] S401: Screen the multi-source modification records of the same text block through the multi-source modification attribution relationship, integrate the modification content of the same text block based on the modified text block number, modification type, and timestamp, analyze the content differences of the text modification, identify the overlapping areas of the modification scope, and obtain the multi-source modification set of the text block;

[0031] S402: Based on the multi-source modification set of the text block, classify and adjust the text block modification records according to the modification type, identify the coverage of deletion, replacement, and format adjustment operations, calculate the coverage index, evaluate the influence scope of the modified text, match the change degree of the text before and after the modification, screen the key modification content, and obtain the classification adjustment result of the modification type;

[0032] S403: Use the classification adjustment result of the modification type to evaluate the priority weight of the modification operation according to the reviewer's permission, modification scope, and modification influence degree, judge the processing order of conflicting modifications, adjust the applicable scope of the modification, and generate the priority allocation result of the modification conflict.

[0033] As a further solution of the present invention, the formula for calculating the coverage index is:

[0034]

[0035] where CA represents the coverage index, M o represents the eigenvalue after the o-th text modification, T o represents the eigenvalue before the o-th text modification, W o represents the weight coefficient of the o-th text, and M represents the number of text segments.

[0036] As a further solution of the present invention, the specific steps for obtaining the review change tracking result are as follows:

[0037] S501: Adopt the priority allocation result of the modification conflict, record the merged modification blocks and unresolved conflicts, extract the modified text block number, modification type, and applicable scope, screen the unresolved conflict records, and analyze the reasons for the unresolved modification conflicts to obtain the merged modification and conflict record data;

[0038] S502: Based on the merged modification and conflict record data, extract the format parameters of the modified text block according to the format inheritance method, evaluate the matching degree of font size, color gradient, paragraph indentation, and line spacing, analyze the consistency between the format adjustment and the inheritance scope, and judge whether the format of the modified text conforms to the format inheritance method to obtain the text format consistency evaluation result;

[0039] S503: Based on the evaluation result of the consistency of the modified text format, combined with the status of the merged modifications and unresolved conflicts, record the format adjustment situation, match the change scope of the modified text block, and generate the review change tracking result.

[0040] On the other hand, the PDF editing review system on the PC side is used to execute the above-mentioned PDF editing review method on the PC side, and the system includes:

[0041] The feature recognition module collects the text data in the PDF document on the PC side, analyzes the positional relationship between the text block and the surrounding text, identifies the title level and numbering order, and establishes a text hierarchy feature index;

[0042] The format inheritance module uses the text hierarchy feature index to set the inheritance rules for the format parameters of the text blocks at the same level in the PDF document on the PC side, analyzes the font size, color gradient, and paragraph indentation, identifies the conflicts in format inheritance, verifies the format consistency, and generates the format inheritance adjustment result;

[0043] The review modification module uses the format inheritance adjustment result to extract the review modification records of the text data in the PDF document on the PC side, identifies the text modification, deletion, and format adjustment information, analyzes the modification conflicts between different reviewers, and generates the multi-source modification attribution relationship;

[0044] The modification conflict classification module sorts out the multi-source modification records based on the multi-source modification attribution relationship, assigns priorities according to the modification type and reviewer permissions, optimizes the overall consistency of the document, and obtains the modification conflict priority assignment result;

[0045] The change recognition module records and merges the conflicting modification blocks according to the modification conflict priority assignment result, evaluates the format consistency of the modified text, and obtains the review change tracking result.

[0046] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:

[0047] By collecting and structurally analyzing the text data in the document, the organization and retrieval efficiency of the text are improved. The detailed analysis of the positional and structural relationships of the text blocks makes the document editing more intuitive and efficient, which is particularly important when processing batch documents. By analyzing the text hierarchy and format inheritance methods, the consistency and professionalism of the document format are ensured, the need for manual format adjustment is reduced, and the automation level of document processing is improved. By identifying and resolving the modification conflicts among multiple reviewers, the collaborative review process is optimized, and the review efficiency and the quality of the document are improved. The intelligent sorting and processing of modification conflicts not only reduce the burden on reviewers, but also ensure more fair and accurate modification adoption, and improve the overall process and output quality of document review. Description of the Drawings

[0048] Figure 1 Schematic diagram of the workflow of the present invention;

[0049] Figure 2 Refined flowchart of S1 of the present invention;

[0050] Figure 3 Refined flowchart of S2 of the present invention;

[0051] Figure 4 Refined flowchart of S3 of the present invention;

[0052] Figure 5 Refined flowchart of S4 of the present invention;

[0053] Figure 6 Refined flowchart of S5 of the present invention;

[0054] Figure 7 System flowchart of the present invention. Detailed implementation manners

[0055] The technical solutions in the present invention will be described below with reference to the accompanying drawings.

[0056] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either of the two can be selected.

[0057] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0058] Please refer to Figure 1 , the embodiments of the present invention provide a PDF editing and review method for the PC side. The processing flow of this method may include the following steps:

[0059] S1: Collect the text data in the PDF document on the PC side, parse the text block structure, analyze the corresponding position relationship between the text block and the surrounding text, and establish a text hierarchy feature index with reference to the title level and numbering order;

[0060] S2: Use the text hierarchy feature index to set the format inheritance method, extract the format parameters of the same-level text blocks, identify the inheritance range of font size, color gradient, paragraph indentation and line spacing, filter format inheritance conflicts, analyze the problem of inconsistent same-level text formats, and obtain the format inheritance adjustment result;

[0061] S3: Adjust the results through format inheritance, extract the PDF review modification records, parse the review notes, text modifications, deletion operations, and format adjustments, record the start and end positions of the modified text blocks, evaluate the matching degree between the modified content and the original text blocks, analyze the text blocks to which the modifications belong, identify the multi-reviewer modification conflict information, and obtain the multi-source modification attribution relationship;

[0062] S4: Adopt the multi-source modification attribution relationship, screen the multi-source modification records of the same text blocks, classify and adjust them according to the modification types, and refer to the reviewer permissions, modification scopes, and modification impact degrees to obtain the modification conflict priority allocation results;

[0063] S5: According to the modification conflict priority allocation results, record the merged modified blocks and the unresolved conflicts, and evaluate the format consistency of the modified text blocks according to the format inheritance method to obtain the review change tracking results;

[0064] The text-level feature index includes hierarchical title encodings, position indexes between text blocks, and text sequence numbers associated with the hierarchy. The format inheritance adjustment results include unified font size parameters, color gradient ranges, and standardized paragraph indentation and line spacing settings. The multi-source modification attribution relationship includes the attributive reviewers of each modification, type identifiers of the modifications, and text block position intervals. The modification conflict priority allocation results include the processing priorities of each modification type, the reviewer permission levels of the modifications, and the classification strategies for conflict resolution. The review change tracking results include the records of the merged modified text blocks, the list of unresolved conflicts, and the evaluation indicators for format consistency.

[0065] Please refer to Figure 2 , and the specific steps for obtaining the text-level feature index are as follows:

[0066] S101: Collect the text data in the PDF documents on the PC side, extract the position information, font features, and paragraph structures of the text blocks, identify the text block areas with reference to the page numbers, coordinate ranges, and text formats, and obtain the spatial distribution of the text blocks;

[0067] The extraction of text data from PDF documents involves multiple steps. It is necessary to obtain the binary data of the PDF file and parse the text content. Text blocks on each page are obtained through text parsing libraries such as pdfplumber or PyMuPDF and classified according to coordinate information. For example, for the standard A4 paper size (210mm×297mm), pixel coordinate mapping can be adopted (corresponding to 595×842 pixels at 72dpi resolution). For different font information, font name, font size, bold, italic and other features are extracted through an OCR engine or a built-in font parsing library. The font size can be calculated through the glyph boundary (boundingbox). For example, if the character height of a certain paragraph of text is 12 pixels, the font size can be inferred as 12pt. All text blocks need to be clustered according to the page number and coordinate range for subsequent spatial distribution analysis. When extracting text blocks, it is necessary to consider whether there is a multi-column layout. The double-column format can determine the dividing line between the two columns by detecting the X coordinate distribution of the text blocks (for example, the X coordinate range of one column of text blocks is 50-250px, and the other column is 300-550px). Finally, for non-pure text data such as tables and text embedded in pictures, table detection algorithms (such as OpenCV edge detection or Tabula's table recognition method) can be combined to obtain the text content and the spatial distribution of the text blocks.

[0068] S102: Based on the spatial distribution of text blocks, analyze the horizontal spacing, vertical spacing and paragraph alignment features of adjacent text blocks, record the arrangement of text blocks, identify the hierarchical relationship of text blocks, judge the logical attribution between text blocks according to the text content connection degree, typesetting rules and page layout information, match the front-back order of text blocks, and obtain the corresponding relationship of text block positions;

[0069] It is necessary to further calculate the horizontal and vertical spacing between adjacent text blocks to determine the logical relationship of the text blocks. If the distance between two text blocks on the X-axis is less than the set paragraph spacing threshold (such as 10px) and the alignment methods are the same, it can be determined that they are part of the same paragraph. For text blocks with obvious indentation (such as the first line indent exceeding 20px), it can be inferred that they are the first lines of paragraphs. For the judgment of paragraph spacing, the minimum line spacing corresponding to different font sizes can be set. The line spacing of text with a font size of 12pt is about 14 - 16px. If the spacing between two text blocks exceeds 1.5 times of this value (such as 24px), it can be inferred that they are in different paragraphs. When identifying the hierarchical relationship of text blocks, it is necessary to combine text formats. For example, headings have a larger font size (such as 14pt or above) and are in bold or centered alignment formats. By matching features, the hierarchical relationship of the text can be established. Taking an example, if there are four heading text blocks on a certain PDF page with font sizes of 16pt, 14pt, 12pt, and 10pt respectively, a four-level heading structure can be constructed, and according to the order of text content connection, match the previous and subsequent text blocks to obtain the corresponding relationship of text block positions.

[0070] S103: Use the corresponding relationship of text block positions to analyze the heading level, numbering order, and hierarchical relationship of text blocks. According to the number of words, font size, and format of the heading text, match the heading hierarchy structure to generate a text hierarchical feature index;

[0071] When analyzing the heading level of text blocks, it is necessary to analyze according to information such as their font size, numbering format, alignment method, etc. For headings numbered with Arabic numerals (such as "1.1"), the hierarchical relationship can be matched according to the numbering format. Common rules include that first-level headings use a larger font size (such as 16pt), second-level headings have a slightly smaller font size (such as 14pt), and third-level headings continue to decrease (such as 12pt). For the automatic matching of heading levels, regular expressions can be used to extract the heading numbers. When extracting the hierarchical numbers, combine font features to determine the hierarchical structure. By constructing a text hierarchical feature index, according to the corresponding relationship between headings and the main text, a structured index similar to a table of contents can be formed. Suppose a certain document contains chapters "1. Data Processing", "1.1 Data Acquisition", "1.1.1 Sensor Data Acquisition", which can be parsed into a tree-like hierarchy and stored as structured data (such as JSON format). At the same time, for the hierarchical division of the main text content, based on the alignment method and indentation of the subsequent text blocks of the heading, if the first line of the main text block indents more than the set threshold (such as 20px), it can be inferred that it is the start of a new paragraph, otherwise it is a continuation of the current paragraph. Complete the parsing of the text structure and generate a text hierarchical feature index.

[0072] Please refer to Figure 3 , and the specific steps for obtaining the format inheritance adjustment result are as follows:

[0073] S201: Use text - level feature indexing to obtain the format parameters of text blocks at the same level, extract the font size, color gradient, paragraph indentation, and line - spacing data of the text blocks, analyze the distribution range of the format parameters in the text blocks at the same level, identify the change trend of the format parameters, and obtain the distribution of the format parameters of the text at the same level.

[0074] When obtaining the format parameters of text blocks at the same level, it is necessary to extract the font size, color gradient, paragraph indentation, and line - spacing data of the text blocks, and analyze the distribution range of the format parameters in the text blocks at the same level. Traverse all text blocks at the same level to obtain the font information of each text block, including font size, font type, styles such as bold or italic. By extracting the coordinate information of the text, calculate the vertical and horizontal spacing between text blocks to identify the line - spacing and paragraph indentation situations. The calculation of the color gradient requires detecting the color information of the text and analyzing its distribution in the RGB or CMYK color models. If the font size of the text blocks at the same level is mainly concentrated between 12pt and 14pt, and the font size of a certain text block is 16pt, it is necessary to record the abnormal situation of this format parameter and analyze the change trend of the format parameters. If the format of the text blocks on a certain page gradually increases or the indentation amplitude changes significantly, it means that there is a hierarchical adjustment or inconsistent typesetting inside the document, and it is necessary to further check whether it conforms to the overall format specification. Based on the statistical results of the format parameters of all text blocks, generate the distribution of the format parameters of the text at the same level.

[0075] S202: Based on the distribution of the format parameters of the text at the same level, set the format inheritance method, identify the format inheritance range between text blocks, analyze the inheritance modes of the font size, color gradient, paragraph indentation, and line - spacing, identify the deviation range of the format parameters, screen the abnormal situations of the format inheritance of the text blocks, and obtain the result of the format inheritance conflict determination.

[0076] When setting the format inheritance method, it is necessary to define the inheritance range of different format parameters based on the distribution of the same-level text format parameters, analyze the inheritance patterns of font size, color gradient, paragraph indentation, and line spacing, and set the reference range for format inheritance. If the font size of the same-level text blocks is mainly between 12pt and 14pt, this range can be set as the default inheritance range. If the font size of a certain text block is 10pt or 16pt, it belongs to a format deviation and it is necessary to further check whether it conforms to the inheritance rules. The inheritance of the color gradient requires analyzing the color change of the text. For example, if the color of the same-level text blocks is mainly dark gray (RGB50, 50, 50), and a certain text block is pure black (RGB0, 0, 0), it will cause inconsistent typesetting and it is necessary to check whether it conforms to the inheritance pattern. The analysis of paragraph indentation and line spacing is mainly used to identify the text alignment method. For example, if the indentation of all text blocks is 20px, and the indentation of a certain text block is 40px, it is a format exception. It is necessary to record this deviation and screen the inheritance exceptions of all format parameters. By comparing the format inheritance relationships of all text blocks, the format inheritance conflict determination result can be obtained.

[0077] S203: Use the format inheritance conflict determination result to screen the problems of inconsistent formats of the same-level text blocks, analyze the priority of format adjustment, adjust the format parameters according to the inheritance relationship, correct the conflicts of font size, color gradient, paragraph indentation, and line spacing, and generate the format inheritance adjustment result;

[0078] When screening the problems of inconsistent formats of the same-level text blocks, it is necessary to analyze the priority of format adjustment and adjust the format parameters according to the inheritance relationship to correct the conflicts of font size, color gradient, paragraph indentation, and line spacing. Set the adjustment priority according to the influence range of the format conflict. For example, the adjustment priority of the font size is higher than that of the color gradient, the adjustment priority of the color gradient is higher than that of the indentation adjustment, and the adjustment priority of the line spacing is the lowest. If the font size of a certain text block exceeds the inheritance range, it is necessary to adjust its font size to the default range. For example, adjust 16pt to 14pt to match the same-level text blocks. The adjustment of the color gradient can use the interpolation method to make it gradually transition to be consistent with the surrounding text. If the color of a certain text block is pure black and the surrounding text is dark gray, its color can be adjusted to an intermediate color level (such as RGB25, 25, 25). The indentation adjustment mainly involves correcting the alignment method. If the indentation of a certain text block is significantly too large, it is necessary to adjust its indentation value to the standard range, such as adjusting 40px to 20px. Through the adjustment of all format parameters and ensuring that the formats of all text blocks conform to the inheritance rules, the typesetting consistency of the document can be improved and the format inheritance adjustment result can be generated.

[0079] Please refer to Figure 4 , the specific steps for obtaining the attribution relationship of multi-source modifications are as follows:

[0080] S301: Adjust the results using format inheritance, extract the review and modification records of the text data in the PDF document on the PC side, parse the review comments, text modifications, deletion operations, and format adjustment content of the text blocks, record the text block numbers and modification timestamps associated with the modification content, and obtain the review and modification dataset;

[0081] The extraction of the review and modification records of the PDF document on the PC side involves multiple steps. It is necessary to parse the annotation layer in the PDF file to obtain all the metadata of the review and modification. For example, by using PyMuPDF to parse the PDF annotations, information such as the text content, modification type (such as addition, deletion, replacement), reviewer name, and timestamp of each modification record can be extracted. At the same time, the numbers of the text blocks are parsed to match the corresponding relationship between the modification content and the original text. If the number of a certain text block is T123, and its modification content involves deleting the original text "data processing" and replacing it with "data optimization", then record the modification content ("data optimization"), the original text ("data processing"), the modification time (such as 2025-02-20 10:30:15), the modification type (replacement), and the modifier ID. For the record of the deletion operation, the text before deletion can be extracted, and the deletion status can be marked in the review and modification dataset. For the format adjustment content, such as font size change, color change, etc., by comparing the parameters of the original format and the modified format, such as the font size is adjusted from 12pt to 14pt, then record the format adjustment content, and store all the associated information of the modifications to form the review and modification dataset.

[0082] S302: Based on the review and modification dataset, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of the modified characters, and the change in the text structure, evaluate the matching degree of the modification content, determine the attribution range of the modified text block, and obtain the matching evaluation result of the modification content;

[0083] Analyze the type of text modification, the length of the modified characters, and the change in the text structure, using the formula:

[0084]

[0085] where, I m represents the text modification intensity index, C new represents the number of characters in the modified text, C orig represents the number of characters in the original text, Δs i represents the change value of the i-th structural modification, and n represents the total number of modification positions;

[0086] Parameter meaning:

[0087] C new is the number of characters in the modified text, obtained by counting the total number of characters in the modified text;

[0088] Corig is the number of characters in the original text, obtained by counting all the characters in the original text;

[0089] Δs i is the modification value of the i-th text structure, representing the measure of the difference in text structure before and after modification, calculated by comparing the text structure features before and after modification (such as the number of paragraphs, sentence length, etc.);

[0090] n is the total number of modification positions in the text, determined by detecting the number of all modification points in the text;

[0091] Calculation and derivation process:

[0092] Suppose the original text contains 1000 characters, the modified text contains 1050 characters, there are 5 structural modifications, and the structure modification values for each modification (for example, the sentence changes from 10 words to 12 words) are 2, 1, 0, 3, 4 respectively;

[0093] Calculate the change in the number of characters:

[0094] |C new -C orig | = |1050 - 1000| = 50;

[0095] Calculate the sum of squares of the structure changes:

[0096]

[0097] Calculate the square root of the sum of squares of the structure changes:

[0098]

[0099] Substitute the values into the formula to calculate the text modification intensity index:

[0100]

[0101] The result shows that the text modification intensity index is 0.05548, reflecting the degree of change from the original text to the modified text. This index reveals the overall intensity of text modification, helps to evaluate the depth and scope of text modification, and is an important measure for evaluating the similarity between the modified text and the original text and the structural changes.

[0102] S303: Use the evaluation result of the modified content matching to number the attribution text blocks of the modified text block, analyze the original source of the modified text, combine the text hierarchy features and format parameters, match the attribution relationship, identify the independent modified text and the merged modified content, and obtain the data of the modified attribution text block;

[0103] It is necessary to number the attribution relationship of the modified text block, match its original source, extract the associated text blocks of the modified text. If the modified text belongs to T101, update the modification record of T101 in the dataset. At the same time, combine the text hierarchy features and format parameters to analyze the format inheritance situation of the modification. For example, if the font size of a modified text block is 14pt while the font size of its attributed text block is 12pt, record the format deviation ΔF = 2pt for subsequent adjustment. For merged modification content, if multiple modified text blocks are adjacent and there is continuity in the modification record, they are merged into one modification block. For example, if the modification contents of T201 and T202 are "Optimize model parameters" and "Improve calculation accuracy" respectively, and the modification time interval is less than the set threshold (such as 5 minutes), they are merged into one modification record "Optimize model parameters and improve calculation accuracy". For independent modified text, if a certain modification content fails to match the attributed text block, it is marked as an independent modification, and the data of the attributed text block of the modification is obtained.

[0104] S304: Based on the data of the attributed text block of the modification, detect the modification situations of multiple reviewers on the same text block, identify the deviations in modification time, content differences, and format adjustments, screen the modification conflict points, analyze the modification priorities, and obtain the multi-source modification attribution relationship;

[0105] In the case of multi-reviewer modifications, it is necessary to detect the modification differences on the same text block, extract all modification records, and group them according to the text block number. For the text block numbered T305, extract the modification contents of all reviewers A, B, and C, and calculate the modification time interval. For example, if A modifies at 10:30, B modifies at 10:32, and C modifies at 10:35, record the time sequence, analyze the content differences, and calculate the text similarity after modification by different reviewers. For example, if A modifies to "Improve calculation efficiency", B modifies to "Optimize calculation method", and C modifies to "Adjust calculation formula", then calculate the A-B similarity of 0.75, the B-C similarity of 0.6, and the A-C similarity of 0.55. If the modification difference is greater than the set threshold (such as 0.7), it is marked as a conflict, sorted according to the modification priorities. If A is the main editor and its modification priority is higher than B and C, then the modification of A is preferred, and the multi-source modification attribution relationship is stored.

[0106] Please refer to Figure 5 , and the specific steps for obtaining the modification conflict priority allocation result are as follows:

[0107] S401: Through the multi-source modification attribution relationship, screen the multi-source modification records of the same text block, integrate the modification contents of the same text block according to the text block number, modification type, and timestamp, analyze the content difference situation of the text modification, and identify the overlapping area of the modification range to obtain the multi-source modification set of the text block;

[0108] When screening multi-source modification records of the same text block, it is necessary to integrate the modification contents of different reviewers and classify them according to the text block number, modification type, and timestamp. Filter all relevant modification records by the text block number (such as T501) and classify them according to the modification type. If there are three modification records A, B, and C for the text block numbered T501, which involve deletion (A), replacement (B), and format adjustment (C) respectively, then store the modification contents in categories. Next, by analyzing the content differences of the text modifications, calculate the similarity of the modification contents, using the formula:

[0109]

[0110] If the modification content of A is "Optimize the calculation method" and the modification content of B is "Improve the calculation efficiency", and the number of intersection characters is 6 and the total number of characters is 12, then calculate the similarity:

[0111]

[0112] If the similarity is lower than the set threshold (such as 0.6), it is determined that the content difference is large, and further analyze the overlapping area of the modification scope. If the modification content of A involves the text "Optimization of the calculation model" and the modification of B involves "Adjustment of the calculation model", then calculate the proportion of the overlapping area between the two. If the overlapping rate of the "calculation model" part is 80%, then identify this part as the overlapping area and form a multi-source modification set of text blocks.

[0113] S402: Based on the multi-source modification set of text blocks, classify and adjust the text block modification records according to the modification type, identify the coverage of deletion, replacement, and format adjustment operations, calculate the coverage index, evaluate the influence range of the modified text, match the change degree of the text before and after the modification, screen the key modification contents, and obtain the classification adjustment result of the modification type;

[0114] The formula for calculating the coverage index is:

[0115]

[0116] Among them, CA represents the coverage index, M o represents the eigenvalue after the o-th text modification, T o represents the eigenvalue before the o-th text modification, W o represents the weight coefficient of the o-th text, and M represents the number of text segments;

[0117] Parameter meaning:

[0118] Text eigenvalue M o and T o Calculation:

[0119] Based on character edit distance: Calculate the edit distance of the corresponding segments before and after text modification. Measure the character-level changes through the Levenshtein distance and normalize it to the range of 0-1 to enable comparison of texts of different scales;

[0120] Based on word vector similarity: Use word embeddings (Word2Vec, BERT) to calculate the cosine similarity between sentences and take the similarity value as the text feature value;

[0121] Based on syntactic structure analysis: Utilize dependency syntactic analysis to calculate the degree of sentence structure change and assign values in combination with the syntactic tree transformation rate;

[0122] Calculation of weight coefficient W o :

[0123] Word frequency weight: Calculate the importance of each text segment based on TFIDF. If the text modification involves high-weight words, assign a high W o value;

[0124] Sentence position weight: The weights of important paragraphs such as the title, the first sentence, and the ending of the text are high. Normalize it to the range of 0-1 with the position factor and use it as the correction factor for the weight coefficient;

[0125] Modification type weight: Operations such as deletion, replacement, and format adjustment of the text have different impacts. To ensure that the contribution degrees of different types of modifications match, set the basic weight values for different modification types;

[0126] Specific example:

[0127] Given the data before and after the modification of a certain text segment as follows: The number of text segments M = 5;

[0128] Feature value T before modification o : T 1 = 0.8, T 2 = 0.6, T 3 = 0.7, T 4 = 0.5, T 5 = 0.9;

[0129] Feature value M after modification o : M 1 = 0.3, M 2 = 0.9, M 3 = 0.6, M 4 = 0.4, M 5 = 0.8;

[0130] Weight coefficient W o : W 1 = 0.5, W 2 = 0.7, W 3 = 0.8, W 4= 0.6, W 5 = 0.9;

[0131] Calculate |M o - T o |:

[0132] |M 1 - T 1 | = |0.3 - 0.8| = 0.5;

[0133] |M 2 - T 2 | = |0.9 - 0.6| = 0.3;

[0134] |M 3 - T 3 | = |0.6 - 0.7| = 0.1;

[0135] |M 4 - T 4 | = |0.4 - 0.5| = 0.1;

[0136] |M 5 - T 5 | = |0.8 - 0.9| = 0.1;

[0137] Calculate the numerator part:

[0138]

[0139] Calculate the denominator part:

[0140]

[0141] Substitute into the formula for calculation:

[0142]

[0143] The result shows that the coverage degree index of the text modification is 0.321, indicating that there are certain differences between the text before and after the modification, but the overall change degree is relatively low. This value can be used as a basis for adjustment in the subsequent text classification and adjustment process.

[0144] S403: Utilize the classification and adjustment results of the modification type, and according to the reviewer's authority, modification scope, and modification impact degree, evaluate the priority weight of the modification operation, determine the processing order of conflicting modifications, adjust the applicable scope of the modification, and generate the priority allocation result of modification conflicts;

[0145] It is necessary to evaluate the priority weights of modification operations to determine the order of handling modification conflicts. Weights are assigned according to the reviewer's permissions. Set the weight of the chief editor's modification \(W = 1.0\), the weight of the ordinary editor's modification \(W = 0.8\), and the weight of the external reviewer's modification \(W = 0.5\). For multiple modifications to the same text block, calculate the weighted modification impact degree \(S\). W As follows:

[0146]

[0147] Among them, \(I\) i represents the influence range of the \(i\)-th modification, and \(n\) is the number of modifications;

[0148] Set the influence ranges of the modifications of three reviewers to be 20, 30, and 50 characters respectively, and the corresponding weights to be 1.0, 0.8, and 0.5. Then the calculation is as follows:

[0149] \(S\) W \(=(1.0×20)+(0.8×30)+(0.5×50)=64\);

[0150] If the \(S\) value of a certain modification W is higher than 20% of the modification, then this modification is preferentially retained. Analyze the modification range. If a certain modification involves the entire text block, while another modification only affects a single word, then the modification with a larger coverage range is preferentially adopted. Considering factors such as the modification influence range, permissions, and timestamps, adjust the applicable range of the modification to generate the priority allocation result of modification conflicts.

[0151] Please refer to Figure 6 , and the specific steps for obtaining the review change tracking results are as follows:

[0152] S501: Adopt the priority allocation result of modification conflicts, record the merged modification blocks and unresolved conflicts, extract the modification text block numbers, modification types, and applicable ranges, screen the unresolved conflict records, analyze the reasons for the unresolved modification conflicts, and obtain the merged modification and conflict record data;

[0153] Screen all text blocks involving modifications and organize them according to the text block number, modification type, and scope of application. For merged modification blocks, it is necessary to extract their modification sources, scope of application, and adopted modification versions. A certain text block T701 has undergone three rounds of modifications. In the first round, some content was deleted. In the second round, some text was replaced. In the third round, the format was adjusted. The text is mainly based on the content modified in the second round, with the format adjustment retained, while the deletion operation was not adopted. Therefore, it is necessary to mark in the merged record that T701 adopted the content modified in the second round and attach the format adjustment information. For unresolved conflict records, it is necessary to identify the reasons for their non-resolution. If there are modification contents from multiple reviewers for a certain text block T702, but there are contradictions in the format, such as one reviewer adjusted the indentation and another reviewer changed the font size, and the two are incompatible, then this modification conflict cannot be automatically merged, and it is necessary to additionally mark this text block as "format conflict pending processing", and it is necessary to further analyze the scope of influence of the conflict. For example, if a certain text block only involves local phrase modifications, while another text block involves the adjustment of the entire paragraph content, then the former conflict is smaller and can wait for manual intervention, while the latter affects the layout of the entire document and needs to be processed first. All merged modification blocks and unresolved conflicts need to be stored in the record for subsequent format matching and modification tracking to obtain the merged modification and conflict record data.

[0154] S502: Based on the merged modification and conflict record data, extract the format parameters of the modified text blocks according to the format inheritance method, evaluate the matching degree of font size, color gradient, paragraph indentation, and line spacing, analyze the consistency between the format adjustment and the inheritance scope, determine whether the format of the modified text conforms to the format inheritance method, and obtain the text format consistency evaluation result;

[0155] Extract the formatting parameters of all modified text blocks, including font size, color gradient, paragraph indentation, and line spacing, and compare them with the format inheritance method. If a text block T801, after modification, has its original font size changed from 12pt to 14pt, while the default format of the same-level text blocks is 12pt, it is necessary to determine whether this modification conforms to the format inheritance method. If 14pt is within the inheritance range, this modification conforms to format inheritance; otherwise, it is necessary to adjust back to the default format or redefine the inheritance range. The evaluation of the color gradient requires detecting the color change of the text after modification. If the original text color is dark gray and it changes to black after modification, it is necessary to determine whether the color gradient change is within the acceptable range. If the change is too large, resulting in inconsistent overall document layout, it is necessary to adjust back to the original color. The matching of the indentation parameter mainly targets the first-line indentation and the overall alignment method. If a text block originally uses the left-alignment format but changes to the center-alignment format after modification, it is necessary to check whether this format conforms to the alignment requirements of the current text level. If not, it needs to be adjusted. The evaluation of the line-spacing parameter is mainly used to check the change in text density. If the original text block uses 1.5 times line spacing and changes to single line spacing after modification, which affects readability, it is necessary to check whether it conforms to the format inheritance standard. By comparing the format parameters of the text block before and after modification, determine that some modifications conform to the inheritance rules and some need to be corrected, and generate the text format consistency evaluation result.

[0156] S503: By modifying the text format consistency evaluation result, combining the status of the merged modifications and the unresolved conflicts, record the format adjustment situation, match the change range of the modified text block, and generate the review change tracking result;

[0157] It is necessary to record the format adjustment situation in combination with the status of merged and modified and unresolved conflicts, match the change range of the modified text block. For all merged modified text blocks, store the adopted modification content and record its format adjustment status. If a text block T901 has been modified multiple times and the content adopted is from the second round of modification, and the format is adjusted to 1.2 times line spacing and 14pt font size, it needs to be stored in the modification record and mark the format adjustment as "confirmed". Secondly, for unresolved conflict text blocks, it is necessary to mark their current status. For example, if T902 is still in the state of unresolved format conflict, mark it as "to be processed" and record its specific problem points, such as "indentation mismatch" or "excessive color gradient change". It is necessary to match the change range of the modified text block to determine the impact of the modification on the overall document. For example, if a modification only involves local phrase adjustment, the impact is small. However, if the modification involves the entire paragraph and affects the surrounding text format, it is necessary to record the range of this modification and evaluate whether additional format adjustment of the text block is required to ensure overall consistency. The status, format adjustment situation and unresolved conflict records of all modified text blocks need to be integrated into the review change tracking data for subsequent format correction and document typesetting to generate the review change tracking result.

[0158] As Figure 7 shown, a PDF editing review system for the PC side, the system includes:

[0159] The feature recognition module collects the text data in the PDF document on the PC side, analyzes the positional relationship between the text block and the surrounding text, identifies the title level and numbering order, and establishes a text hierarchy feature index;

[0160] The format inheritance module uses the text hierarchy feature index to set inheritance rules for the format parameters of text blocks at the same level in the PDF document on the PC side, analyzes the font size, color gradient and paragraph indentation, identifies the conflicts in format inheritance, verifies the format consistency, and generates a format inheritance adjustment result;

[0161] The review modification module uses the format inheritance adjustment result to extract the review modification record of the text data in the PDF document on the PC side, identifies the text modification, deletion and format adjustment information, analyzes the modification conflicts between different reviewers, and generates a multi-source modification attribution relationship;

[0162] The modification conflict classification module classifies and organizes the multi-source modification records based on the multi-source modification attribution relationship, assigns priorities according to the modification type and reviewer permissions, optimizes the overall consistency of the document, and obtains the modification conflict priority assignment result;

[0163] The change recognition module records and merges the conflicting modification blocks according to the modification conflict priority assignment result, evaluates the format consistency of the modified text, and obtains the review change tracking result.

[0164] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A PDF editing and reviewing method on a PC, characterized in that: The following steps are involved: S1: Collect text data from PDF documents on the PC, parse the text block structure, analyze the corresponding position relationship between the text block and the previous and following texts, and establish a text level feature index by referring to the title level and numbering sequence; S2: using the text-level feature index, setting the format inheritance mode, extracting the format parameters of the text blocks at the same level, identifying the inheritance range of paragraph indentation and line spacing, screening format inheritance conflicts, and obtaining the format inheritance adjustment result; S3: extracting PDF review modification records through the format inheritance adjustment results, parsing review comments, text modifications and format adjustments, recording the start and end positions of the modified text blocks, evaluating the matching degree between the modified content and the original text blocks, identifying the reviewer modification conflict information, and obtaining the multi-source modification attribution relationship; S4: using the multi-source modification attribution relationship, screening the multi-source modification records of the same text block, classifying and adjusting according to the modification type, and obtaining the modification conflict priority allocation result by referring to the reviewer authority, modification scope and modification impact degree; S5: According to the modification conflict priority allocation result, the merged modification blocks and unresolved conflicts are recorded, the format consistency of the modification text blocks is evaluated, and the review change tracking result is obtained.

2. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The text level feature index includes hierarchical title codes, position indexes between text blocks, and hierarchically associated text serial numbers; the format inheritance adjustment results include unified font size parameters, color gradient ranges, and standardized paragraph indentation and line spacing settings; the multi-source modification attribution relationship includes the reviewer to whom each modification is attributed, the type identifier of the modification, and the text block position interval; the modification conflict priority allocation result includes the processing priority of each modification type, the reviewer authority level of the modification, and the classification strategy for conflict resolution; the review change tracking result includes the merged modified text block record, a list of unresolved conflicts, and format consistency evaluation indicators.

3. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The steps for obtaining the text-level feature index are specifically as follows: S101: Collect text data in the PDF document on the PC, extract location information, font features and paragraph structure of text blocks, identify text block areas with reference to page numbers, coordinate ranges and text formats, and obtain spatial distribution of text blocks; S102: Based on the spatial distribution of the text blocks, analyze the horizontal spacing, vertical spacing and paragraph alignment features of adjacent text blocks, record the arrangement of the text blocks, identify the superior-subordinate relationship of the text blocks, determine the logical belonging of the text blocks according to the degree of cohesion of the text content, typesetting rules and page layout information, match the front-back order of the text blocks, and obtain the corresponding relationship between the text block positions; S103: using the text block position correspondence, parsing the title level, numbering sequence and hierarchical relationship of the text block, matching the title hierarchy according to the number of words, font size and format of the title text, and generating a text hierarchical feature index.

4. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The steps for obtaining the format inheritance adjustment result are specifically as follows: S201: using the text level feature index, obtaining format parameters of text blocks of the same level, extracting font size, color gradient, paragraph indentation and line spacing data of the text blocks, analyzing the distribution range of format parameters in text blocks of the same level, identifying the change trend of format parameters, and obtaining the distribution of format parameters of texts of the same level; S202: Based on the format parameter distribution of the text at the same level, a format inheritance mode is set, the format inheritance range between text blocks is identified, the inheritance mode of font size, color gradient, paragraph indentation and line spacing is analyzed, the deviation interval of the format parameters is identified, the abnormal situation of the format inheritance of the text block is screened, and the format inheritance conflict determination result is obtained; S203: Using the format inheritance conflict determination result, screening the format mismatch problems of the same-level text blocks, analyzing the priority of format adjustment, adjusting the format parameters according to the inheritance relationship, correcting the conflicts of font size, color gradient, paragraph indentation and line spacing, and generating a format inheritance adjustment result.

5. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The steps for obtaining the multi-source modification attribution relationship are specifically as follows: S301: extracting the review modification records of the text data in the PDF document on the PC by using the format inheritance adjustment result, parsing the review comments, text modification, deletion operations and format adjustment contents of the text block, recording the text block number and modification timestamp associated with the modification contents, and obtaining the review modification data set; S302: Based on the manuscript review modification data set, identify the content similarity between the modified text and the original text block, analyze the type of text modification, the length of the modified characters and the text structure change, evaluate the matching degree of the modified content, determine the belonging scope of the modified text block, and obtain the modified content matching evaluation result; S303: using the modified content matching evaluation result, numbering the text blocks to which the modified text blocks belong, analyzing the original source of the modified texts, combining text level features and format parameters, matching the attribution relationships, identifying the independent modified texts and merging the modified contents, and obtaining the modified attribution text block data; S304: Based on the modification attribution text block data, detect the modification status of the same text block by multiple reviewers, identify the modification time, content difference and format adjustment deviation, screen the modification conflict points, analyze the modification priority, and obtain the multi-source modification attribution relationship.

6. The method for editing and reviewing PDF manuscripts on a PC according to claim 5, characterized in that: Analyze the type of text modification, the length of modified characters and the change in text structure, using the formula: Among them, I m represents the text modification intensity index, C new Represents the number of characters in the modified text, C orig Represents the number of characters in the original text, Δs i represents the change value of the structural modification at the i-th location, and n represents the total number of modified locations.

7. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The steps for obtaining the result of modifying the conflict priority allocation are specifically as follows: S401: Filter the multi-source modification records of the same text block through the multi-source modification attribution relationship, integrate the modification contents of the same text block according to the modified text block number, modification type and timestamp, analyze the content differences of the text modifications, identify the overlapping areas of the modification ranges, and obtain the multi-source modification set of the text block; S402: Based on the multi-source modification set of the text block, the modification record of the text block is adjusted according to the modification type classification, the coverage degree of the deletion, replacement, and format adjustment operations is identified, the coverage degree index is calculated, the impact scope of the modified text is evaluated, the change degree of the text before and after the modification is matched, the key modification content is screened, and the modification type classification adjustment result is obtained; S403: Using the modification type classification adjustment result, the priority weight of the modification operation is evaluated according to the reviewer authority, modification scope and modification impact, the processing order of the conflicting modifications is determined, the modification application scope is adjusted, and the modification conflict priority allocation result is generated.

8. The method for editing and reviewing PDF manuscripts on a PC according to claim 7, characterized in that: The formula for calculating the coverage index is: Among them, CA represents the coverage index, M o represents the feature value after the text at position o is modified, T o represents the feature value before the text at position o is modified, W o represents the text weight coefficient at the oth position, and M represents the number of text fragments.

9. The method for editing and reviewing PDF manuscripts on a PC according to claim 1, characterized in that: The specific steps for obtaining the review change tracking results are as follows: S501: using the modification conflict priority allocation result, recording the merged modification blocks and unresolved conflicts, extracting the modification text block number, modification type and applicable scope, screening the unprocessed conflict records, analyzing the unresolved reasons for the modification conflicts, and obtaining the merged modification and conflict record data; S502: Based on the merged modification and conflict record data, according to the format inheritance mode, extract the format parameters of the modified text block, evaluate the matching degree of font size, color gradient, paragraph indentation and line spacing, analyze the consistency of format adjustment and inheritance range, determine whether the format of the modified text conforms to the format inheritance mode, and obtain the text format consistency evaluation result; S503: Based on the modified text format consistency assessment result, combined with the status of merged modifications and unresolved conflicts, the format adjustment status is recorded, the change scope of the modified text block is matched, and the review change tracking result is generated.

10. A PC-based PDF editing and review system, characterized in that: According to the method for editing and reviewing PDF manuscripts on a PC according to any one of claims 1 to 9, the system comprises: The feature recognition module collects text data from the PDF document on the PC, analyzes the positional relationship between the text block and the surrounding text, identifies the heading level and numbering sequence, and establishes a text-level feature index; The format inheritance module uses the text level feature index to set inheritance rules for format parameters of text blocks of the same level in the PDF document on the PC, analyzes font size, color gradient and paragraph indentation, identifies conflicts in format inheritance, verifies format consistency, and generates format inheritance adjustment results; The review and modification module uses the format inheritance adjustment result to extract the review and modification records of the text data in the PDF document on the PC side, identify the text modification, deletion and format adjustment information, analyze the modification conflicts between different reviewers, and generate multi-source modification attribution relationships; The modification conflict classification module classifies and sorts the multi-source modification records based on the multi-source modification attribution relationship, assigns priorities according to the modification type and reviewer authority, optimizes the overall consistency of the document, and obtains the modification conflict priority assignment result; The change identification module records and merges the conflicting modification blocks according to the modification conflict priority allocation result, evaluates the format consistency of the modified text, and obtains the review change tracking result.

Citation Information

Patent Citations

  • Method for editing PDF (Portable Document Format) document and terminal equipment

    CN107977346A

  • Text format auditing module for financial long text rechecking system

    CN114691919A

  • Parallel editing method for splitting Word document into multiple rich texts

    CN117252164A

  • Content extraction-based contract intelligent auditing method and device, equipment and medium

    CN119249109A

  • Web-based collaborative document review system

    US20130283147A1