Document checking method and system suitable for large language model

By integrating the document dynamic segmentation module and the parsing module into the large language model, the problems of low recognition rate and poor accuracy in document review are solved, and efficient and accurate recognition and review of document format information are achieved.

CN120832879APending Publication Date: 2025-10-24GUANGXI ZHUANG AUTONOMOUS REGION WATER CONSERVANCY & ELECTRIC POWER SURVEY DESIGN & RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510804271.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies have low recognition rates and poor accuracy in document review, and cannot effectively identify semantically related format errors and typos. They also require high coding skills from users, and the recognition rate of multimodal large models is low and cannot accurately identify information such as document font size and spacing.

Method used

The interface between the document dynamic segmentation module and the document parsing module is set up using MCP technology. Combined with AI agent technology, it is integrated into a large language model. The document is segmented and parsed through dynamic segmentation strategy and maximum utilization strategy to generate proofread draft.

Benefits of technology

It improves the efficiency and accuracy of document segmentation and parsing, reduces the professional requirements of users, and enables accurate identification and verification of document format information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832879A_ABST
    Figure CN120832879A_ABST
Patent Text Reader

Abstract

The invention discloses a document checking method and system suitable for a large language model, and belongs to the field of artificial intelligence technology and natural language process.The method comprises the steps that an interface of a document dynamic segmentation module and an interface of a document analysis module are set based on the MCP technology; integrating the document dynamic segmentation module and the document analysis module into a large language model based on an AI agent technology; and obtaining the analysis intention and the to-be-processed document, and processing based on the integrated large language model to output a check draft. According to the invention, the interfaces of the document dynamic segmentation module and the document analysis module are set based on the MCP technology, so that the reasonable segmentation and analysis functions of the document can be realized; based on the AI agent technology, the document dynamic segmentation module and the document analysis module are integrated into a large language model, the segmentation and analysis efficiency and accuracy can be improved, and the professional requirement of a user is lowered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence technology and natural language processing, in particular to a document proofreading method and system suitable for large language models. BACKGROUND

[0002] In modern office scenarios, the standardization of various documents (including academic papers, contract files, technical reports, etc.) has become a rigid requirement. Currently, traditional solutions rely on manually preset regular expression matching rules, which have problems such as inability to effectively identify semantic association format errors, inability to identify misspelled words, exponential growth of maintenance costs with format rule complexity, and certain requirements for user coding techniques.

[0003] With the breakthroughs in LLM (such as PatSnap 2.5) and ai applications (such as workflow, agent, MCP), it has become a feasible direction to use ai technology for document proofreading. The underlying of a word document is implemented by xml, containing information such as text, style, indentation, etc. However, existing LLM models only understand the text of the document when transmitting the file, and cannot directly understand the document style and other format information. Moreover, there are problems such as high hardware cost of long context models. SUMMARY

[0004] The present application proposes a document proofreading method and system suitable for large language models to solve the problems of low recognition rate and poor accuracy in existing technologies.

[0005] To achieve the above purpose, the technical solution adopted by the present application is:

[0006] The document proofreading method suitable for large language models comprises:

[0007] Setting the interface of the document dynamic segmentation module and the document parsing module based on MCP technology;

[0008] Integrating the document dynamic segmentation module and the document parsing module into a large language model based on AI agent technology;

[0009] Obtaining the parsing intent and the document to be processed, and processing based on the integrated large language model to output the proofreading manuscript.

[0010] Further, the document dynamic segmentation module is configured to perform document segmentation based on a dynamic segmentation strategy and a maximum utilization strategy. The dynamic segmentation strategy comprises dynamically selecting an optimal segmentation strategy according to a preset context window and a document type, and the segmentation strategy comprises semantic integrity priority, key information protection, cross-page paragraph processing, and table paragraph processing. The maximum utilization strategy comprises performing document segmentation close to an upper limit of a context length under the condition that document information loss is less than a specified threshold according to a context limit and document content, so as to reduce the number of subsequent calls to an LLM model and maximize the utilization of model performance.

[0011] Further, the document dynamic segmentation module is configured to extract an element stream of a document and mark the element stream, form a buffer area through a temporary storage stack, and store the element stream. The buffer area is initialized according to a maximum context length of a preset large language model, a maximum threshold of a certain elastic interval is set, and then the element stream is traversed according to the maximum utilization strategy, and a corresponding segmentation strategy is used according to the type of the mark. After processing the buffer area, the segmentation rationality is detected, and if the segmentation rationality is destroyed, a more optimal boundary is searched forward or backward, the remaining or subsequent segments before the segmentation are merged, and it is ensured that the merged segments do not exceed the maximum threshold of the buffer area. The processed segments after segmentation are generated piece by piece to assemble a new segmented document.

[0012] Further, the document dynamic segmentation module and the document parsing module are connected through an interface based on the MCP technology, and the interface comprises that the document parsing module proposes a docx format document parsing algorithm to parse underlying data of a document and convert corresponding format features into digital language.

[0013] Further, the document parsing module is configured to use a semantic paragraph reorganization algorithm for text content in a document, and extract character-level attributes and paragraph-level attributes by traversing a document style tree for a style layer.

[0014] Further, the document dynamic segmentation module and the document parsing module are configured to perform binary unpacking, identify a document format, read a zip stream of a.docx file, locate positions of word / document.xml and word / styles.xml, and decompress key xml to a memory for parsing. XML preprocessing is performed, a namespace is registered, a DOM tree or a SAX parser is constructed, and elements of input xml are parsed. A style table is constructed by using a parser to traverse styles.xml in depth, a mapping table of style IDs and attributes is recorded, and a default style benchmark is established. Document cores are parsed by using a parser to traverse all <w:p>The node extracts paragraph-level attributes and performs priority processing; calculates line spacing, performs alignment mode analysis and indentation analysis, and then traverses the child node to extract direct format attributes, associate character style chains and merge inherited attributes to generate a JSON structure file to be processed; unit conversion engine: secondary processing of format style features in the JSON file to be processed, including unit conversion, color value conversion, and special symbol decoding; structured output: generating a JSON file readable by the LLM model based on the processed data.

[0015] Further, the AI agent-based technology integrates the document dynamic segmentation module and the document parsing module into a large language model, including: providing a question and answer interface to upload a file and determining the user's parsing intention; correspondingly, the obtained parsing intention and the document to be processed are processed based on the integrated large language model to output a proofreading draft, including: querying the number of pages of the uploaded file, obtaining the maximum context length allowed by the large language model at the execution time, calling the dynamic segmentation strategy, and generating a document cutting array; cutting the uploaded file based on the cutting array to parse the document; submitting the parsed JSON file to the large language model and performing sequential circulation to generate a document proofreading result; integrating the document proofreading result to obtain the proofreading draft.

[0016] Further, the character-level attributes include font, font size, font color, and bold; the paragraph-level attributes include first-line indentation, line spacing, paragraph before distance, paragraph after distance, and page break; the context length includes 4k, 8k, 16k, and custom; and the document type includes technical documents, legal documents, and scientific papers.

[0017] Further, the semantic integrity priority includes maintaining paragraph integrity if the current element spans pages; the key information protection includes prohibiting segmentation within an element if the current element contains key information; the cross-page paragraph processing includes moving the entire paragraph to the next segment if the current element is larger than the remaining space; and the table paragraph processing includes forcing the entire table to be retained if the current element is a table.

[0018] The document proofreading system suitable for a large language model includes:

[0019] The first module is configured to set the interfaces of the document dynamic segmentation module and the document parsing module based on the MCP technology;

[0020] The second module is configured to integrate the document dynamic segmentation module and the document parsing module into a large language model based on the AI agent technology;

[0021] The third module is configured to obtain a parsing intention and a document to be processed, and process based on the integrated large language model to output a proofreading draft.

[0022] By adopting the technical solutions described above, the present application has the following beneficial effects:

[0023] 1. The present application can achieve reasonable segmentation and parsing functions of documents by setting the interface of the document dynamic segmentation module and the document parsing module based on MCP technology; based on AI agent technology, the document dynamic segmentation module and the document parsing module are integrated into a large language model, which can improve the efficiency and accuracy of segmentation and parsing, and reduce the professional requirements of users. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 The present application proposes a document checking method suitable for a large language model. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0026] As shown in the document checking method suitable for a large language model, comprising: Figure 1

[0027] S1, setting the interface of the document dynamic segmentation module and the document parsing module based on MCP technology;

[0028] S2, integrating the document dynamic segmentation module and the document parsing module into a large language model based on AI agent technology;

[0029] S3, obtaining a parsing intention and a document to be processed, and processing based on the integrated large language model to output a checked draft.

[0030] Model Context Protocol (MCP) is an open-source protocol proposed by Anthropic, aiming to realize the integration of large language models with external data sources and tools, and to establish a secure two-way connection between large models and data sources. Let AI models interact seamlessly with different data sources and tools. It is like a USB-C interface, providing a standardized method to connect AI models to various data sources and tools.

[0031] ​By the MCP technology, a suitable interface can be set to play a control function and a data transmission function, improving the efficiency of data transmission and utilization. And through the existence of the interface, the dynamic segmentation function / document parsing of the self-set / external document is integrated, which can also be the segmentation / parsing function of the large language model to improve the efficiency.

[0032] By the AI agent technology, the document dynamic segmentation module and the document parsing module are integrated into the large language model, which can improve the efficiency of using the large language model compared with direct human control.

[0033] The intention is obtained and analyzed, which can improve the accuracy of document processing.

[0034] The document dynamic segmentation module and the document parsing module are formed by the MCP technology.

[0035] The document dynamic segmentation module performs document segmentation based on a dynamic segmentation strategy and a maximum utilization strategy. The dynamic segmentation strategy includes dynamically selecting an optimal segmentation strategy according to a pre-set context window and a document type, and the segmentation strategy includes semantic integrity priority, key information protection, cross-page paragraph processing, and table paragraph processing. The maximum utilization strategy includes using a context length upper limit close to the context length upper limit for document segmentation under the condition that the document information loss is less than a specified threshold according to the context limit and the document content, so as to reduce the number of subsequent LLM model calls and maximize the utilization of model performance.

[0036] The document dynamic segmentation module is used to extract and mark the element stream of the document, form a buffer area through a temporary storage stack, and store the element stream. The buffer area is initialized according to the maximum context length of the pre-set large language model, a maximum threshold of a certain elastic interval is set, then the element stream is traversed according to the maximum utilization strategy, and the corresponding segmentation strategy is used according to the type of the mark. After processing the buffer area, the segmentation rationality is detected, if there is damage, the more optimal boundary is searched forward / backward, the remaining or subsequent fragments before and after are merged, and it is ensured that the merged fragments do not exceed the maximum threshold of the buffer area. The processed fragments after segmentation are generated one by one to assemble into a new segmented document.

[0037] The document dynamic segmentation module:

[0038] The current market large model has context length limit, and the data generated by document parsing is generally very large. The hardware cost of long context model is high, and the longer context requires higher performance of large model and is more difficult to correct. Therefore, direct parsing and format correction of documents cannot be realized by the current large model on the market. The present application provides a document dynamic segmentation interface for LLM model, which dynamically and intelligently segments the document according to the context length limit of the large model. The interface first dynamically identifies the length required for one segmentation according to the preset context length.

[0039] The interface adopts a two-level processing strategy:

[0040] 1. Dynamic segmentation strategy: based on the preset context window (4k / 8k / 16k configurable), combined with the document type (technical document / law document / scientific paper), the optimal segmentation strategy is dynamically selected, supporting semantic integrity priority, key information protection, cross-page paragraph processing, and table paragraph processing.

[0041] 2. Maximum utilization strategy: according to the context limit and the document content, the document is segmented as close to the upper limit of the context length as possible while preserving the document information without loss, so as to reduce the number of subsequent LLM model calls and maximize the utilization of model performance.

[0042] The interface execution process is divided into four steps:

[0043] 1. Document structure parsing: first extract the document element stream, find the document paragraph, table, title, header and footer information, then mark the element data type, length estimation, and whether it is cross-page, etc., and store it in the temporary storage stack to provide data support for subsequent processing and ensure the completeness of the element information.

[0044] 2. Dynamic segmentation engine: initialize the buffer according to the maximum context length of the preset LLM model, set the maximum threshold of a certain elastic interval, then traverse the element stream according to the maximum utilization strategy, if the current element is a table, force the whole table to be retained; if the current element is cross-page, prefer to keep the paragraph complete, for example, if the current element contains key information, it is prohibited to be segmented within the element; if it is a regular text paragraph, predict whether it will exceed the limit after joining, and execute intelligent segmentation decision. When the buffer is close to full, execute the segmentation decision logic to determine whether the element is larger than the remaining space, if it is larger, move the whole segment to the next segment. According to the different types of elements, special processing is performed to ensure that the information is complete while maximizing the utilization and intelligent segmentation decision.

[0045] 3. Backtracking optimization: after processing the buffer, check the segmentation rationality, check whether the element integrity is damaged, verify the completeness of tables, paragraphs, titles, etc., search forward / backward for better boundaries, merge the remaining or subsequent segments, and ensure that the merged segments do not exceed the maximum threshold of the buffer.

[0046] 4. Output the segmented fragments: according to the optimized buffer, generate segmented processing fragments one by one, assemble into a new segmented document, and provide it to the subsequent stage for processing.

[0047] The interface between the document dynamic segmentation module and the document parsing module based on the MCP technology comprises: the document parsing module proposes a docx format document parsing algorithm to parse the underlying data of the document and convert the corresponding format features into digital language, i.e. JSON format data structure, and the font, font size and other format features are embodied by JSON fields.

[0048] The document parsing module is configured to: for text layer extraction, adopt a semantic paragraph reorganization algorithm; and for style layer, extract character-level attributes and paragraph-level attributes by traversing the document style tree.

[0049] For the text content in the document, a semantic paragraph reorganization is adopted to extract paragraph information. For the format style in the document, character-level attributes and paragraph attributes are extracted by traversing the underlying data of the document.

[0050] The document dynamic segmentation module and the document parsing module are configured to: binary unpacking: identify the document format, read the.docx zip stream, locate the word / document.xml and word / styles.xml positions, and decompress the key xml into the memory for parsing; XML preprocessing: register the namespace, build the DOM tree or SAX parser, and parse the elements of the input xml; stylesheet construction: use the parser to traverse styles.xml in depth, record the mapping table of style ID and attribute, and establish the default style benchmark; parse the document core: use the parser to traverse all <w:p>Node, extract paragraph-level attributes, and perform priority processing; calculate line spacing, analyze alignment and indentation, and then traverse the child nodes to extract direct formatting attributes, associate character style chains, and merge inherited attributes to generate a json structure file to be processed (encapsulate the formatting style features such as line spacing, alignment, indentation, font, font size, etc. in the document into a structured JSON file.); unit conversion engine: convert the units of the processed json, convert Twip to pounds, color value conversion, and special symbol decoding; structured output: generate a json file readable by the LLM model based on the processed data.

[0051] Document parsing module:

[0052] Currently, directly transmitting Office / WPS document original files to LLM large models has format parsing bottlenecks. Due to the complex binary structure of the underlying file format, the unprocessed document cannot effectively extract the layout metadata, resulting in the lack of key structured information in the format review process.

[0053] To solve this problem, a document style parsing interface based on Java language is developed. Its core function is to parse the underlying data structure of docs, wps and other format files, identify, extract and font, font size, line spacing and other style format related attribute value information, and provide accurate data information for subsequent LLM model processing.

[0054] During the parsing process, the text layer extraction uses a semantic paragraph reorganization algorithm to intelligently process complex layouts such as columns, headers and footers while preserving the original paragraph structure; the table layer accurately restores cell merging, border style and nested table structure; the style layer extracts various character-level attributes including font (family), font size, font color, bold / italic, as well as first-line indentation, line spacing, paragraph-level attributes such as paragraph spacing, page breaks, etc. The document segmentation module and the parsing module are deeply coupled in the preprocessing stage to achieve paragraph segmentation and parsing processing. Finally, the parsing results are converted into JSON files that can be recognized by the LLM model through the Gson customized serialization component.

[0055] The specific process is as follows:

[0056] 1. Binary unpacking: identify the document format, read the.docx zip stream, locate the word / document.xml and word / styles.xml positions, and decompress the key xml into memory for parsing.

[0057] 2. XML preprocessing: register namespaces, build DOM trees or SAX parsers, and parse elements for incoming xml.

[0058] 3. Style sheet construction: The parser is used to traverse styles.xml, recording the mapping of style IDs to attributes, and establishing the default style baseline.

[0059] 4. Parsing the document core: The parser is used to traverse all of the elements in document.xml, recording the mapping of element IDs to attributes, and establishing the default element baseline. <w:p>Node, extract paragraph-level attributes, prioritize processing, and format directly to character style, paragraph style, and default style. Calculate line spacing, alignment resolution, and indentation resolution, then traverse child nodes, extract direct format attributes, associate character style chains, and merge inherited attributes to generate a JSON structure for processing.

[0060] 5. Unit conversion engine: secondary processing of format style features in the JSON file to be processed, including unit conversion, color value conversion, and special symbol decoding.

[0061] 6. Structured output: processed data in JSON format readable by LLM model.

[0062] The AI agent-based technology integrates the document dynamic segmentation module and the document parsing module into a large language model, including providing a question and answer interface to upload files and determining the user's parsing intent. Correspondingly, the parsing intent and the document to be processed are obtained based on the integrated large language model processing to output a proofreading draft, including querying the number of pages of the uploaded file, and performing a preset cutting mode according to the context length limit of the large language model, analyzing the document length to generate a cutting array. Based on the cutting array, the uploaded file is cut to perform document parsing. The parsed JSON file is submitted to the large language model and sequentially cycled to generate a document proofreading result. The document proofreading result is integrated to obtain the proofreading draft.

[0063] Using AiAgent technology, the above document dynamic segmentation interface and document style parsing interface are integrated with LLM model for application. Through uploading files and logical judgment of user requirements in the question and answer interface, the LLM model automatically calls the intelligent agent for document proofreading. The intelligent agent first queries the number of pages of the document, and performs a preset cutting mode according to the context length limit of the large model, analyzes the document length to generate a cutting array. The dynamic segmentation function cuts the document according to the array and submits it to the document parsing function for document parsing. After generating the parsed JSON, it is submitted to the large model for sequential circulation to proofread the JSON, identify spelling errors, and check the format and layout specifications to generate a document proofreading result. Finally, the result is integrated to provide a download of the proofreading report. Some commonly used document preset rules are provided, and users can also input custom proofreading descriptions. The intelligent agent generates proofreading rules based on user descriptions to proofread the document.

[0064] The character-level attributes include font, font size, font color, and bold; the paragraph-level attributes include first-line indentation, line spacing, paragraph spacing, and page breaks; the context length includes 4k, 8k, 16k, and custom; and the document types include technical documents, legal documents, and scientific papers.

[0065] The semantic integrity priority includes preferentially maintaining paragraph integrity if the current element crosses a page; the key information protection includes prohibiting segmentation within an element if the current element contains key information; the cross-page paragraph processing includes moving an entire paragraph to the next segment if the current element is larger than the remaining space; and the table paragraph processing includes forcibly retaining an entire table if the current element is a table.

[0066] The document proofreading system suitable for a large language model comprises:

[0067] A first module is configured to set an interface of a document dynamic segmentation module and a document parsing module based on MCP technology;

[0068] A second module is configured to integrate the document dynamic segmentation module and the document parsing module into a large language model based on AI agent technology;

[0069] A third module is configured to acquire a parsing intention and a document to be processed, and process based on the integrated large language model to output a proofreading draft.

[0070] The running process of the document dynamic segmentation module comprises:

[0071] Step 1. Acquire a file to be segmented;

[0072] Step 2. Extract a document element stream, find document paragraphs, tables, titles, headers and footers and the like, then mark the data types, length estimates, whether to cross a page and the like of these element data, and store them in a temporary storage stack to provide data support for subsequent processing and ensure the integrity of the element information;

[0073] Step 3. Dynamically initialize a buffer according to a preset maximum context length of an LLM model, and set a certain elastic interval maximum threshold;

[0074] Step 4. Traverse the element stream according to a maximum utilization strategy, forcibly retain an entire table if the current element is a table, preferentially maintain paragraph integrity if the current element crosses a page, prohibit segmentation within an element if the current element contains key information, and predict whether to exceed the limit after joining a regular text paragraph, and execute intelligent segmentation decision. When the buffer is close to full, execute segmentation decision logic, judge whether the element is larger than the remaining space, and move an entire paragraph to the next segment if it is larger. Special processing is performed according to different types of elements to ensure that the information is retained while maximizing utilization, and intelligent segmentation is performed;

[0075] Step 5. Backtracking optimization verification, after processing the buffer, detect the rationality of segmentation, check whether the integrity of tables, paragraphs, titles and the like is damaged, search forward / backward for a better boundary, merge the previous remaining or subsequent segments, and ensure that the merged segments do not exceed the maximum threshold of the buffer;

[0076] Step 6. According to the optimized buffer described above, generate the segmented processing fragments one by one, assemble into a new segmented document, and output the segmented document for subsequent functions.

[0077] The running process of the document dynamic segmentation module includes:

[0078] Step 21. Obtain the file to be parsed;

[0079] Step 22. Judge whether the document format and type are legal. If not, prompt that the document cannot be parsed. If legal, continue to the next step;

[0080] Step 23. Perform binary unpacking on the docx document, read the.docx zip stream, locate the word / document.xml and word / styles.xml positions, and decompress the key xml into the memory for parsing;

[0081] Step 24. Build a DOM tree and a SAX parser, parse the elements of the input xml, and pass the styles.xml parsing content to the style table building step and the document.xml content to the document parsing core step;

[0082] Step 25. Use the parser to traverse styles.xml in depth, record the mapping table of style ID and attribute, establish a default style benchmark, and generate a style parsing json according to the benchmark. This step is performed synchronously with step 6;

[0083] Step 26. Use the parser to traverse all <w:p>Node, extract paragraph-level attributes, prioritize processing, and format directly to character style, paragraph style, and default style. Calculate line spacing, alignment resolution, and indentation resolution, then traverse child nodes, extract direct formatting attributes, associate character style chains, and merge inherited attributes. Finally, generate the JSON structure for processing.

[0084] Step 27. Convert the above-mentioned JSON to units, convert Twip to pounds, color value conversion, and special symbol decoding.

[0085] Step 28. Process the data to generate a JSON format readable by the LLM model and output.

[0086] Integrate intelligent agent document review process, including:

[0087] Step 31. User uploads file in ai application;

[0088] Step 32. LLM model analyzes user request information, if user uploads document for review is preset analysis type, automatically calls preset document parsing rules; if user raises parsing rules, automatically generates parsing rules according to user's description;

[0089] Step 33. LLM model automatically calls document segmentation module, uses dynamic segmentation strategy and maximum utilization strategy, automatically selects optimal segmentation size according to model context length, dynamically and intelligently segments document, based on preset context window (4k / 8k / 16k configurable), combined with document type (technical document / law document / scientific paper), dynamically selects optimal segmentation strategy, supports semantic integrity priority, key information protection, cross-page paragraph processing, and table paragraph processing. According to context limit and document content, try to maximize the upper limit of context length while preserving document information, to reduce the number of subsequent LLM model calls and maximize model performance;

[0090] Step 34. Transmit segmented document into platform parsing module, call document parsing module, generate high-fidelity structured data through multi-dimensional document parsing technology, and convert parsing result to JSON file recognizable by LLM model;

[0091] Step 35. Submit generated document JSON to LLM model for cyclic parsing in batches, LLM model synchronously performs error correction, punctuation, format layout, content specification, etc. according to preset layout rules or generated temporary rules, generates output result, if this part is not the last part of the document, continue to execute the above steps;

[0092] Step 36. Output all output results to user client interface.

[0093] The above description is for the preferred embodiment of the present application, but the embodiment is not intended to limit the scope of the patent application of the present application. Any equivalent changes or modifications made under the technical spirit of the present application should be covered by the patent scope of the present application.< / w:p> < / w:p> < / w:p> < / w:p>

Claims

1. A document proofreading method suitable for a large language model, characterized by, The application relates to a document dynamic segmentation module and a document parsing module based on an MCP technology, and an interface between the two modules. The document dynamic segmentation module and the document parsing module are integrated into a large language model based on an AI agent technology. The application also relates to a method for processing a document, which comprises the following steps: obtaining an analysis intention and a to-be-processed document, and processing the to-be-processed document based on the integrated large language model to output a proofreading draft. The document dynamic segmentation module is used for performing document segmentation based on a dynamic segmentation strategy and a maximum utilization strategy.

2. The document proofreading method suitable for a large language model according to claim 1, characterized in that, The dynamic segmentation strategy comprises the following steps: dynamically selecting an optimal segmentation strategy according to a preset context window and a document type, and the segmentation strategy comprises semantic integrity priority, key information protection, cross-page paragraph processing and table paragraph processing. The maximum utilization strategy comprises the following steps: according to context restrictions and document content, when document information loss is less than a specified threshold, performing document segmentation close to an upper limit of a context length, so as to reduce the number of subsequent LLM model calls and maximize the utilization of model performance. The document dynamic segmentation module is used for the following steps: extracting an element stream of a document and marking, forming a buffer area through a temporary storage stack, and storing the element stream; 3.The document proofreading method suitable for a large language model according to claim 2, characterized in that, According to a maximum context length of a preset large language model, the buffer area is initialized, a maximum threshold of a certain elastic interval is set, then the element stream is traversed according to the maximum utilization strategy, and corresponding segmentation strategies are adopted according to the types of the marks; After processing the buffer area, the segmentation rationality is detected, if the segmentation rationality is destroyed, a more optimal boundary is searched forwardly or backwardly, the remaining previous or subsequent segments are combined, and it is ensured that the combined segments do not exceed the maximum threshold of the buffer area; The processed segments after segmentation are generated piece by piece to assemble a new segmented document. The interface between the document dynamic segmentation module and the document parsing module based on the MCP technology comprises the following steps: The document parsing module proposes a docx format document parsing algorithm to parse document underlying data and convert corresponding format features into digital language.

4. The document proofreading method suitable for a large language model according to claim 3, characterized in that, The document parsing module is used for the following steps: A semantic paragraph reorganization algorithm is adopted for text content in a document; 5. The document proofreading method suitable for a large language model according to claim 4, characterized in that, Character-level attributes and paragraph-level attributes are extracted by traversing a document style tree for a style layer. The document dynamic segmentation module and the document parsing module are used for the following steps: Binary unpacking: identifying a document format, reading a.zip stream of a.docx, positioning a word / document.xml and a word / styles.xml position, decompressing key xml into memory for parsing; 6. The document proofreading method suitable for a large language model according to claim 5, characterized in that, XML preprocessing: registering a namespace, building a DOM tree or a SAX parser, and performing element parsing on the input xml; Style table construction: using a parser to deeply traverse the styles.xml, recording a mapping table of style IDs and attributes, and establishing a default style benchmark; Nodes are extracted for priority processing; Line spacing is calculated, alignment mode parsing and indentation parsing are performed, then child nodes are traversed, direct format attributes, associated character style chains and inherited attributes are extracted, and a to-be-processed JSON structure file is generated; Parsing the document core: The parser walks through all the <w:p>​< / w:p> ​ unit conversion engine: secondary processing of the format style features in the to-be-processed JSON file, including unit conversion, color value conversion, and special symbol decoding; structured output: generating a JSON file readable by an LLM model based on the processed data.

7. The document proofreading method suitable for a large language model according to claim 6, characterized in that, The AI agent technology-based integration of the document dynamic segmentation module and the document parsing module into a large language model includes providing a question and answer interface to upload a file and determining the user's parsing intention. Correspondingly, the acquisition of the parsing intention and the to-be-processed document based on the integrated large language model processing to output a proof includes: querying the number of pages of the uploaded file, acquiring the maximum context length allowed by the large language model at the execution time, calling a dynamic segmentation strategy, and generating a document cutting array; cutting the uploaded file based on the cutting array to perform document parsing; submitting the parsed JSON file to the large language model and performing sequential circulation to generate a document proofreading result; integrating the document proofreading result to obtain the proof.

8. The document proofreading method suitable for a large language model according to claim 7, characterized in that, The character-level attributes include font, font size, font color, and bold; The paragraph-level attributes include first-line indentation, line spacing, paragraph spacing, and page breaks. The context window size includes 4k, 8k, 16k, and custom; The document types include technical documents, legal documents, and scientific papers.

9. The document proofreading method suitable for a large language model according to claim 8, characterized in that, The semantic integrity priority includes maintaining paragraph integrity if the current element spans multiple pages. The key information protection includes prohibiting segmentation within an element if the current element contains key information. The cross-page paragraph processing includes moving an entire paragraph to the next segment if the current element is larger than the remaining space. The table paragraph processing includes forcing the entire table to be retained if the current element is a table.

10. A document proofreading system suitable for large language models, characterized by, It includes: A first module for setting the interface of the document dynamic segmentation module and the document parsing module based on MCP technology; A second module for integrating the document dynamic segmentation module and the document parsing module into a large language model based on AI agent technology; A third module for acquiring a parsing intention and a to-be-processed document based on the integrated large language model processing to output a proof.

Citation Information

Cited By

  • Intelligent document checking method based on large language model

    CN121859891A