An efficient document segmentation method based on stack mechanism

By adopting a stack-based heading level classification method, the problems of inappropriate document segmentation granularity and semantic relevance are solved, achieving efficient and lightweight document segmentation that can adapt to the needs of different document formats.

CN120995981AInactive Publication Date: 2025-11-21BEIJING GUODIAN ZHISHEN CONTROL TONGDY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510983731.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing document segmentation methods suffer from problems such as inappropriate segmentation granularity, inability to adapt to different document formats, and inconsistency in segmentation results due to ignoring the semantic relationships between paragraphs. Furthermore, advanced natural language processing techniques are resource-intensive and complex to operate.

Method used

A stack-based approach is adopted to divide document blocks by heading levels, construct a heading tree by using Markdown format conversion and heading abstraction, and flexibly set the granularity of segmentation to achieve efficient and semantically coherent document segmentation.

Benefits of technology

It enables flexible adjustment of segmentation granularity based on document type and user needs, ensuring semantic coherence and lightweight operation, improving document processing efficiency, and is applicable to various document formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995981A_ABST
    Figure CN120995981A_ABST
Patent Text Reader

Abstract

The application discloses a high-efficiency document segmentation method based on a stack mechanism, which comprises the following steps: performing document preprocessing, screening the title based on a title paradigm, and segmenting the document according to stack elements; wherein the screening of the title based on the title paradigm comprises grading the title, after traversing the title list, judging the title grade one by one, confirming the title grade, and updating the title grade of the line. The document is continuously divided into blocks of different sizes according to the title grade, so that the high-efficiency, accurate and semantically coherent document segmentation can be realized according to different document types, language structures and user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically a high-efficiency document segmentation method based on a stack mechanism. Background Technology

[0002] With the development of technology, document content is becoming increasingly large and complex. In actual production environments, a large number of traditional documents require analysis and processing. Currently, commonly used traditional document segmentation methods include fixed-length segmentation, manual segmentation based on semantic units, machine recognition segmentation, and segmentation based on headings and content. These methods each have their advantages, but they also suffer from the following problems to varying degrees: the segmentation granularity is too coarse or too fine, thus affecting the efficiency of subsequent information extraction; they cannot adapt to the unified processing of documents in different formats (such as PDF, Word, and HTML); they ignore the semantic relationships between paragraphs, leading to logically inconsistent segmentation results; and advanced natural language processing technologies and some model support require certain hardware resources and specialized technical skills from operators.

[0003] Therefore, there is an urgent need to propose a document segmentation method that is highly adaptive, lightweight, easy to operate, and ensures the logical coherence and semantic relevance of the segmentation blocks to the greatest extent possible, so as to improve the overall efficiency and quality of document processing. Summary of the Invention

[0004] To address the problems existing in the background technology, this invention provides an efficient document segmentation method based on a stack mechanism. This method can achieve efficient, accurate, and semantically coherent document segmentation according to different document types, language structures, and user needs. Its core lies in continuously dividing the document into blocks of different sizes according to heading levels. The technical solution includes:

[0005] Step 1: Perform document preprocessing, including:

[0006] Step 11: Read the document content;

[0007] Step 12: Analyze the document's type and format, convert the document to Markdown format, and clean up any redundancy in the converted content;

[0008] Step 13: Split the document by line breaks, then divide the document into lines; and remove any extra line breaks.

[0009] Step 2: Filter titles based on title paradigms, including:

[0010] Step 21: Initialize variables. Abstract the title into an object, which contains four basic attributes: title content, line number of the title in the document, title format, and title level.

[0011] Step 22: Traverse the text lines, then check each line to see if it is a heading, and determine the starting position of the table of contents.

[0012] Step 23: Initialize the title list;

[0013] Step 24: Divide the headings into different levels;

[0014] Step 3: Divide the document into chunks based on stack elements, including:

[0015] Step 31: Construct the title tree: Initialize an empty stack, push the first element of the title list directly onto the stack, and continue pushing elements onto the stack in the order of the title list;

[0016] Step 32: Segment the text according to heading level: Determine whether there are elements in the current stack with the same heading level as the pushed element. If they exist, the condition for segmenting the text block is met, and some elements are popped from the stack. If they do not exist, no popping is performed.

[0017] Step 33: Generate the segmented content and output the segmentation result.

[0018] Step 1 includes:

[0019] A1. Determine the target Markdown structure;

[0020] A2. Analyze the original document structure;

[0021] A3. Use appropriate tools or methods to split the document into lines;

[0022] A4. Manual adjustment and formatting;

[0023] A5. Proofreading and testing;

[0024] A6. Save.

[0025] Step 24 includes:

[0026] Step 241: Extract title information;

[0027] Step 242: Store title information;

[0028] Step 243: After traversing the list of headings, determine the heading level of each heading, and then confirm the heading level and update the heading level of the row.

[0029] The process of determining the title level is as follows: push the first element in the title list onto the stack. The title level of the first element pushed onto the stack is 1. Continue pushing elements onto the stack in the order of the title list. The key is to determine whether there is a title pattern of the currently pushed element in the stack, and then determine the title level.

[0030] In step 3, when no element is pushed onto the stack, the last text block is determined based on the top element of the stack and the length of the text content.

[0031] The beneficial effects of this invention are as follows:

[0032] 1. This invention enables efficient, accurate, and semantically coherent document segmentation based on different document types, language structures, and user needs. Its core lies in continuously dividing the document into blocks of varying sizes according to heading levels. Current traditional segmentation methods based on headings and content do not distinguish the hierarchical relationships between headings, nor do they consider the logical relationships and semantic connections between headings and body text.

[0033] 2. The method proposed in this embodiment divides the document into several blocks according to the heading level, which can ensure that the segmentation results can greatly avoid problems such as semantic loss and logical incoherence. However, for any segmentation method, the problem of segmentation granularity is inevitable. If the segmentation granularity is too fine, the resulting document blocks will be too small and too numerous. Conversely, if the segmentation granularity is too coarse, the document blocks will be too large.

[0034] 3. In this method, the segmentation granularity corresponds to the classification of heading levels. Therefore, this method flexibly sets the segmentation granularity by setting a parameter (i.e., setting a heading level and its content as the smallest block unit), thereby ensuring applicability to various application scenarios and meeting user needs.

[0035] 4. The granularity of segmentation can be flexibly adjusted according to actual scenario requirements; segmenting according to heading level can largely ensure the logical coherence of semantics; it has wide applicability for processing different document types; it is lightweight and easy to operate, improving the overall efficiency of document processing while ensuring the quality of segmentation. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating an embodiment of an efficient document segmentation method based on a stack mechanism according to the present invention. Detailed Implementation

[0037] The present invention will be further described in detail below with reference to the accompanying drawings.

[0038] like Figure 1 The embodiment of the present invention shown includes:

[0039] Step 1: First, perform document preprocessing, including:

[0040] Step 11: Read the document content

[0041] Step 12: Analyze the document type and format, use appropriate technical means to convert the document into Markdown format, and clean up redundancy in the converted content, such as removing extra blank lines to form text content.

[0042] Step 13: Split the document by line breaks, then divide the document into lines; and remove any extra line breaks.

[0043] Step 2: Filter titles based on heading paradigms. Abstract the title into an object containing four basic attributes: title content, line number of the title in the document, heading paradigm, and heading level. Then, propose a stack mechanism to determine the heading level. Add a heading level attribute to the title, initialize a stack, and push the first element from the heading list onto the stack. The first element pushed onto the stack has a heading level of 1. Continue pushing elements onto the stack in order. The key is to check if the stack already contains the heading paradigm of the currently pushed element to determine the heading level. A simple example: ['5.1.1 CPU-related faults', 26, '^.^.^', 3], where '^.^.^' belongs to the heading paradigm of 5.1.1, and 3 is the heading level; including:

[0044] Step 21: Initialize variables and pre-set all possible heading formats. Markdown format documents have a clear distinction between headings and body text, for example, all headings are marked with a # symbol at the beginning.

[0045] Step 22: Traverse the text lines, then check each line to see if it is a heading, and determine the starting position of the table of contents.

[0046] Step 23: Initialize the title list. After traversing the text lines, determine whether each line is a title. If it is, identify the title content and extract the title information.

[0047] Step 24: Divide the headings into levels, including:

[0048] Step 241: Extract title information;

[0049] Step 242: Store title information;

[0050] Step 243: After traversing the list of headings, determine the heading level of each heading, and then confirm the heading level and update the heading level of the row.

[0051] Step 3: Propose a stack-based mechanism to segment the document using stack elements. Pre-set the segmentation granularity parameter `SPLIT_DEPTH`, initialize an empty stack, and push the first element from the title list directly onto the stack, continuing to push elements according to the title list order. The key is to determine if there are elements in the current stack with the same title level as the pushed element. If so, the condition for segmenting the text into blocks is met, and some elements are popped from the stack; otherwise, no popping is performed. When no elements are pushed onto the stack, the last text block is determined based on the top element of the stack and the length of the text content. Specifically, this includes:

[0052] Step 31: Construct a title tree based on the stack elements;

[0053] Step 32: Segment the text according to heading level;

[0054] Step 33: Generate the segmented content and output the segmentation result.

[0055] Step 1: In the specific process of document preprocessing,

[0056] Converting various document formats to Markdown typically involves several steps, the specific steps of which depend on the original document's format (such as Word, PDF, HTML, etc.). Below are some general steps for converting common document formats to Markdown:

[0057] A1. Determine the target Markdown structure;

[0058] A2. Analyze the original document structure;

[0059] A3. Use appropriate tools or methods to split the document into lines;

[0060] A4. Manually adjust and format (remove newline characters);

[0061] A5. Proofreading and testing;

[0062] A6. Save.

[0063] Step 2, specifically during the title filtering and saving process,

[0064] Pre-define all heading formats as much as possible. Markdown documents will mark all recognized headings with a # symbol at the beginning. This process may mistakenly identify some content that is not a heading. Each heading has a corresponding heading format, which can filter out redundant false headings. Save all possible headings in a heading list. When saving a heading, in addition to the heading content, you also need to save other information about the heading, such as the line number of the heading and the heading format to which the heading belongs. Abstract the heading as an object with four basic attributes: heading content, the line number of the heading in the document, the heading format to which the heading belongs, and the heading level. Here is a simple example: ['5.1.1 CPU Class Fault',26,'^.^.^'], where '^.^.^' belongs to the heading format of 5.1.1, and 26 is the line number of the text corresponding to the heading.

[0065] Step 24: In the specific process of dividing the heading hierarchy,

[0066] The key to defining heading levels is determining if the heading level of the currently pushed element exists in the stack. First, an empty stack is initialized. The first element pushed onto the stack has a heading level of 1. Headings are pushed onto the stack in order, and their heading levels are recorded. All elements in the stack are traversed. If the heading level of a pushed element matches that of an element `i` in the current stack, the heading level of the pushed element is set to that of element `i`. Then, all elements from `i` to the top of the stack are popped, and the pushed element is pushed back onto the stack. If the heading level of a pushed element differs from that of any element in the current stack, it is pushed directly onto the stack, and its heading level is set to the level of the top element plus one. No pop operation is performed. As you can see, heading levels are determined during the pushing process. Heading elements from the heading list are pushed onto the stack in order until all elements have been pushed, thus determining the heading levels of all headings. A simple example: ['5.1.1 CPU-related faults',26,'^.^.^',3], where '^.^.^' belongs to the heading paradigm of 5.1.1, and 3 is the heading level.

[0067] Step 4, in the specific process of document segmentation,

[0068] Initialize an empty stack and pre-set the segmentation granularity parameter SPLIT_DEPTH (if SPLIT_DEPTH=3, it means that the smallest granularity of this segmentation is the third-level heading and its content). Push the first element in the heading list onto the stack without processing it. Then continue to push onto the stack according to the order of the heading list. When pushing onto the stack, first check whether there is an element i in the current stack with the same heading level as the pushed element. If there is, then perform the pop operation. There are two cases when popping from the stack: (1) If the element i is the top element of the stack, directly pop the top element of the stack as the heading. Then determine the text block corresponding to this heading according to the row number index1 of the top element of the stack and the row number index2 of the pushed element. (2) If element i is not the top element of the stack, the top element is still popped and used as the title. Then, the text block corresponding to this title is determined based on the line number (index1) of the top element and the line number (index2) of the pushed element. The pop operation continues until element i is popped before the current pushed element is pushed onto the stack. If the title does not exist, it is not popped. In this case, it is necessary to check whether the difference between the line number (index1) of the top element and the line number (index2) of the pushed element exceeds 1. If it does, the text block can be determined, and the range of the text block is between line 1 and line 2. Then, the current pushed element is pushed onto the stack. If it does not exceed 1, there is no need to determine the text block, and the current pushed element is pushed onto the stack directly. Then, the push operation continues according to the order of the title list.

[0069] Until no more elements are pushed onto the stack, the top element of the stack is the last element in the title list. At this point, the segmented content is generated, which requires special handling. Specifically, it forms the last text block, whose range is from line index 1 of the top element's line number to the end of the document. The line number of the end of the document can be determined using the function len(texts). Finally, the text block is output.

Claims

1. A high-efficiency document segmentation method based on a stack mechanism, characterized in that, include: Step 1: Perform document preprocessing, including: Step 11: Read the document content; Step 12: Analyze the document's type and format, convert the document to Markdown format, and clean up any redundancy in the converted content; Step 13: Split the document by line breaks, then divide the document into lines; and remove any extra line breaks. Step 2: Filter titles based on title paradigms, including: Step 21: Initialize variables. Abstract the title into an object, which contains four basic attributes: title content, line number of the title in the document, title format, and title level. Step 22: Traverse the text lines, then check each line to see if it is a heading, and determine the starting position of the table of contents. Step 23: Initialize the title list; Step 24: Divide the headings into different levels; Step 3: Divide the document into chunks based on stack elements, including: Step 31: Construct the title tree: Initialize an empty stack, push the first element of the title list directly onto the stack, and continue pushing elements onto the stack in the order of the title list; Step 32: Segment the text according to heading level: Determine whether there are elements in the current stack with the same heading level as the pushed element. If they exist, the condition for segmenting the text block is met, and some elements are popped from the stack. If they do not exist, no popping is performed. Step 33: Generate the segmented content and output the segmentation result.

2. The efficient document segmentation method based on a stack mechanism according to claim 1, characterized in that, Step 1 includes: A1. Determine the target Markdown structure; A2. Analyze the original document structure; A3. Use appropriate tools or methods to split the document into lines; A4. Manual adjustment and formatting; A5. Proofreading and testing; A6. Save.

3. The efficient document segmentation method based on a stack mechanism according to claim 1, characterized in that, Step 24 includes: Step 241: Extract title information; Step 242: Store title information; Step 243: After traversing the list of headings, determine the heading level of each heading, and then confirm the heading level and update the heading level of the row.

4. The efficient document segmentation method based on a stack mechanism according to claim 3, characterized in that, The process of determining the title level is as follows: push the first element in the title list onto the stack. The title level of the first element pushed onto the stack is 1. Continue pushing elements onto the stack in the order of the title list. The key is to determine whether there is a title pattern of the currently pushed element in the stack, and then determine the title level.

5. The efficient document segmentation method based on a stack mechanism according to claim 1, characterized in that, In step 3, when no element is pushed onto the stack, the last text block is determined based on the top element of the stack and the length of the text content.