A method for merging and splitting the content of document paragraphs based on a combined tree structure

Through the document paragraph content merging and segmentation method based on the combined tree structure, the problems of incomplete identification and semantic segmentation when dividing document content in the prior art are solved, and semantic rich and complete document block generation are achieved.

CN119918505BActive Publication Date: 2025-06-17SHANGHAI ICEKREDIT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413265.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-17
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

When the existing document content segmentation method utilizes paragraph hierarchy structure, there are problems such as incomplete identification, loss of information and semantic separation, and fixed-length segmentation results in incomplete semantic information of document blocks.

Method used

The document paragraph content merging and segmentation method based on a combination tree structure is adopted. By traversing the document content line by line, identifying the title hierarchy, and building a content structure tree, pruning and sentence-level segmentation are performed to generate semantic rich document blocks.

Benefits of technology

More accurate paragraph hierarchical recognition is achieved, reducing information loss and semantic fragmentation, and ensuring the integrity of sentences and semantic information in document blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918505B_ABST
    Figure CN119918505B_ABST
Patent Text Reader

Abstract

This application designs a method for merging and splitting document paragraph content based on a combined tree structure, including: S1. Obtain the document content using a document parsing tool; S2. Traverse the document content line by line to obtain a list of line text contents secs; Assign a title level to each line of content to obtain a list of level recognition results levels; S3. Use the obtained list of line text contents secs, list of level recognition results levels, and the processing depth depth defined by the program to generate a list of paragraph title organization groups cks; S4. According to the list of paragraph title organization groups cks, group the title path information to construct a paragraph title information group cks-group; S5. Organize the paragraph title information group cks-group into the form of a content structure tree; S6. Prune and merge the content structure tree to merge the text content at a larger level; S7. Process the pruned content structure tree to generate the document block content of the current file. This application can efficiently construct a document block to be matched with rich and complete semantics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and particularly relates to a method for merging and splitting document paragraph content based on a combined tree structure. Background Art

[0002] In the current Retrieval-Augmented Generation (RAG) intelligent question-and-answer application, one of the prerequisites is to segment the content of the documents uploaded by users. Currently, the mainstream document content segmentation schemes include fixed-length segmentation, sliding window segmentation, paragraph recognition segmentation, or semantic segmentation model segmentation, etc. Fewer schemes consider segmenting by utilizing the paragraph hierarchy structure to which the content itself belongs. The existing document content paragraph hierarchy recognition and merging schemes only complete the content recognition of each paragraph level based on heuristic title rules and binary search methods.

[0003] The recognized content of the existing paragraph hierarchy scheme does not fully comply with the content order of the document, there is a long-tail disordered situation, and in the case where the heuristic title rules are not fully covered, there is a situation of overly fine segmentation granularity, resulting in a certain degree of information loss and semantic fragmentation in the document blocks obtained by subsequent content segmentation.

[0004] The currently commonly used solution is as follows: The paragraph content merging and splitting method combines the heuristic rule-based title recognition and the traditional splitting method. The general steps of this method are as follows: Step S1, the computer uses a file parsing tool to traverse and obtain the file content line by line from the text file, and stores the content line by line to obtain a list of line texts, and synchronously retains the order id of the line texts. Step S2, the computer traverses the above list of line texts in order according to the pre-defined title recognition rule group to obtain the title level to which each line text belongs, and organizes the order of appearance of the line texts according to the level to which the line texts belong. For example, if there are a total of 5 recognition rules in the pre-defined title recognition rule group, then the finest level obtained for the text that hits the rule is 5, and the line texts that do not hit the recognition rule are uniformly assigned to level 6 (that is, this line text is not a title but a simple content text). Finally, a level information dictionary is obtained, where the key of the dictionary is the title level, and the value of the dictionary is the order id of the line texts included in this level. Step S3, the computer program uses the obtained level information dictionary to find the paragraph title to which each line text belongs at each level in order according to a certain backtracking depth. During the loop, according to the current order id of the line text and the list of order ids of the parent level line texts, the binary search algorithm is used to obtain the id of the paragraph line text to which the current line text belongs. Step S4, the computer program starts the above loop from the lowest level (level 6 in the above example) until the set backtracking depth is reached (assuming the backtracking depth is 3, then the program only loops through the line texts at levels 6, 4, and 5 to find the paragraph levels to which the line texts at these levels belong), and at this time, the paragraph level path of each line text can be obtained. Step S5, the computer program organizes all the line text contents according to the paragraph level path, that is, splices the line text contents belonging to the same paragraph level together, and then splits the text content in a fixed length manner to obtain the final document blocks.

[0005] However, in the above method, when directly using the paragraph levels obtained by rule recognition to organize the text content, due to the limitations of rule coverage and the possible non-uniqueness of the order id attribution of the line texts (that is, the current line text may be both the paragraph title of a certain line text and the text content of the current level), the document blocks obtained by the existing method cannot fully comply with the orderliness of the content in the original document; and when directly using the paragraph levels obtained by rule recognition to organize the text content and perform splitting, due to the deviation between the pre-defined rules and the title distribution of the actual parsed file, there are cases where the granularity of the line text content at some paragraph levels is too fine, resulting in the lack of context information and semantics in the content of the final obtained document blocks; finally, when dividing the file content by the fixed length splitting method to obtain the final document blocks, it may lead to the situation where the first / last sentences of the current document block are split, thereby making the semantic information of the current document block incomplete.

[0006] Therefore, how to design a better method for merging and splitting paragraph-level content to obtain document blocks with more complete and rich semantics has become a current design difficulty. Summary of the Invention

[0007] To solve the above problems, this application designs a method for merging and splitting document paragraph content based on a combined tree structure, aiming to efficiently construct document blocks to be matched with rich and complete semantics by combining the hierarchical recognition effect of heuristic title rules, sentence-level content splitting, and dynamic paragraph-level content merging based on the tree structure.

[0008] A method for merging and splitting document paragraph content based on a combined tree structure includes the following steps:

[0009] Step S1: Use a document parsing tool to obtain the document content;

[0010] Step S2: Traverse the document content line by line to obtain a list of line text contents secs; assign a title level to each line of content and obtain a list of hierarchical recognition results levels;

[0011] Step S3: Use the list of line text contents secs, the list of hierarchical recognition results levels, and the processing depth depth defined by the program obtained in Step S2 to generate a list of paragraph title organization groups cks;

[0012] Step S4: According to the list of paragraph title organization groups cks, group the title path information to construct a paragraph title information group cks-group;

[0013] Step S5: Organize the paragraph title information group cks-group into the form of a content structure tree;

[0014] Step S6: Prune and merge the content structure tree to merge text content at a larger level;

[0015] Step S7: Process the pruned content structure tree to generate the document block content of the current file.

[0016] Preferably, Step S2 includes:

[0017] Step S21: Traverse the document content line by line to obtain a list of line text contents secs;

[0018] Step S22: Sequentially traverse each line of text content, use a predefined set of regular expressions to perform title level recognition on each line of text, and store the recognition results in the list of hierarchical recognition results levels.

[0019] Preferably, the specific method of Step S22 includes:

[0020] Step S221: Combine K title recognition regular expressions into a regular expression group;

[0021] Step S222: Initialize a nested list with a size of K + 1;

[0022] Step S223: Use the regular expression group to match each line of text;

[0023] If the regular expression successfully matches the line of text, assign the line of text to the level corresponding to the regular expression, and add the order of the line of text to the hierarchical recognition result list levels;

[0024] If the regular expression does not match the line of text, assign the line of text to the default last level K + 1, and add the order of the line of text to the hierarchical recognition result list levels.

[0025] Preferably, in step S3, the method for generating the paragraph title organization grouping list cks includes:

[0026] Step S31: Reverse the content of the hierarchical recognition result list levels, and start identifying the parent title of the line of text from the text of the last level k + 1;

[0027] Step S32: Traverse each level and each line of text within the level in the hierarchical recognition result list levels until the level reaches the specified depth;

[0028] Step S33: Use the binary search algorithm to find the parent title paragraph corresponding to the line of text in the current level levels[l i in all previous levels levels[l j ; and combine the order r idi of the current line of text, and the order r idk of the text corresponding to the parent title paragraph into a tuple (r idk , r idi ), add it to the paragraph title organization grouping list cks, and initialize the paragraph title organization grouping list cks;

[0029] Step S34: Compare the content of the line of text recorded in the paragraph title organization grouping list cks with the content in the line of text content list secs to find the unprocessed line of text;

[0030] Step S35: For each unprocessed line of text, use the binary search algorithm to determine its parent title and fill it into the paragraph title organization grouping list cks.

[0031] Preferably, in step S4, the dictionary key of the paragraph title information group cks-group is all the identified title paths, and the dictionary value of the paragraph title information group cks-group is the text content included under the paragraph title path.

[0032] Preferably, the specific method of step S5 includes:

[0033] Step S51: Initialize a multi-way content tree structure MultiWayTree, where the node elements in the tree are dictionaries. If the node is a non-leaf node, i.e., a title element, the value of the dictionary is the sub-title node; if the node is a leaf node, the value of the dictionary is a list of line text contents; the key of the dictionary is the title content;

[0034] Step S52: Traverse the paragraph title information group cks-group, and insert the title path and the corresponding line text content into the content tree MultiWayTree. During the insertion process, to ensure the validity of the node order, based on the line id of the current line text and the line id stored under the current path and the line id list of adjacent paths, determine whether the insertion of the current line text will affect the orderliness and whether to continue adding under the current path.

[0035] Preferably, the specific method of step S6 includes:

[0036] Step S61: Depth-first traverse each path of the content structure tree, calculate the current path depth, and use the path depths of all paths in the current subtree to calculate the median of the path depths, and record this median;

[0037] Step S62: For paths with a path depth exceeding the median, perform pruning and content merging.

[0038] Preferably, the specific method of step S7 includes:

[0039] Step S71: Traverse the pruned structure tree, process each branch, and generate a description text for the paragraph section according to the title and paragraph text content stored in the node;

[0040] Step S72: Use a sentence-level segmentation scheme to process the description text of the paragraph section to obtain all the final document block contents of the file.

[0041] The advantages and effects of this application are as follows:

[0042] This application designs a method for merging and splitting document paragraph content based on a combined tree structure, including: S1. Obtain the document content using a document parsing tool; S2. Traverse the document content line by line to obtain a list of line text contents secs; Assign a title level to each line of content to obtain a list of level recognition results levels; S3. Use the obtained list of line text contents secs, list of level recognition results levels, and the processing depth depth defined by the program to generate a list of paragraph title organization groups cks; S4. Group the title path information according to the list of paragraph title organization groups cks to construct a paragraph title information group cks-group; S5. Organize the paragraph title information group cks-group into a content structure tree form; S6. Prune and merge the content structure tree to merge the text content at a larger level; S7. Process the pruned content structure tree to generate the document block content of the current file. This application can efficiently construct a document block to be matched with rich and complete semantics.

[0043] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, so that it can be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following takes the preferred embodiments of this application and combines the drawings to describe in detail as follows.

[0044] According to the following detailed description of the specific embodiments of this application in conjunction with the drawings, those skilled in the art will be more clear about the above and other purposes, advantages and features of this application. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale.

[0046] Figure 1 Flowchart of a method for merging and splitting document paragraph content based on a combined tree structure designed for this application;

[0047] Figure 2 Reference pseudo-code for step S3 designed for this application;

[0048] Figure 3 Example diagram of node ids for ordered merging designed for this application;

[0049] Figure 4Example diagram of the document block content of path depth median merging and sentence-level segmentation designed for this application. Detailed implementation manners

[0050] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. In the following description, specific details such as specific configurations and components are provided only to assist in a comprehensive understanding of the embodiments of this application. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Additionally, descriptions of known functions and structures are omitted for clarity and conciseness.

[0051] It should be understood that the "one embodiment" or "this embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of this application. Therefore, the "one embodiment" or "this embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0052] In addition, this application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0053] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously. The term " / and" in this document is a description of another association object relationship, indicating that there can be two relationships. For example, A / and B can represent: A exists alone, and both A and B exist. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.

[0054] The term "at least one" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, at least one of A and B can represent: A exists alone, both A and B exist simultaneously, and B exists alone.

[0055] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion.

[0056] Example 1: Please refer to Figure 1 , this example mainly introduces a method for merging and splitting document paragraph content based on a combined tree structure, including the following steps:

[0057] Step S1: Use a document parsing tool to obtain the document content;

[0058] Step S2: Traverse the document content line by line to obtain a list of line text contents secs; assign a title level to each line of content and obtain a list of level recognition results levels; specifically: according to the recognition rules (heuristic title rules), assign each line of text to different levels, and record the line text and the corresponding level in a list;

[0059] Step S3: Use the list of line text contents secs, the list of level recognition results levels, and the processing depth depth defined by the program obtained in Step S2 to generate a list of grouped paragraph title organizations cks;

[0060] Step S4: Group the title path information according to the list of grouped paragraph title organizations cks to construct a group of paragraph title information cks-group;

[0061] Step S5: Organize the group of paragraph title information cks-group into the form of a content structure tree;

[0062] Step S6: Prune and merge the content structure tree to merge the text content at a larger level;

[0063] Step S7: Process the pruned content structure tree to generate the document block content of the current file.

[0064] Furthermore, the said Step S2 includes:

[0065] Step S21: Traverse the document content line by line to obtain a list of line text contents secs;

[0066] Step S22: Traverse each line text content in sequence, use a predefined set of regular expressions to identify the title level of each line text, and store the recognition results in the list of level recognition results levels.

[0067] Furthermore, the specific method of the said Step S22 includes:

[0068] Step S221: Combine K title recognition regular expressions into a regular expression group;

[0069] Step S222: Initialize a nested list with a size of K + 1;

[0070] Step S223: Use the regular expression group to match each line of text;

[0071] If the regular expression successfully matches the line of text, assign the line of text to the level corresponding to the regular expression, and add the order of the line of text to the hierarchical recognition result list levels;

[0072] If the regular expression does not match the line of text, assign the line of text to the default last level K + 1, and add the order of the line of text to the hierarchical recognition result list levels.

[0073] Specifically, assume that the order of the current line of text is r-id i ; (if the current line of text is the first line of text, then r-id i is 0), then there are the following matching situations:

[0074] If the regular expression successfully matches the line of text (that is, the title pattern defined by the regular expression can be matched in the current line of text, e.g., the regular expression "Title \d" can match "Title 1"), assign the line of text to the level corresponding to the regular expression - add the order of the current line of text to the list levels[l j ;

[0075] If the match is not successful, assign the line of text to the default last level K + 1 - add the order of the current line of text to the list levels[K + 1].

[0076] Furthermore, please refer to Figure 2 , Figure 2 the reference pseudo-code of Step S3 designed for this application; in the said Step S3, the method for generating the paragraph title organization grouping list cks includes:

[0077] Step S31: Reverse the content of the hierarchical recognition result list levels, and start recognizing the parent title to which the line of text belongs from the text of the last level k + 1;

[0078] Step S32: Traverse each level and each line of text within each level in the hierarchical recognition result list levels until the level reaches the specified depth;

[0079] For hierarchical traversal, depth controls the hierarchical depth d for identifying parent headings, where d ≤ depth. That is, the program will identify the parent headings for which levels of line text - assuming depth is 3, the program will only identify the parent headings for line text belonging to levels k+1, k, and k-1. This is done to introduce a controllable way to avoid introducing headings as ordinary content text into the final result.

[0080] Step S33: Use the binary search algorithm to find the parent heading paragraph corresponding to the line text of the current level levels[l i among all previous levels; and combine the order r j of the current line text, and the text order r idi corresponding to the parent heading paragraph into a tuple (r idk , r idk , r idi ), add it to the paragraph heading organization grouping list cks, and initialize the paragraph heading organization grouping list cks.

[0081] For unprocessed line text, use the binary search algorithm to find the parent heading paragraph of the current line text among all previous levels levels[l i , (l i< l j ).

[0082] Specifically, according to the order r idi of the current line text, use the binary search algorithm to find the insertable position jj of r idi in levels[l i . Using the size comparison property of the binary search algorithm, the program takes the line text at position jj in levels[l i as the parent heading of the current line text and adds it to the list of parent headings involved in the current line text.

[0083] Step S34: Compare the line text recorded in the paragraph heading organization grouping list cks with the content in the line text content list secs to find unprocessed line text;

[0084] Step S35: For each unprocessed line text, use the binary search algorithm to determine its parent heading and fill it into the paragraph heading organization grouping list cks.

[0085] For each unprocessed line text, use the binary search idea to determine its parent heading and fill it into cks. The final content of cks is as follows: cks[i] = [(s a , s b ),...], where s bFor the order of the lines traversed, s a Is the line text content.

[0086] Furthermore, in the step S4, the dictionary key of the paragraph title information group cks-group is all the recognized title paths, and the dictionary value of the paragraph title information group cks-group is the text content included under the paragraph title path.

[0087] Traverse cks to construct the paragraph title information group cks-group, where the key of the dictionary is all the recognized title paths (composed of all the parent titles in cks[i]), and the value of the dictionary is the text content included under the paragraph title path. a Consisting), and the value of the dictionary is the text content included under the paragraph title path.

[0088] Note: Take cks[i]=[(s a , s b ),..., (s a+n , s b+n )] as an example, s a+1 To s a+n The line text content of constitutes the title path, and s a Is the paragraph content under the title path.

[0089] Furthermore, the specific method of the step S5 includes:

[0090] Step S51: Initialize a multi-way content tree structure MultiWayTree, where the node elements in the tree are dictionaries. If the node is a non-leaf node, i.e., a title element, the value of the dictionary is the sub-title node. If the node is a leaf node, the value of the dictionary is a list of line text content; the key of the dictionary is the title content;

[0091] Step S52: Traverse the paragraph title information group cks-group, and insert the title path and the corresponding line text content into the content tree MultiWayTree. During the insertion process, to ensure the validity of the node order, according to the line id of the current line text and the line id stored under the current path and the line id list of the adjacent paths, determine whether the insertion of the current line text will affect the orderliness and whether to continue adding under the current path.

[0092] Furthermore, the specific method of the step S6 includes:

[0093] Step S61: Perform a depth-first traversal of each path of the content structure tree, calculate the current path depth, and use the depth of all paths of the current subtree to calculate the median of the path depths, and record this median;

[0094] Step S62: For paths with a path depth exceeding the median, perform pruning and content merging.

[0095] Further, the specific method of step S7 includes:

[0096] Step S71: Traverse the pruned structure tree, process each branch, and generate descriptive text for the paragraph section according to the title and paragraph text content stored in the node;

[0097] Step S72: Use the sentence-level segmentation scheme to process the descriptive text of the paragraph section to obtain all the final document block contents of the file.

[0098] Through this step, the sentence-level content segmentation scheme can, to a certain extent, alleviate the problem of semantic fragmentation of document blocks. The mainstream document segmentation scheme does not consider the integrity of sentences within document blocks - there is a phenomenon of in-sentence cutting, resulting in incomplete / fragmented semantic information in the embeddings obtained for document blocks. The sentence-level content segmentation scheme can ensure that each sentence within a document block is complete, and combined with the overlap (block content overlap) technique, it can better ensure the integrity of the semantic information of document blocks.

[0099] Example of the parsing result of this method in the opening report of a certain paper in Baidu Library:

[0100] Please refer to Figure 3 , Figure 3 as an example of the node IDs for ordered merging, Figure 4 and Figure 4 as an example of the document block contents after path depth median merging and sentence-level segmentation. As can be seen from

[0101] This application designs a method for merging and segmenting document paragraph contents based on a combined tree structure, including: S1: Obtain the document content using a document parsing tool; S2: Traverse the document content line by line to obtain a list of line text contents secs; Assign a title level to each line content to obtain a list of level recognition results levels; S3: Use the obtained list of line text contents secs, list of level recognition results levels, and the processing depth depth defined by the program to generate a list of paragraph title organization groups cks; S4: Group the title path information according to the list of paragraph title organization groups cks to construct a paragraph title information group cks-group; S5: Organize the paragraph title information group cks-group into a content structure tree form; S6: Prune and merge the content structure tree to merge the text content at a larger level; S7: Process the pruned content structure tree to generate the document block contents of the current file. This application can efficiently construct semantically rich and complete document blocks to be matched.

[0102] The above are only the preferred embodiments of the present invention, and thus do not limit the protection scope of the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any changes, modifications, substitutions, integrations, and parameter changes to these embodiments that are made through conventional substitutions or that can achieve the same functions without departing from the principle and spirit of the present invention fall within the protection scope of the present invention.

Claims

1. A document paragraph content merging and segmenting method based on a combined tree structure, characterized in that: The following steps are involved: Step S1, using a document parsing tool to obtain document content; Step S2, traverse the document content line by line to obtain a line text content list secs; assign a title level to each line content, and obtain a level recognition result list levels; Step S3, using the line text content list secs obtained in step S2, the level recognition result list levels and the processing depth depth defined by the program, to generate a paragraph title organization grouping list cks; Step S4, grouping the title path information according to the paragraph title organization group list cks, and constructing the paragraph title information group cks-group; Step S5, organizing the paragraph title information group cks-group into a content structure tree format; Step S6, pruning and merging the content structure tree, merging text content at a larger level; Step S7, processing the pruned content structure tree to generate the document block content of the current file; In step S3, the method of generating the paragraph title organization group list cks includes: Step S31, reverse the content of the level recognition result list levels, and start recognizing the parent level title to which the line text belongs from the text of the last level k+1; Step S32, traversing each level and each line of text in the level in the level recognition result list levels until the level reaches a specified depth; Step S33: Use a binary search algorithm to find all previous levels [l i ] to find the current level levels[l j ] the parent title paragraph corresponding to the line text; and change the order of the current line text to r idi , and the text order corresponding to the parent title paragraph r idk Composite tuple (r idk , r idi ), add it to the paragraph title organization grouping list cks, and initialize the paragraph title organization grouping list cks; Step S34, comparing the line texts recorded in the paragraph title organization grouping list cks with the content in the line text content list secs, to find out the line texts that have not been processed; Step S35: For each unprocessed line of text, use a binary search algorithm to determine its parent title and fill it into the paragraph title organization group list cks.

2. According to the method for merging and segmenting document paragraph contents based on a combined tree structure according to claim 1, it is characterized in that: The step S2 comprises: Step S21, traverse the document content line by line to obtain a line text content list secs; Step S22, sequentially traverse each line of text content, use a predefined regular expression set to perform title level recognition on each line of text, and store the recognition results in the level recognition result list levels.

3. A document paragraph content merging and segmenting method based on a combined tree structure according to claim 2, characterized in that: The specific method of step S22 includes: Step S221, combining K title recognition regular expressions into a regular expression group; Step S222, initialize a nested list of size K+1; Step S223, using a regular expression group to match each line of text; If the regular expression successfully matches the line of text, the line of text is assigned to the level corresponding to the regular expression, and the order of the line of text is added to the level recognition result list levels; If the regular expression does not match the line of text, the line of text is assigned to the default last level K+1, and the order of the line of text is added to the level recognition result list levels.

4. The method for merging and segmenting document paragraph contents based on a combined tree structure according to claim 1, characterized in that: In the step S4, the dictionary keys of the paragraph title information group cks-group are all the identified title paths, and the dictionary values ​​of the paragraph title information group cks-group are the text content contained in the paragraph title path.

5. The method for merging and segmenting document paragraph contents based on a combined tree structure according to claim 1, characterized in that: The specific method of step S5 includes: Step S51, initialize a multi-way content tree structure MultiWayTree, wherein the node elements in the tree are dictionaries, if the node is a non-leaf node, i.e., a title element, the value of the dictionary is a subtitle node, if the node is a leaf node, the value of the dictionary is a row text content list; the key of the dictionary is the title content; Step S52, traverse the paragraph title information group cks-group, insert the title path and the corresponding line text content into the content tree MultiWayTree. During the insertion process, in order to ensure the validity of the node order, according to the line id of the current line text and the line ids already stored under the current path and the line id list of the adjacent path, it is determined whether the insertion of the current line text will affect the orderliness, and whether to continue adding to the current path.

6. The method for merging and segmenting document paragraph contents based on a combined tree structure according to claim 1, characterized in that: The specific method of step S6 includes: Step S61: Depth-first traverse each path of the content structure tree, calculate the current path depth, and use all path depths of the current subtree to calculate the median of the path depths, and record the median; Step S62: For paths whose path depth exceeds the median, pruning and content merging are performed.

7. The method for merging and segmenting document paragraph contents based on a combined tree structure according to claim 1, characterized in that: The specific method of step S7 includes: Step S71, traverse the pruned structure tree, process each branch, and generate a description text of the paragraph chapter according to the title and paragraph text content stored in the node; Step S72: Use a sentence-level segmentation scheme to segment the description text of the paragraphs and chapters into blocks to obtain the final content of all document blocks of the file.

Citation Information

Patent Citations

  • Document title tree construction method and device, electronic equipment and storage medium

    CN111460083A

  • Document segmentation method and device, computer equipment and storage medium

    CN119474250A