Multistage document segmentation method adopting self-adaptive dynamic partitioning algorithm

By using an adaptive dynamic segmentation algorithm, information density and overlap are calculated based on document type and chapter information, and the segmentation strategy is dynamically adjusted, which solves the problem of inaccurate segmentation in existing technologies and achieves high-quality and efficient document segmentation.

CN120951941APending Publication Date: 2025-11-14DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202510979371.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing document segmentation methods cannot adapt to changes in the content and structure of different documents, resulting in inaccurate segmentation results that affect subsequent document processing and analysis.

Method used

An adaptive dynamic segmentation algorithm is adopted. By determining the document type and chapter information, calculating the information density and overlap, and dynamically adjusting the segmentation strategy, the segmented document chapter content is ensured to be complete and logically coherent.

Benefits of technology

It improves the quality and efficiency of document segmentation, avoids information loss or fragmentation of chapter content, meets actual usage needs, and facilitates subsequent management and information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951941A_ABST
    Figure CN120951941A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level document segmentation method adopting an adaptive dynamic partitioning algorithm, and relates to the technical field of document segmentation, the method comprises the following steps: determining a target document type and target chapter information of a to-be-segmented document; when the target document type is a standard document, calculating a target information density corresponding to the to-be-segmented document based on the target chapter information; and based on the target information density, determining a target chapter overlapping degree corresponding to the to-be-segmented document, and segmenting each target chapter according to the target chapter overlapping degree to complete segmentation processing of the to-be-segmented document. According to the method and device, it is ensured that the chapter content of the segmented document is complete and logically coherent, the problems of information loss or chapter content segmentation and the like caused by blind segmentation are solved, the document segmentation quality and efficiency are effectively improved, and the segmented document better meets the actual use requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of document segmentation technology, and in particular to a multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm. Background Technology

[0002] Retrieval Augmented Generation (RAG) is a technique applied to the field of knowledge content retrieval. It combines a vector database and a large language model, aiming to enhance the generative capabilities of the large model through retrieval techniques. The construction process of the vector database involves: document preparation → document segmentation → slice vectorization → storage. Document segmentation optimizes the construction of the vector database and improves the generation quality of the large language model. In this technique, the large model uses the segmented document content after retrieval ranking as a reference for its responses; therefore, document segmentation is particularly important for the quality of the large model's responses.

[0003] In terms of document segmentation techniques in retrieval enhancement generation, previous methods mainly included two approaches: segmentation based on a fixed number of characters and segmentation based on specified punctuation marks.

[0004] Fixed-word-count segmentation methods often fail to adapt well to variations in document content and structure. Because paragraph lengths and content differ significantly across documents, fixed-word-count segmentation can lead to inaccurate results, impacting subsequent document processing and analysis. Summary of the Invention

[0005] The technical problem to be solved by this invention is to address the shortcomings of existing technologies, specifically by providing a multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm, as detailed below: 1) In a first aspect, the present invention provides a multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm, the specific technical solution of which is as follows: Determine the target document type and target chapter information of the document to be segmented; When the target document type is a standard document, the target information density corresponding to the document to be segmented is calculated based on the target chapter information. The target information density represents the average information entropy corresponding to each chapter. Based on the target information density, the overlap of the target chapters corresponding to the document to be segmented is determined. According to the overlap of the target chapters, each target chapter is segmented to complete the segmentation process of the document to be segmented.

[0006] The beneficial effects of the multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm provided by this invention are as follows: By clearly defining the target document type and target chapter information, a clear guide and basis are provided for subsequent segmentation operations. When the target document is a standard document, the target information density, i.e., the average information entropy corresponding to each chapter, is calculated based on the target chapter information. This process can quantitatively assess the information richness and importance of each chapter, thus providing a scientific and reasonable reference for determining the overlap of target chapters. Then, each target chapter is segmented according to the overlap of target chapters, ensuring that the content of the segmented document chapters is complete and logically coherent. This avoids problems such as information loss or fragmentation of chapter content caused by blind segmentation, effectively improving the quality and efficiency of document segmentation. The segmented document is more in line with actual usage needs and facilitates subsequent document management and information retrieval operations.

[0007] Based on the above solution, the present invention can be further improved as follows.

[0008] Furthermore, when the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

[0009] Furthermore, when the target document type is an extra-long document, the document to be segmented is processed according to a preset fixed block size.

[0010] Furthermore, the method for determining the overlap of target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

[0011] 2) Secondly, the present invention also provides a multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm, the specific technical solution of which is as follows: The acquisition module is used to: determine the target document type and target chapter information of the document to be segmented; The calculation module is used to: when the target document type is a standard document, calculate the target information density corresponding to the document to be segmented based on the target chapter information. The target information density represents the average information entropy corresponding to each chapter. The segmentation module is used to: determine the overlap of target chapters corresponding to the document to be segmented based on the target information density, and segment each target chapter according to the overlap of target chapters to complete the segmentation process of the document to be segmented.

[0012] Based on the above solution, the present invention can be further improved as follows.

[0013] Furthermore, when the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

[0014] Furthermore, when the target document type is an extra-long document, the document to be segmented is processed according to a preset fixed block size.

[0015] Furthermore, the method for determining the overlap of target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

[0016] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to perform any of the methods described above.

[0017] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above methods.

[0018] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0019] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating a multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm according to an embodiment of the present invention. Figure 2 This is a structural framework diagram of an electronic device according to the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0021] like Figure 1 As shown in the figure, a multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm according to an embodiment of the present invention includes the following steps: S1, Determine the target document type and target chapter information of the document to be segmented; S2, when the target document type is a standard document, calculate the target information density corresponding to the document to be segmented based on the target chapter information. The target information density represents the average information entropy corresponding to each chapter. S3. Based on the target information density, determine the overlap of the target chapters corresponding to the document to be segmented. Based on the overlap of the target chapters, segment each target chapter to complete the segmentation process of the document to be segmented.

[0022] The beneficial effects of the multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm provided by this invention are as follows: By clearly defining the target document type and target chapter information, a clear guide and basis are provided for subsequent segmentation operations. When the target document is a standard document, the target information density, i.e., the average information entropy corresponding to each chapter, is calculated based on the target chapter information. This process can quantitatively assess the information richness and importance of each chapter, thus providing a scientific and reasonable reference for determining the overlap of target chapters. Then, each target chapter is segmented according to the overlap of target chapters, ensuring that the content of the segmented document chapters is complete and logically coherent. This avoids problems such as information loss or fragmentation of chapter content caused by blind segmentation, effectively improving the quality and efficiency of document segmentation. The segmented document is more in line with actual usage needs and facilitates subsequent document management and information retrieval operations.

[0023] It's important to note that before segmenting the document, the document type and chapter information need to be identified and extracted. Document types can be categorized as standard documents, basic documents, and very long documents, etc. Target chapter information includes structural information such as chapters, sections, and subsections within the document, which can be identified through document formatting, heading styles, etc.

[0024] The process of identifying the target document type includes: Natural Language Processing (NLP) techniques are used to determine the document type by combining document format features (such as page number, font size, heading style, etc.).

[0025] Document types are identified using a predefined rule base, which contains feature descriptions of different document types.

[0026] The process of extracting target chapter information includes: For documents that support structured formats (such as Word, PDF, etc.), document parsing tools (such as ApachePOI, PDFBox, etc.) can be used to extract the document's chapter information.

[0027] For plain text documents, chapter information can be extracted by matching heading patterns (such as "Chapter 1", "1.1", etc.) using regular expressions.

[0028] The extracted chapter information includes the chapter's start and end positions, as well as the chapter title.

[0029] For S2, the standard document has a clear chapter structure, and its information density can be calculated by analyzing the content of each chapter. Information density represents the average information entropy of each chapter, reflecting the complexity and richness of the chapter's content.

[0030] In another embodiment of this solution, the method further includes: text preprocessing, namely, cleaning the text content of each chapter to remove useless symbols and spaces; and segmenting the text to obtain word or phrase sequences.

[0031] In another embodiment of this solution, based on the target information density and combined with the preset correspondence between the number of chapters, information density, preset overlap ratio, and chapter overlap degree, the target chapter overlap degree corresponding to the target information density is found. If the target information density is high, it indicates that the chapter content is relatively complex, and the overlap degree can be appropriately increased to avoid semantic breaks.

[0032] Each target chapter is segmented based on its overlap with other target chapters.

[0033] When segmenting, the overlapping parts between adjacent chapters are retained, and the length of the overlapping parts is calculated based on the overlap degree of the target chapters.

[0034] For example, if the target chapter overlap is 20% and a chapter is 1000 words long, then the overlap between adjacent chapters is 200 words long.

[0035] After segmentation, multiple semantically coherent document fragments are obtained.

[0036] Furthermore, when the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

[0037] It should be noted that the document is divided into multiple segments based on the number of chapters in the target chapter information. Each segment corresponds to one chapter, and the segmentation is performed directly according to the start and end positions of the chapters.

[0038] For example, if a document has 5 chapters, then the document is directly divided into 5 segments.

[0039] Furthermore, when the target document type is an extra-long document, the document to be segmented is processed according to a preset fixed block size.

[0040] It should be noted that the document is divided into segments according to a fixed block size, starting from the beginning.

[0041] If the size of the last block is smaller than the fixed block size, the remaining part is treated as a separate block.

[0042] For example, if the total length of the document is 5000 words and the fixed block size is 1000 words, then the document will be divided into 5 blocks.

[0043] In another embodiment of this solution, when the target document type is an excessively long document, the document to be segmented can also be processed using the following methods, the specific process of which is as follows: First method: Layered processing Use document parsing tools (such as PDFBox, Apache POI, etc.) to extract the document's structural information, including chapter titles, sub-chaps, appendices, etc.

[0044] Natural Language Processing (NLP) techniques are used to identify the topics and modules within a document. Based on the extracted structural information (topics and modules), the document is divided into different levels of blocks: "chapter" level, "sub-chapter" level, and "paragraph" level. Each level of block is processed separately, including extracting key information and generating a summary. Based on the extracted key information and the generated summary, an index is generated for each level of block, resulting in a chapter index, a sub-chapter index, and a keyword index. The different levels of blocks are stored as independent files or data structures.

[0045] The second method is summary extraction and compression. Use natural language processing techniques (such as text summarization algorithms and keyword extraction algorithms) to extract key sentences and keywords from documents. This includes, but is not limited to, using the TextRank algorithm to extract important sentences from documents (important sentences are those whose keywords account for a preset percentage).

[0046] The extracted key information is combined into a summary document. The length of the summary can be adjusted as needed, for example, generating a summary that is 10% or 20% of the original document length. The summary document is then further compressed, for example, by removing redundant information and merging similar content. Text compression algorithms (such as LZ77 or Huffman coding) are used to encode the summary document to reduce storage space.

[0047] The generated summaries can be stored as separate files or data structures for easy browsing and retrieval.

[0048] The third method is parallel processing. Divide excessively long documents into multiple blocks of similar size, each of which can be processed independently. This includes, but is not limited to, dividing a document into multiple blocks based on a fixed number of pages or a fixed number of words.

[0049] Each block is assigned to a different processor or compute node for processing. Distributed computing frameworks (such as Apache Spark and Hadoop) are used to manage task allocation and execution.

[0050] Each processor or compute node processes its allocated blocks, performing tasks such as information extraction, index generation, and analysis. Multithreading or distributed computing techniques can be used to improve efficiency during processing. The processing results from all processors or compute nodes are then combined into a single, complete output.

[0051] The fourth type is semantic segmentation. Natural language processing techniques (such as BERT and GPT) are used to perform semantic analysis on the document. This involves identifying the topics, concepts, and semantic boundaries within the document. Based on the semantic analysis results, the document is divided into multiple semantically related blocks, i.e., into different topic paragraphs or concept modules. Each semantic block is then optimized, for example, by removing redundant information and merging similar content. This ensures that each block is semantically independent and complete. Each semantic block is stored as an independent file or data structure for easy subsequent processing.

[0052] Fifth, dynamic block division and adaptive processing Analyze the document's content features, such as information density, frequency of topic changes, and structural complexity. Extract feature information using natural language processing techniques. Dynamically adjust block sizes based on content features. This dynamic adjustment process can involve dividing high-information-density sections into smaller blocks and low-information-density sections into larger blocks. Use adaptive algorithms (such as the sliding window algorithm) to dynamically partition the document. Select an appropriate processing method based on the content features of each block. For example, perform detailed analysis on high-information-density blocks and perform simple processing on low-information-density blocks. Use machine learning models (such as decision trees and neural networks) to select processing strategies based on block features. Integrate the processing results of all blocks into a single, complete output.

[0053] Sixth, incremental processing The process divides extremely long documents into multiple blocks of similar size, each of which can be processed independently. For example, the document can be divided into blocks based on a fixed number of pages or words. One block is processed at a time, and the results are stored as intermediate results. These intermediate results are accumulated to gradually build the complete processing result. The processing strategy for subsequent blocks is dynamically adjusted based on problems or needs discovered during the process.

[0054] Seventh type: Interactive processing Before processing begins, user needs and preferences are gathered through the user interface. Users specify the types of information to be extracted, the level of detail required, etc. The document is initially processed based on user needs, generating preliminary results. These preliminary results are then presented to the user, and feedback is collected. The processing strategy is dynamically adjusted based on user feedback, and the document is reprocessed. Finally, the final processing result is generated based on the user's final requirements.

[0055] Eighth type, mixed treatment Analyze the characteristics and processing requirements of extremely long documents to determine the necessary combined processing methods. For an extremely long document containing numerous charts and technical details, semantic segmentation and parallel processing can be combined. Select appropriate processing methods based on requirements, including but not limited to: a combination of hierarchical processing, parallel processing, and summary extraction. Combine the selected processing methods and process the document step by step. Optimize the results of the combined processing to ensure accuracy and completeness. Generate the final processed result to meet the user's needs.

[0056] Furthermore, the method for determining the overlap of target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

[0057] It should be further explained that basic documents refer to documents with simple structure, single content, and short length (less than the first preset value). They usually do not have complex chapter divisions, the content is relatively direct, and the format is standardized and uniform. They are mainly used for the transmission or recording of simple information, such as personal resumes, brief meeting minutes, or project progress reports.

[0058] A standard document is a document with a standardized structure, relatively complex content, and a moderate length (the word count is between the first and second preset values). It has clear chapter divisions, such as preface, main body, conclusion, and appendix. The content includes various types of text, charts, formulas, and citations. The format follows certain standards and is often used in academic papers, corporate annual reports, or technical manuals.

[0059] Extra-long documents are documents that are extremely long (exceeding the third preset value) with very complex content and multiple structural levels. They are usually over 100 pages long, and can even reach hundreds of pages. They contain a large amount of text, charts, formulas, citations, appendices, etc. Structurally, they have multiple levels of chapters and indexes and need to follow strict format standards. They are suitable for complex scenarios such as academic works, large corporate reports, or technical standard documents.

[0060] Example 1: Different segmentation strategies are adopted for documents of different levels.

[0061] A. Basic Document (average word count per chapter is less than 2000 characters) Processing: Divide directly into chapters. Standard document (average word count per chapter is between 2000-5000) A dynamic segmentation method is adopted, specifically: (1) Calculate the initial ratio init_ratio The initial ratio `init_ratio` (default 0.6) is adjusted by calculating the average information entropy of each chapter, i.e., the information density of the text. The higher the information density, the higher this value. When the information density is less than 0.3, the init_ratio is set to 0.5. When the information density is between 0.3 and 0.8, the init_ratio is set to 0.6. When the information density is greater than 0.8, the init_ratio is set to 0.7. (2) An adaptive dynamic block partitioning algorithm based on information density is adopted: Where: Overlap_num represents the amount of overlap. This represents the length of a specific chapter after the document has been split into sections, where N represents the number of chapters the document has been split into. 'o' represents the initial ratio (calculated from information density), and 'overlap_ratio' represents the overlap ratio, which defaults to 20% (a 20% overlap ratio can cover the typical context span in natural language; experiments show that 15-25% overlap can repair more than 90% of the boundary information loss problem). Calculation example: 1. Dynamic Chunk_size calculation logic: The initial proportion of the document length is determined based on the average information density of the document chapters. For example, if the average information density is 0.7, then 0.6 is used as the initial proportion value.

[0062] 2. Overlap calculation: For example, if the average character length of the document after being split into chapters is 3000, then there are 3000 * 0.6 * 0.2 = 360 characters as overlapping characters. 3. Sliding window segmentation: Each time, retrieve the length of the text. Chunk_size= -overlap refers to the length of the chapter and the amount of overlap. like: Suppose a document has three chapters, with lengths of 2000, 3000, and 4000 words respectively. The dynamic Chunk_size should then be: First segment: 0~2000 Second block: 1640 (2000-360) ~ 4640 (2000+3000-360) (360 characters overlapping with the preceding and following blocks) Third block: 4280 (4640-360) ~ 9000 tokens C. Extremely long documents (average chapter length 5000+ words) Processing: Segment directly according to a fixed chunk size. Therefore, a multi-level dynamic segmentation strategy using an adaptive dynamic block algorithm is introduced to retain more related information, avoid excessive fragmentation, and improve semantic coherence. This effectively solves the problems of token waste and information loss in traditional document segmentation methods, improving document processing efficiency and quality.

[0063] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0064] This invention also provides a multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm, the specific technical solution of which is as follows: The acquisition module is used to: determine the target document type and target chapter information of the document to be segmented; The calculation module is used to: when the target document type is a standard document, calculate the target information density corresponding to the document to be segmented based on the target chapter information. The target information density represents the average information entropy corresponding to each chapter. The segmentation module is used to: determine the overlap of target chapters corresponding to the document to be segmented based on the target information density, and segment each target chapter according to the overlap of target chapters to complete the segmentation process of the document to be segmented.

[0065] Based on the above solution, the present invention can be further improved as follows.

[0066] Furthermore, when the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

[0067] Furthermore, when the target document type is an extra-long document, the document to be segmented is processed according to a preset fixed block size.

[0068] Furthermore, the method for determining the overlap of target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

[0069] It should be noted that the beneficial effects of the multi-level document segmentation system using an adaptive dynamic block-based algorithm provided in the above embodiments are the same as those of the multi-level document segmentation method using an adaptive dynamic block-based algorithm described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0070] like Figure 2 As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned methods. Specifically: The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The memories 310 store at least one computer program 330, which is loaded and executed by the processors 320 to enable the electronic device 300 to implement the multi-level document segmentation method using an adaptive dynamic block-based algorithm provided in the above embodiments. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. It may also include other components for implementing device functions, which will not be elaborated upon here.

[0071] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-described methods.

[0072] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0073] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the methods described above.

[0074] It should be noted that the terms "first" and "second" in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0075] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0076] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0077] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A multi-level document segmentation method employing an adaptive dynamic block-segmentation algorithm, characterized in that, include: Determine the target document type and target chapter information of the document to be segmented; When the target document type is a standard document, the target information density corresponding to the document to be segmented is calculated based on the target chapter information, and the target information density represents the average information entropy corresponding to each chapter. Based on the target information density, the overlap of the target chapters corresponding to the document to be segmented is determined. According to the overlap of the target chapters, each target chapter is segmented to complete the segmentation process of the document to be segmented.

2. The multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm according to claim 1, characterized in that, When the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

3. The multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm according to claim 1, characterized in that, When the target document type is an ultra-long document, the document to be segmented is segmented according to a preset fixed block size.

4. The multi-level document segmentation method using an adaptive dynamic block-segmentation algorithm according to claim 1, characterized in that, The method for determining the overlap of the target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

5. A multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm, characterized in that, include: The acquisition module is used to: determine the target document type and target chapter information of the document to be segmented; The calculation module is used to: when the target document type is a standard document, calculate the target information density corresponding to the document to be segmented based on the target chapter information, wherein the target information density represents the average information entropy corresponding to each chapter; The segmentation module is used to: determine the overlap of target chapters corresponding to the document to be segmented based on the target information density, and segment each target chapter according to the overlap of target chapters to complete the segmentation process of the document to be segmented.

6. A multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm according to claim 5, characterized in that, When the target document type is a basic document, it is directly segmented according to the number of chapters in the target chapter information.

7. A multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm according to claim 5, characterized in that, When the target document type is an ultra-long document, the document to be segmented is segmented according to a preset fixed block size.

8. A multi-level document segmentation system employing an adaptive dynamic block-segmentation algorithm according to claim 5, characterized in that, The method for determining the overlap of the target chapters is as follows: The target chapter overlap degree corresponding to the target information density is determined by the correspondence between the number of chapters, information density, preset overlap ratio and chapter overlap degree.

9. An electronic device, characterized in that, The electronic device includes a processor coupled to a memory storing at least one computer program, which is loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to perform the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and system for dividing txt document content chapters

    CN117473984A

  • Document segmentation processing method and device, computer equipment and readable storage medium

    CN119066037A

  • Information retrieval method and system based on RAG and large language model

    CN119597870A

  • Document classification system, method thereof, and apparatus for application

    KR1020140046556A

  • Processing heterogeneous content in language models

    WO2025125288A1

Cited By

  • Document fragmentation method, system and equipment based on semantic boundary and medium

    CN121659954A