A semantic boundary-based document shredding method, system, device, and medium

CN121659954BActive Publication Date: 2026-09-15WUHAN DAMENG DATABASE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511866854.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-09-15
Estimated Expiration
2045-12-11

AI Technical Summary

Technical Problem

[0004]然而,这些现有技术方案仍存在显著不足

Benefits of technology

[0016]The beneficial effects of this invention are as follows: The document segmentation method based on semantic boundaries provided by this invention firstly avoids segmentation at semantically disruptive locations such as within sentences or between words by using multi-level semantic boundary recognition and intelligent segmentation point selection based on weight calculation. This improves the semantic integrity of the document segments generated from the segmented document, ensuring that each document segment is a logically self-consistent semantic unit. Furthermore, through a context-preserving processing mechanism, key contextual information is analyzed, extracted, and embedded, allowing the segments to maintain logical connections with other parts of the original text while being used independently, effectively preserving the contextual coherence of the segmented content. Furthermore, by evaluating the quality of document segments, the system automatically diagnoses and adjusts strategy parameters for reprocessing when the quality is substandard, enabling the system's self-optimization capabilities. This ensures the high quality and stability of the segmentation results and allows the system to adapt to documents of different structures and domains (such as technical reports, academic papers, news, and code), without requiring manual reconfiguration of rules for each type. Furthermore, since the output fragments have extremely high semantic integrity and good contextual hints, they directly provide high-quality input for downstream applications such as large language model (LLM) processing, precise retrieval, and knowledge base construction. This can reduce the ambiguity of the model's understanding, improve the precision and recall of information retrieval, and significantly improve the overall effect and user experience of document-based AI applications, while reducing the cost of manual verification and post-processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659954B_ABST
    Figure CN121659954B_ABST
Patent Text Reader

Abstract

The application provides a semantic boundary-based document slicing method, system, device and medium, and belongs to the technical field of databases.The method comprises the following steps: identifying a plurality of semantic boundaries in a document to be processed; generating a plurality of candidate cutting points based on the semantic boundaries; calculating the semantic weight of each candidate cutting point based on the preset weight corresponding to the boundary type; selecting a target cutting point from the plurality of candidate cutting points based on the semantic weight; performing a cutting operation at the target cutting point and performing context maintenance processing on the text before and after the target cutting point to generate an independent document slice; and performing quality evaluation on the document slice, and if the slice is determined to be unqualified according to the quality evaluation result, adjusting the slicing parameters and re-performing the slicing processing based on the adjusted parameters. The application realizes intelligent, automatic and high-quality slicing of documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database-related technologies, specifically to a document fragmentation method, system, device, and medium based on semantic boundaries. Background Technology

[0002] In the digital information age, processing massive amounts of documents has become a fundamental task in fields such as artificial intelligence, information retrieval, and big data analysis. To adapt to limitations in computing resources and improve processing efficiency, large-scale documents typically need to be segmented into appropriately sized fragments for subsequent operations. Traditionally, document fragmentation relies on simple rules, such as mechanical division based on a fixed number of bytes, characters, or natural paragraphs.

[0003] To address these needs, existing technologies have proposed several document segmentation methods. The most direct is fixed-length segmentation, such as dividing the document evenly by a specified number of characters or bytes. To slightly improve semantic coherence, some methods introduce boundary recognition based on punctuation marks (such as periods and semicolons) or spaces. Furthermore, some research attempts to combine natural language processing (NLP) techniques, using simple sentence detection models to locate sentence ends as reference points for segmentation.

[0004] However, these existing technical solutions still have significant shortcomings. First, fixed segmentation methods completely ignore the inherent semantic structure of documents, easily breaking them apart in the middle of sentences or even words, severely damaging the integrity of semantic units. Second, even segmentation based on punctuation lacks consideration of deep semantic connections within the context, resulting in the loss of logical relationships before and after the segmentation point, creating information silos. Third, existing methods generally lack effective segmentation quality evaluation mechanisms, failing to quantify and provide feedback control over the quality of segmentation results. Finally, these methods employ simplistic strategies, lacking adaptability to the structural and semantic characteristics of different document types such as technical documents, news, and code, leading to poor versatility and unsatisfactory performance in diverse real-world application scenarios. Summary of the Invention

[0005] In view of this, it is necessary to provide a document segmentation method, system, device and medium based on semantic boundaries to solve the technical problem that existing document segmentation methods based on fixed length or simple punctuation will destroy the integrity of semantic structure and sever contextual associations.

[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a document fragmentation method based on semantic boundaries, comprising: Identify multi-level semantic boundaries in the document to be processed; Multiple candidate segmentation points are generated based on the semantic boundaries; Based on the preset weights corresponding to the boundary types, the semantic weights of each candidate segmentation point are calculated. Based on the semantic weights, a target segmentation point is selected from the plurality of candidate segmentation points; The splitting operation is performed at the target splitting point, and the text before and after the target splitting point is processed to preserve the context, so as to generate independent document fragments; The document fragments are evaluated for quality. If the fragments are deemed unqualified based on the evaluation results, the fragmentation parameters are adjusted and the fragmentation process is repeated based on the adjusted parameters.

[0007] In one possible implementation, identifying multi-level semantic boundaries in the document to be processed includes: Paragraph boundaries are determined by detecting paragraph identifiers, analyzing paragraph structure, and identifying paragraph topics. Sentence boundaries are determined through sentence structure analysis and punctuation detection at the end of sentences. Word boundaries are determined through word separator detection and word integrity checks. Character boundaries are determined through character type recognition and character encoding detection; The semantic boundaries include the paragraph boundaries, sentence boundaries, word boundaries, and character boundaries.

[0008] In one possible implementation, calculating the semantic weights of each candidate segmentation point based on a preset weight corresponding to the boundary type includes: Determine the corresponding preset weights based on the semantic boundary types of the candidate segmentation points; Obtain the recognition confidence of the semantic boundary type; Based on the preset weights and the recognition confidence level, the semantic weights are calculated according to the adjustment factor; the adjustment factor is determined according to the document type of the document to be processed.

[0009] In one possible implementation, selecting the target segmentation point from the plurality of candidate segmentation points based on the semantic weights includes: The optimal candidate segmentation point is subjected to quality verification; the optimal candidate segmentation point is the candidate segmentation point with the largest semantic weight, and the quality verification includes semantic integrity verification and contextual relevance verification after simulated segmentation; If the quality verification passes, the optimal candidate segmentation point is determined as the target segmentation point; If the quality verification fails, a suboptimal candidate segmentation point is selected for verification, or the segmentation parameters are adjusted and a new candidate segmentation point is regenerated until the target segmentation point is selected; the suboptimal candidate segmentation point is the candidate segmentation point with the second largest semantic weight.

[0010] In one possible implementation, the context-preserving processing of the text before and after the target segmentation point includes: Perform contextual semantic analysis on the text located within a first preset range before the target segmentation point to extract key information; Perform contextual semantic analysis on the text located within a second preset range after the target segmentation point to identify related information; Based on the analysis results of the key information and related information, one of the following is selected as the target context preservation strategy: complete retention strategy, key retention strategy, and necessary retention strategy; the context preservation range of the complete retention strategy is greater than that of the key retention strategy, and the context preservation range of the key retention strategy is greater than that of the necessary retention strategy. According to the target context preservation strategy, the corresponding range of context information is embedded into the segmented text generated by segmenting before and after the target segmentation point to obtain the document segment.

[0011] In one possible implementation, the quality assessment of the document fragments includes: The document fragments are subjected to semantic integrity assessment to obtain a first score; A second score is obtained by evaluating the contextual association of the document fragments; A third score is obtained by performing a structural integrity assessment on the document fragments; The overall quality score of the document segment is calculated by weighting the first score, the second score, and the third score. The overall quality score is compared with a first quality threshold and a second quality threshold; the first quality threshold is greater than the second quality threshold. If the overall quality score is greater than or equal to the first quality threshold, the document segmentation is determined to be qualified and the segmentation result is output. Then, the next target segmentation point is selected for segmentation. If the overall quality score is less than the second quality threshold, the document segmentation is determined to be unqualified, the segmentation parameters are adjusted, and the segmentation process is re-performed based on the adjusted parameters.

[0012] In one possible implementation, prior to identifying the multi-level semantic boundaries in the document to be processed, the following steps are included: Identify the document type of the document to be processed; The document to be processed is preprocessed according to the document type.

[0013] Secondly, the present invention also provides a document fragmentation system based on semantic boundaries, comprising: The input processing module is used to identify multi-level semantic boundaries in the document to be processed; The intelligent segmentation module is used to generate multiple candidate segmentation points based on the semantic boundary, calculate the semantic weight of each candidate segmentation point based on the preset weight corresponding to the boundary type, select a target segmentation point from the multiple candidate segmentation points based on the semantic weight, and perform a segmentation operation at the target segmentation point, and perform context-preserving processing on the text before and after the target segmentation point to generate independent document segments. The quality assessment module is used to assess the quality of the document segments. If the segment is determined to be unqualified based on the quality assessment results, the segmentation parameters are adjusted and the segmentation process is repeated based on the adjusted parameters.

[0014] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the semantic boundary-based document fragmentation method described in any of the above implementations.

[0015] Fourthly, the present invention also provides a computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, can implement the steps in the semantic boundary-based document fragmentation method described in any of the above implementations.

[0016] The beneficial effects of this invention are as follows: The document segmentation method based on semantic boundaries provided by this invention firstly avoids segmentation at semantically disruptive locations such as within sentences or between words by using multi-level semantic boundary recognition and intelligent segmentation point selection based on weight calculation. This improves the semantic integrity of the document segments generated from the segmented document, ensuring that each document segment is a logically self-consistent semantic unit. Furthermore, through a context-preserving processing mechanism, key contextual information is analyzed, extracted, and embedded, allowing the segments to maintain logical connections with other parts of the original text while being used independently, effectively preserving the contextual coherence of the segmented content. Furthermore, by evaluating the quality of document segments, the system automatically diagnoses and adjusts strategy parameters for reprocessing when the quality is substandard, enabling the system's self-optimization capabilities. This ensures the high quality and stability of the segmentation results and allows the system to adapt to documents of different structures and domains (such as technical reports, academic papers, news, and code), without requiring manual reconfiguration of rules for each type. Furthermore, since the output fragments have extremely high semantic integrity and good contextual hints, they directly provide high-quality input for downstream applications such as large language model (LLM) processing, precise retrieval, and knowledge base construction. This can reduce the ambiguity of the model's understanding, improve the precision and recall of information retrieval, and significantly improve the overall effect and user experience of document-based AI applications, while reducing the cost of manual verification and post-processing. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic flowchart of an embodiment of the document fragmentation method based on semantic boundaries provided by the present invention; Figure 2 For the present invention Figure 1 A schematic diagram of an embodiment of S101; Figure 3 For the present invention Figure 1 A schematic diagram of an embodiment of S103; Figure 4 For the present invention Figure 1 A schematic diagram of an embodiment of S104; Figure 5 For the present invention Figure 1 A schematic diagram of an embodiment of S105; Figure 6 For the present invention Figure 1A schematic diagram of an embodiment of S106; Figure 7 For the present invention Figure 1 A flowchart of an embodiment prior to S101; Figure 8 A schematic diagram of an embodiment of the document fragmentation system based on semantic boundaries provided by the present invention; Figure 9 A schematic diagram of an embodiment of the electronic device provided by the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0021] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.

[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0023] This invention provides a document fragmentation method, system, device, and medium based on semantic boundaries, which are described below.

[0024] Figure 1 This is a schematic diagram of an embodiment of the document fragmentation method based on semantic boundaries provided by the present invention, as shown below. Figure 1 As shown, document fragmentation methods based on semantic boundaries include: S101. Identify the multi-level semantic boundaries in the document to be processed.

[0025] It's important to clarify that a semantic unit refers to the smallest linguistic fragment or structural block in a document text that can carry a relatively independent and complete meaning; it is the object protected and separated by semantic boundaries. Examples include a complete argument, a complete sentence, or an independent conceptual word. A semantic boundary refers to the naturally occurring dividing position or transition point between semantic units in a document text; in short, a semantic boundary is the natural separating line between semantic units. Examples include the end of a paragraph, a period, and spaces between words. Identifying multi-level semantic boundaries in a document involves dividing and identifying semantic boundaries of different granularities (coverage size and importance) according to the document's content or the semantic integrity of the text.

[0026] S102. Generate multiple candidate segmentation points based on the semantic boundary.

[0027] It should be noted that the system aggregates all semantic boundaries detected and identified by S101, and adds them to the candidate segmentation point set as candidate segmentation points. For example, the end of a paragraph, because it is simultaneously a paragraph boundary, sentence boundary, and word boundary, will become a high-priority candidate segmentation point. The system records the corresponding boundary type and detection confidence score for each candidate segmentation point.

[0028] S103. Calculate the semantic weight of each candidate segmentation point based on the preset weight corresponding to the boundary type.

[0029] It should be noted that different semantic boundary types correspond to different preset weights. For example, the preset weight for a paragraph boundary is 0.9, for a sentence boundary it is 0.7, for a word boundary it is 0.4, and for a character boundary it is 0.1. For each candidate segmentation point, its corresponding semantic weight is calculated based on the preset weight corresponding to the boundary type of that candidate segmentation point.

[0030] S104. Based on the semantic weights, select the target segmentation point from the plurality of candidate segmentation points.

[0031] It should be noted that, based on the semantic weights corresponding to each candidate segmentation point, the target segmentation point is selected from the plurality of candidate segmentation points according to steps S401 to S403 of the following embodiments.

[0032] S105. Perform a segmentation operation at the target segmentation point and perform context-preserving processing on the text before and after the target segmentation point to generate independent document fragments.

[0033] It should be noted that the segmentation operation is performed at the target segmentation point to divide the document into two parts. Semantic analysis is then performed on the text before and after the segmentation point, specifically analyzing the preceding and following text (identifying related references and logical connections). Based on the analysis results, selected contextual information (such as a summary sentence or a "therefore" conjunction) is seamlessly embedded into the generated segmented text content (e.g., appending a key sentence from the preceding text to the beginning of the following segment), ensuring that each document segment is semantically consistent and naturally connected when read independently.

[0034] S106. Perform a quality assessment on the document segments. If the segment is determined to be unqualified based on the quality assessment results, adjust the segmentation parameters and re-process the segmentation based on the adjusted parameters.

[0035] It should be noted that: for each generated document fragment, a quality assessment is performed. If the quality assessment result determines that the fragment is qualified, the fragmentation result is output and the fragmentation of the target split point is completed. If the quality assessment result determines that the fragment is unqualified, the parameters are adjusted, the semantic weights are recalculated, and S103 to S106 are re-executed.

[0036] In summary, the document segmentation method based on semantic boundaries provided in this invention firstly avoids segmentation at semantically disruptive locations such as within sentences or between words by using multi-level semantic boundary recognition and intelligent segmentation point selection based on weight calculation. This improves the semantic integrity of the document segments generated from the segmented document, ensuring that each segment is a logically self-consistent semantic unit. Furthermore, through a context-preserving processing mechanism, key contextual information is analyzed, extracted, and embedded, allowing the segments to maintain logical connections with other parts of the original text while being used independently, effectively preserving the contextual coherence of the segmented content. Moreover, by evaluating the quality of document segments, the system automatically diagnoses and adjusts strategy parameters for reprocessing when quality is substandard, enabling the system's self-optimization capabilities. This ensures high quality and stability of the segmentation results and allows the system to adapt to documents with different structures and domains (such as technical reports, academic papers, news, and code), eliminating the need for manual rule reconfiguration for each type. Furthermore, since the output fragments have extremely high semantic integrity and good contextual hints, they directly provide high-quality input for downstream applications such as large language model (LLM) processing, precise retrieval, and knowledge base construction. This can reduce the ambiguity of the model's understanding, improve the precision and recall of information retrieval, and significantly improve the overall effect and user experience of document-based AI applications, while reducing the cost of manual verification and post-processing.

[0037] To comprehensively and accurately identify semantic boundaries at all levels within a document, from macroscopic structure to microscopic units, and to avoid the core defects of incomplete and inaccurate boundary identification, in some embodiments of the present invention, such as... Figure 2 As shown, step S101 includes: S201. Determine paragraph boundaries by detecting paragraph identifiers, analyzing paragraph structure, and identifying paragraph topics.

[0038] It should be noted that the process involves scanning the format and symbol features of the document to be processed, and identifying explicit paragraph separators for paragraph identifier detection. The detected paragraph identifiers include newline character combinations (such as two consecutive newlines), first-line indentation symbols, and specific paragraph numbers (such as "1.", "(a)").

[0039] The content blocks of suspected paragraphs are structurally analyzed. The paragraph structure analysis mainly includes examining the paragraph length (abnormally short paragraphs may be titles or list items), detecting the characteristics of the first and last sentences (to determine whether they are topic sentences or summary sentences), and detecting the coherence between sentences (by analyzing the word vector similarity between sentences to determine whether a content block is topic-focused).

[0040] By using topic modeling (such as LDA) or keyword extraction techniques (such as TF-IDF), each candidate paragraph is analyzed to identify its core keywords. When the core keyword sets of adjacent paragraphs change significantly, their intersection is strongly identified as a paragraph topic boundary. For example, transitioning from a paragraph discussing "system architecture" to a paragraph discussing "technology selection".

[0041] S202. Sentence boundaries are determined through sentence structure analysis and punctuation detection at the end of sentences.

[0042] It's important to note that using Natural Language Processing (NLP) tools to analyze the sentence structure of the document being processed involves parsing the sentence's syntax tree to accurately determine the subject-verb-object structure and the nesting relationships of clauses. Simultaneously, punctuation marks such as periods, question marks, and exclamation marks are used, while eliminating interference from abbreviations (e.g., "Fig.") and numbering (e.g., "1.") to determine sentence boundaries. Alternatively, NLP parsers (such as dependency parsing) can be used to analyze the text, determining the true end of a grammatically complete sentence by analyzing subject-verb-object structures, the start and end positions of clauses (e.g., clauses introduced by subordinating conjunctions), and the pairing relationships of quotation marks and parentheses. Detecting sentence boundaries using final punctuation marks such as periods, question marks, and exclamation marks is also possible. However, it's crucial to consider the context and structural analysis to eliminate interference when determining sentence boundaries based on final punctuation marks. For example, this involves distinguishing between periods in English abbreviations ("Dr."), periods in numbering ("1."), and true sentence-ending marks.

[0043] S203. Determine word boundaries through word separator detection and word integrity check.

[0044] It should be noted that for each potential lexical boundary identified by delimiters in the document to be processed, a lexical integrity check is performed. This involves using a dictionary or lexical analysis model to verify whether potential segmentation points would disrupt the integrity of the vocabulary. For example, the system will recognize the compound word "state-of-the-art" or the variable name userAuthenticationToken as a complete semantic unit, avoiding incorrect segmentation in the middle. The system can also identify common lexical delimiters such as spaces and specific punctuation marks (e.g., commas, semicolons, and pauses) in the document text to determine lexical boundaries.

[0045] S204. Determine character boundaries through character type recognition and character encoding detection; The semantic boundaries include the paragraph boundaries, sentence boundaries, word boundaries, and character boundaries.

[0046] It should be noted that the function identifies the character types in the document text to be processed, such as distinguishing between letters, numbers, punctuation marks, ideographic characters (such as Chinese characters), and control characters. It can also detect or confirm the character encoding format of the document (such as UTF-8, GBK, ISO-8859-1). During the segmentation operation, especially when processing by byte position, ensure that the segmentation point falls on the legal character boundaries specified by the encoding, avoiding cutting a multi-byte character (such as a UTF-8 encoded Chinese character) in the middle, which would result in garbled text.

[0047] In this embodiment, boundary recognition is performed through four key levels: paragraph, sentence, word, and character. This improves the completeness and accuracy of boundary recognition, enabling the system to not only recognize explicit format boundaries but also understand implicit grammatical and semantic boundaries. This fundamentally solves the problem of missed and incorrect boundary recognition, laying the foundation for obtaining high-quality document fragments by segmenting them at the correct locations.

[0048] To accurately measure the merits of different candidate segmentation points in maintaining semantic integrity, and thus intelligently select the optimal target segmentation point to improve the quality of subsequent document segmentation, in some embodiments of the present invention, such as... Figure 3 As shown, step S103 includes: S301. Determine the corresponding preset weights based on the semantic boundary type of the candidate segmentation points.

[0049] It should be noted that the system internally maintains a semantic boundary importance lookup table, which stores the preset weights corresponding to different boundary types. These preset weights are set based on their inherent ability to maintain semantic integrity. For example, paragraph boundaries have the highest preset weight (e.g., 0.9), where segmentation preserves a complete logical or narrative unit. Sentence boundaries have the second highest preset weight (e.g., 0.7), where segmentation preserves a complete statement or expression. Lexical boundaries have a medium preset weight (e.g., 0.4), where segmentation only ensures word integrity but disrupts sentence fluency. Character boundaries have the lowest preset weight (e.g., 0.1), where segmentation only ensures correct encoding. When calculating the semantic weight for a candidate segmentation point, the system first queries all identified boundary types for that candidate segmentation point and then reads the corresponding preset weights from the semantic boundary importance lookup table.

[0050] S302. Obtain the recognition confidence of the semantic boundary type.

[0051] It should be noted that the recognition confidence score represents the degree of confidence that the algorithm believes the candidate boundary point is indeed the claimed boundary type. The recognition confidence score is usually a value between 0.0 and 1.0.

[0052] Paragraph boundary identification typically combines both formatting and semantic cues, and its identification confidence level is derived accordingly. Formatting detection relies on obvious formatting markers, such as blank lines, first-line indentation, and specific heading styles (e.g., #, ...). <h1>The closer the match to obvious formatting markers, the higher the confidence level of formatting recognition for paragraph boundaries. The matching score can be mapped to a formatting matching confidence level. Semantic model-based recognition confidence typically uses topic models (such as LDA) or text embedding models to calculate the topic similarity or semantic vector cosine similarity between adjacent text blocks (e.g., the first three sentences and the last three sentences). Lower similarity indicates a more obvious topic shift, resulting in higher semantic recognition confidence for paragraph boundaries. The similarity value can be mapped to a semantic recognition confidence level. The final paragraph boundary recognition confidence level is obtained by weighted averaging of the formatting recognition confidence level and the semantic recognition confidence level.

[0053] Sentence boundary identification requires handling complex punctuation and grammar, necessitating more refined confidence assessment. One approach is to use rule-based models based on punctuation and context to determine sentence boundary identification confidence. This involves using regular expressions or predefined rule sets in conjunction with the context. For example, when encountering a period ".", check if the preceding word is a known abbreviation (e.g., "Dr.", "eg") and if the following word is uppercase (usually indicating the start of a new sentence in English). A high confidence score (e.g., 1.0) is assigned when the rule is a perfect match (e.g., "!", "?" followed by a space and an uppercase letter). A lower confidence score (e.g., 0.3-0.6) is assigned in ambiguous cases (e.g., a period followed by a number or lowercase letter). Alternatively, a machine learning model based on sequence labeling can be used to obtain sentence boundary identification confidence. This involves employing models such as Conditional Random Fields (CRF) or Bidirectional Long Short-Term Memory Networks (Bi-LSTM) + CRF, treating sentence boundary detection as a sequence labeling task (labeling each character / word as "boundary" or "non-boundary"). These types of models typically output the predicted probability for each label. For a location predicted as a "sentence boundary," its predicted probability value can be directly used as the confidence score for sentence boundary identification. Alternatively, sentence boundary identification confidence can be obtained based on a pre-trained language model, such as BERT, which performs mask prediction at positions like periods or utilizes its next-sentence prediction (NSP) capability. For example, by inserting a [SEP] marker at a suspected sentence end, the model can determine whether the preceding and following parts form a coherent sentence; the model's output probability of incoherence can be used as the confidence score for identifying that location as a true sentence boundary.

[0054] Lexical boundaries primarily ensure that words are not segmented within their own words. The confidence level for lexical boundary recognition is typically based on validation results. One approach is to use a dictionary / dictionary validation method. This involves initial separation using spaces or punctuation, followed by querying a pre-defined general dictionary or domain-specific terminology dictionary (such as a technical terminology database). If a match is found in the dictionary, the lexical boundary is assigned a high confidence level (e.g., 1.0). For words that do not match in the dictionary, a moderate confidence level (e.g., 0.7) may be assigned based on word formation rules. Alternatively, word segmentation models can be used. For languages ​​like Chinese without natural spaces, or when dealing with unknown words, word segmenters such as BPE or WordPiece can be used. These segmenters provide their own segmentation schemes. A boundary representing a complete word segmented by a word segmenter generally has a high confidence level.

[0055] S303. Based on the preset weight and the recognition confidence, the semantic weight is calculated according to the adjustment factor; the adjustment factor is determined according to the document type of the document to be processed.

[0056] It should be noted that the semantic weight is calculated by substituting the preset weights and recognition confidence levels corresponding to different boundary types, combined with the adjustment factor, into the following formula: Semantic weight = Σ(preset weight × recognition confidence) + adjustment factor, where i is the label of the boundary type.

[0057] The preset weights are pre-set according to the boundary type. The preset weight of paragraph boundaries is greater than that of sentence boundaries, the preset weight of sentence boundaries is greater than that of word boundaries, and the preset weight of word boundaries is greater than that of character boundaries.

[0058] For a candidate segmentation point, it may be identified as multiple boundary types simultaneously (such as both the end of a paragraph and the end of a sentence). The products of (preset weight × recognition confidence) of all identified boundary types are summed. In this way, the more important the boundary detectors that confirm a candidate segmentation point with high recognition confidence, the higher its semantic weight.

[0059] The adjustment factor is dynamically determined based on the document type. In the initial stage, the system can identify the document type of the document to be processed (such as technical documents, news, legal documents, and code), and each document type corresponds to a fine-tuning tendency for various boundary weights. Of course, the adjustment factor can also be dynamically determined based on the results of contextual correlation analysis.

[0060] For example, for a document to be processed that is a technical document, the paragraph boundary adjustment factor is set to +0.05. Assuming that the candidate segmentation points in the technical document are identified as paragraph boundaries (preset weight 0.9, recognition confidence 1.0) and sentence boundaries (preset weight 0.7, recognition confidence 1.0), the calculated semantic weight is: (0.9×1.0)+(0.7×1.0)+0.05=1.65.

[0061] In this embodiment, semantic weights calculated using preset weights, identification confidence levels, and adjustment factors transform the fuzzy concept of semantic integrity into a precisely calculable and comparable scalar score, making the subsequent selection of the optimal segmentation point fully automated and objective. Furthermore, the semantic weights calculated by comprehensively considering multiple analytical dimensions such as paragraphs, sentences, and words improve fault tolerance and robustness. Moreover, by using dynamically changing adjustment factors based on document type, semantic weights can be adaptively calculated according to document type, allowing for adaptation to different document types and producing high-quality segmentation results in various scenarios, greatly enhancing the method's versatility and practicality. Furthermore, the calculated semantic weights can not only be used for ranking, but their composition (which boundaries contributed how many points, and why adjustments were made) is traceable, providing a clear and interpretable data foundation for system debugging, effect analysis, and subsequent parameter adjustment feedback loops. This makes the entire intelligent segmentation system a transparent and optimizable closed loop, providing an interpretable decision-making basis for global optimization.

[0062] To ensure that the final selected target segmentation points can withstand scrutiny in actual segmentation results and improve the quality and reliability of the fragmentation results, in some embodiments of the present invention, such as... Figure 4 As shown, step S104 includes: S401. Perform quality verification on the optimal candidate segmentation point; the optimal candidate segmentation point is the candidate segmentation point with the largest semantic weight, and the quality verification includes semantic integrity verification and contextual relevance verification after simulated segmentation.

[0063] It should be noted that the system does not perform actual segmentation. Instead, it uses the optimal candidate segmentation point as the hypothetical segmentation location and creates two virtual text segments in memory: a pre-segment (from the beginning of the document to this point) and a post-segment (from this point to the end of the document). Both virtual segments are then quickly scanned for semantic integrity verification. Semantic integrity verification includes boundary integrity checks and coherence assessment. Boundary integrity checks verify whether the end of the pre-segment terminates at a complete sentence, paragraph, or topic unit, and whether the beginning of the post-segment starts at a complete sentence or paragraph. Special attention is paid to checking for incomplete sentences (e.g., ending with conjunctions like "because" or "although"), broken words, or unclosed parentheses or quotation marks. Coherence assessment uses a lightweight topic model or keyword extraction to determine whether the topics within each virtual segment are focused and whether there are abrupt or unfinished semantic jumps.

[0064] Contextual relevance verification is performed on the preceding and following virtual segments. This verification includes semantic similarity checking and logical analysis. Semantic similarity checking involves calculating the cosine similarity between the semantic vectors of the last part of the preceding virtual segment (e.g., the last 2-3 sentences) and the beginning part of the following virtual segment (e.g., the first 2-3 sentences). If the similarity is below a threshold, it indicates a potentially abrupt semantic connection. Logical analysis includes checking if the beginning of the following virtual segment contains pronouns ("this," "that," "it") or summarizing phrases ("the aforementioned issue," "based on this") that refer to the end of the preceding virtual segment, and confirming whether their reference to the end of the preceding virtual segment is explicit. Simultaneously, the use of logical connectors ("therefore," "however," "then") is analyzed for appropriateness.

[0065] S402. If the quality verification passes, the optimal candidate split point is determined as the target split point.

[0066] It should be noted that the quality verification is considered passed only when both semantic integrity verification and contextual relevance verification results meet the preset strict standards (e.g., no broken sentences or words, coherent theme, clear referents, and logically sound). Once the quality verification is passed, the system officially determines the optimal candidate segmentation point as the final target segmentation point and passes it to subsequent steps for actual segmentation.

[0067] S403. If the quality verification fails, select a suboptimal candidate segmentation point for verification, or adjust the segmentation parameters and regenerate a new candidate segmentation point until the target segmentation point is selected; the suboptimal candidate segmentation point is the candidate segmentation point with the second largest semantic weight.

[0068] It should be noted that if either the semantic integrity verification or the contextual relevance verification fails to meet the preset strict criteria, the optimal candidate segmentation point verification is deemed to have failed. The system does not immediately declare failure but instead activates a fallback plan: selecting the second-best candidate segmentation point with the second-highest semantic weight and performing the same quality verification process as in S401. If the second-best candidate segmentation point verification still fails, or if verification fails consecutively a certain number of times, it indicates that under the current parameter configuration, the entire candidate segmentation point set may not be able to produce qualified segments. At this point, the system will trigger a parameter adjustment mechanism. The adjustments may include increasing the confidence threshold for boundary detection, modifying the adjustment factor in the semantic weight calculation in S303, and relaxing the tolerance for segment size. After adjusting the parameters, the system will return to step S102 or S103, regenerate the candidate segmentation point set based on the new parameters, calculate the semantic weights, and then re-enter the verification loop from S401 to S403 until a target segmentation point that can pass the strict quality verification is selected.

[0069] In this embodiment, by introducing simulated segmentation and quality verification of the optimal candidate segmentation point, the system effectively prevents theoretically high-scoring but poorly performing segmentation (e.g., segmenting in the middle of a strong logical chain) candidate segmentation point from being mistakenly selected as the best choice, greatly reducing the risk of erroneous segmentation. Furthermore, when the quality verification of the optimal candidate segmentation point fails, a suboptimal candidate segmentation point is selected for quality verification, attempting the optimal alternative at extremely low cost, significantly improving the success rate of the first attempt and enhancing the system's robustness and practicality when dealing with complex and unconventional documents. Moreover, when the quality verification of the suboptimal candidate segmentation point still fails, parameter adjustments are triggered and candidate segmentation points are regenerated, enabling the entire system to possess self-diagnostic and evolutionary capabilities, allowing it to adapt to documents of varying difficulty and style. Furthermore, since the final selected target segmentation point is the only segmentation position that simultaneously passes both theoretical calculations (optimal) and simulated verification (qualified), the quality of the generated document fragments is greatly improved.

[0070] To ensure that each document fragment is semantically independent and complete while maintaining the overall coherence of the document fragments to the greatest extent possible, in some embodiments of the present invention, such as... Figure 5 As shown, step S105 includes: S501. Perform semantic analysis on the text within a first preset range before the target segmentation point to extract key information.

[0071] It should be noted that the first preset range is typically 1 to 3 paragraphs, or the first N sentences, preceding the target segmentation point. This first preset range can be dynamically adjusted based on the average paragraph length of the document. The system uses natural language processing technology to perform in-depth analysis of the text within the first preset range preceding the target segmentation point, extracting key information. Algorithms such as TF-IDF and TextRank can be used to identify and extract keywords and high-weight sentences (usually the first, last, or transitional sentences) representing the core theme of this range. Additionally, it identifies and extracts key entities such as names of people, organizations, locations, and technical terms. Through sentence templates (such as "...refers to...", "In summary...") or semantic role labeling, it locates and extracts core definitions, conclusions, or summarizing statements to obtain key sentences. Finally, it outputs a structured list of key information, including the core theme, key entities, and key sentences.

[0072] S502. Perform contextual semantic analysis on the text located within a second preset range after the target segmentation point to identify related information.

[0073] It should be noted that the second preset range is typically one to two paragraphs, or M sentences, following the target segmentation point. This second preset range can be dynamically adjusted based on the average paragraph length of the document. The system uses natural language processing technology to perform in-depth analysis of the text within the second preset range following the target segmentation point, identifying and obtaining related information. It can identify pronouns (such as "it," "this," "the") or noun phrases (such as "the above method," "this system") appearing in the following text, and clarify which specific entity or concept in the preceding text they refer to, thus establishing a referential relationship. It detects and analyzes logical connectors used at the beginning of the following text (such as "therefore," "however," "for example," "specifically"), clarifying the causal, adversative, or exemplifying relationships they express, and determining whether their logical premise is connected by a logical connector in the preceding text. It calculates the similarity between the keywords in the following text and the keywords in the preceding text (especially the ending section), determining whether the topic is a continuation, a transition, or a new topic, thus obtaining semantic continuity information. Finally, it outputs a structured list of related information, including referential relationships, logical connectors, and semantic continuity information.

[0074] S503. Based on the analysis results of the key information and related information, select one of the following as the target context preservation strategy: the complete retention strategy, the key retention strategy, and the necessary retention strategy; the context preservation range of the complete retention strategy is greater than that of the key retention strategy, and the context preservation range of the key retention strategy is greater than that of the necessary retention strategy.

[0075] It should be noted that the density of key information measures the concentration of core semantic content within a pre-defined range of text before the segmentation point (e.g., the first 1-3 paragraphs). Higher density indicates a greater and more concentrated amount of information in the preceding text that is crucial for understanding the following text. The process of calculating the density of key information is as follows: For the text within the first pre-defined range before the target segmentation point (hereinafter referred to as the preceding text range), a candidate list of key information is generated, where each key information item has its own importance weight. Key information is extracted using TF-IDF or TextRank algorithms, with the weights directly provided by the algorithm (e.g., TF-IDF values). The weights of all extracted key information items within the preceding text range are then summed using a weighted average. Named entities of different types are identified using NER (e.g., technical terms have higher weights than common names). Key sentences are identified by position (first sentence, last sentence), sentence structure (summary sentence, definition sentence), or summarization algorithms (e.g., LexRank) and assigned weights. The normalization function is set as Sum_keyweight = Σ(weight of each key information item). The density of key information is calculated as (Sum_key weight) / (text length or total number of units in the preceding range). The result of this calculation is then mapped to a standard range (e.g., between 0 and 1.0) to determine the final density of the key information.

[0076] For example, suppose there are 5 sentences in the preceding text, from which 3 key pieces of information are extracted: a core term (weight 0.8), a defining sentence (weight 0.9), and a concluding sentence (weight 0.7). The numerator would then be approximately 0.8 + 0.9 + 0.7 = 2.4. If the number of sentences is used as the denominator, the density would be 2.4 / 5 = 0.48. This value will be compared with a preset first density threshold (e.g., 0.7) and a second density threshold (e.g., 0.3).

[0077] The strength of relevance information measures the tightness of semantic dependencies and logical connections between the text following the segmentation point (e.g., the last 1-2 sentences / paragraphs) and the text preceding the segmentation point. Higher strength indicates a greater risk of disruption to coherence caused by the segmentation. The process of calculating the strength of relevance information involves calculating one or more feature values ​​and then combining them into a strength score. Feature values ​​include the strength of referential dependency, the strength of logical connection, and the strength of semantic continuation. The strength of relevance information is calculated by weighting the above features (referential dependency strength S_coref, logical connection strength S_logic, and semantic continuation strength S_sim) to obtain the strength of relevance information = Wt1×S_coref + Wt2×S_logic + Wt3×S_sim. Here, Wt1, Wt2, and Wt3 are weights, which can be adjusted according to the document type. The output strength of relevance information is also normalized to the range of 0-1. The process of calculating the strength of referential dependency involves identifying referential terms (pronouns, anaphoric nouns) at the beginning of the following text. Through referential resolution, the success rate and distance of these referencing specific entities in the preceding text are confirmed. Successful resolution and close referential distance (e.g., the previous sentence) result in high strength (e.g., 0.9); failure to resolve results in a strength of 0. The process of calculating logical connection strength involves detecting logical connectors at the beginning of the following text ("therefore," "however," "for example"). Preset strength values ​​are set based on the logical components of the logical connectors (e.g., "therefore" indicates strong causality, strength 0.8; "for example" indicates weak exemplification, strength 0.4; no connector results in 0). The process of calculating semantic continuation strength involves calculating the cosine similarity of the semantic vectors of the beginning of the following text and the end of the preceding text. The semantic vector cosine similarity itself (0~1) can be directly used as a strength feature.

[0078] For example, the following text begins with "Therefore, the architecture is highly scalable." "The architecture" explicitly refers to the preceding "layered architecture," so S_coref = 0.9. The conjunction "therefore" emphasizes causality, so S_logic = 0.8. The semantic similarity to the ending of the preceding text is calculated, so S_sim = 0.75. Assuming weighted average, the strength of the related information = (0.9 + 0.8 + 0.75) / 3 ≈ 0.82. This value will be compared with a preset first strength threshold (e.g., 0.8) and a second strength threshold (e.g., 0.4).

[0079] A first density threshold, a second density threshold, a first intensity threshold, and a second intensity threshold are pre-defined, with the first density threshold > the second density threshold and the first intensity threshold > the second intensity threshold. If the density of key information is greater than or equal to the first density threshold, and the intensity of related information is greater than or equal to the first intensity threshold, then the full context preservation strategy is selected. If the density of key information is greater than or equal to the second density threshold but less than the first density threshold, or the intensity of related information is greater than or equal to the second intensity threshold but less than the second intensity threshold, then the key context preservation strategy is selected. If the density of key information is less than the second density threshold, and the intensity of related information is less than the second intensity threshold, then the necessary context preservation strategy is selected.

[0080] S504. According to the target context preservation strategy, the context information of the corresponding range is embedded into the segmented text generated by segmenting before and after the target segmentation point to obtain the document segment.

[0081] It should be noted that: when the density of key information is extremely high (e.g., containing an indivisible core theorem or complete case) and the correlation information is extremely strong (existing tight causal chains or continuous step descriptions), i.e., the density of key information ≥ the first density threshold and the strength of correlation information ≥ the first strength threshold, the complete retention strategy is triggered, and the system will retain a large range of original text before and after the segmentation point. When there is a clear summary sentence in the key information list, or a clear referent or logical connection in the correlation information list (i.e., the second density threshold ≤ the density of key information < the first density threshold, and the second strength threshold ≤ the strength of correlation information < the first strength threshold), the key retention strategy is triggered, and the system only precisely embeds the extracted key sentences or related phrases. When the analysis results of key information and correlation information show that the semantics of the preceding and following text are relatively independent and the correlation is weak (i.e., the density of key information < the second density threshold and the strength of correlation information < the second strength threshold), the necessary retention strategy is triggered, and the system only embeds a minimum number of transition words or repeated core nouns. In summary, if a key retention strategy is adopted, the extracted "key information" (such as the concluding sentence of the preceding text) will be appended to the beginning of the next segment; or the identified "related information" (such as referential phrases in the following text) will be inserted to the end of the previous segment. If a complete retention strategy is adopted, a complete paragraph before the segmentation point may be appended to the next segment as an introduction, or a complete paragraph after the segment may be placed before the previous segment as an appendix. If a necessary retention strategy is adopted, phrases such as "in addition," "this part..." will only be added at the segment boundaries. After the embedding operation is completed, the system will output the adjusted text content as the final deliverable document segments. Each segment contains the necessary context to make its own semantics more complete and its logic more coherent with adjacent segments.

[0082] In this embodiment, quantifiable thresholds are set based on the density of key information and the strength of related information. By comparing the analysis results of S501 and S502 with these thresholds, a preset decision tree or rule engine is executed to automatically select the most economical and effective strategy. If the analysis results show that the semantic relationship before and after the segmentation point is extremely close or the density of key information is extremely high, a full context preservation strategy is selected to embed a large range of original text content in the segment. If the analysis results show that the key information is clear or there is specific related information, a key context preservation strategy is selected to embed core sentences or phrases directly related to the key information or related information in the segment. If the analysis results show that the semantics of the text before and after the segmentation point are relatively independent, a necessary context preservation strategy is selected to embed a minimum number of words or conjunctions in the segment to maintain basic coherence. Through bidirectional semantic analysis of the preceding and following text, the system can accurately determine what information needs to be supplemented in the fragments and how much, ensuring that each fragment remains highly comprehensible even after being separated from the original document. This allows each document fragment to be both independent (relying on key information from the preceding text) and smoothly connected (relying on related information from the following text), thus achieving an optimal balance between fragment independence and overall document coherence. Furthermore, the defined three-level retention strategy (complete, key, and necessary) dynamically selects the most economical and effective retention strategy based on real-time analysis results. While meeting semantic coherence requirements, it minimizes the introduction of redundant information, significantly improving the information density of the fragments and the efficiency of processing, storage, and transmission. Moreover, since document fragments contain carefully selected context, this greatly facilitates downstream natural language processing tasks such as text understanding, question answering, and summarization. For example, when performing question answering based on fragments, Large Language Models (LLMs) can obtain key background information without additional tracing of the original text, significantly improving the accuracy and relevance of the answers and reducing the complexity and error rate of task processing.

[0083] In some embodiments of the present invention, such as Figure 6 As shown, step S106 includes: S601. Perform semantic integrity assessment on the document fragments to obtain a first score.

[0084] It should be noted that the semantic integrity assessment of document fragments checks the completeness of semantic units, semantic logical consistency, and semantic expression within each fragment, and outputs a first score. Specifically, the semantic unit integrity check examines whether there are incomplete sentences or broken words within the fragment, and whether it begins and ends with a complete sentence or paragraph. The semantic logical consistency check uses a topic model to analyze the topic concentration of the text within the fragment, checking for abrupt insertions of irrelevant topics, ensuring that the content revolves around one or a few closely related core topics. The semantic expression integrity check identifies whether the fragment contains a complete semantic expression structure, such as "problem-analysis-solution" or "argument-evidence-conclusion," and determines whether this structure is completely preserved within the fragment. Based on the combined results of the semantic unit integrity, semantic logical consistency, and semantic expression integrity checks, a first score (S) between 0 and 100 is generated through rule-based scoring or lightweight model prediction.

[0085] S602, Perform contextual evaluation on the document fragments to obtain a second score.

[0086] It's important to note that evaluating the contextual relevance of document fragments essentially involves calculating the relationship between the current fragment and its preceding and following fragments in the original document. Contextual relevance evaluation includes assessing the preceding contextual relevance (i.e., the semantic similarity between the document fragment and its preceding adjacent fragment), the following contextual relevance (i.e., the semantic similarity and logical coherence between the document fragment and its following adjacent fragment), and the overall coherence (i.e., the logical coherence between the current document fragment and its adjacent fragments), and outputting a second score. Assessing the preceding contextual relevance involves calculating the text vector similarity between the beginning of the current document fragment and the end of its preceding fragment (as described above). Assessing the following contextual relevance involves calculating the similarity between the end of the current document fragment and the beginning of its following fragment, taking the average or lowest value as the coherence index. Assessing the overall coherence involves checking whether logical connectors (such as "therefore," "however") and referential relationships (such as "this method," "the aforementioned system") between fragments are properly handled, and whether there are any unclear referents or logical breaks. Based on the combined evaluation results of the preceding and following contexts and overall coherence, a second score (C) between 0 and 100 is generated through rule-based scoring or lightweight model prediction.

[0087] S603. Perform a structural integrity assessment on the document fragments to obtain a third score.

[0088] It should be noted that the document fragments undergo structural integrity assessment. This assessment includes checking the grammatical structure integrity, punctuation integrity, and formatting consistency of the document fragments, and outputting a third score. Specifically, the grammatical structure integrity check includes verifying whether sentence grammar within a document fragment is broken due to fragmentation; punctuation integrity checks include verifying whether parentheses and quotation marks within a document fragment are paired; and formatting element integrity checks include verifying the appropriateness of heading hierarchy (e.g., isolated subheadings should not appear), the completeness of list items, whether tables or code blocks are cut in the middle, and whether image titles are separated from images. Based on the combined results of these checks on grammatical structure integrity, punctuation integrity, and formatting consistency, a third score (T) between 0 and 100 is generated using rule-based scoring or lightweight model prediction, depending on the degree and type of structural disruption.

[0089] S604. Based on the first score, the second score, and the third score, a weighted comprehensive quality score for the document segment is calculated.

[0090] It should be noted that the calculation formula is: Overall Quality Score Q = Wq1 × S + Wq2 × C + Wq3 × T. The weight values ​​Wq1, Wq2, and Wq3 (where Wq1 + Wq2 + Wq3 = 1) are not fixed but dynamically configured based on the document type. For example, technical / academic documents prioritize semantics and structure, and might have Wq1 = 0.5, Wq2 = 0.3, and Wq3 = 0.2. News / narrative documents prioritize coherence, and might have Wq1 = 0.4, Wq2 = 0.4, and Wq3 = 0.2. Code documents place extreme emphasis on structure, and might have Wq1 = 0.2, Wq2 = 0.2, and Wq3 = 0.6.

[0091] S605. Compare the comprehensive quality score with the first quality threshold and the second quality threshold; the first quality threshold is greater than the second quality threshold.

[0092] S606. If the comprehensive quality score is greater than or equal to the first quality threshold, determine that the document segmentation is qualified and output the segmentation result, and select the next target segmentation point for segmentation operation.

[0093] S607. If the overall quality score is less than the second quality threshold, the document segmentation is determined to be unqualified, the segmentation parameters are adjusted, and the segmentation process is re-performed based on the adjusted parameters.

[0094] It should be noted that: the comprehensive quality score is compared with the first quality threshold and the second quality threshold respectively in size, and whether the document fragment obtained by segmentation according to the target segmentation point is qualified is determined according to the comparison result. Assume that the first quality threshold Th_high=80, the second quality threshold Th_low=60, and Th_high>Th_low. If the comprehensive quality score Q≥Th_high, it is determined as a high-quality fragment, the system confirms that the document fragment is qualified and adds it to the final output result. Subsequently, the process automatically filters the next target segmentation point from the remaining documents and continues the segmentation processing (S606). If Th_low≤Q<Th_high, it is determined as a medium-quality fragment, the system confirms that the document fragment can be output, but the system will record its score and possible problem dimensions for generating an optimization report, which can be used as a reference for subsequent manual or automatic batch optimization. If Q<Th_low, it is determined as a low-quality fragment and unqualified. At this time, the system will not output the problematic fragment, but immediately trigger the parameter adjustment mechanism (S607). The system analyzes the main reasons for the low score (e.g., low semantic score may be due to improper boundary weight, low structural score may be due to incorrect detection range), and adjusts relevant parameters accordingly (such as the recognition confidence threshold used in semantic boundary detection, the preset weights assigned to different semantic boundary types, the text range covered by the preceding semantic analysis or the following semantic analysis in context retention processing, the selection threshold of the context retention strategy, the constraint range of the fragment size, etc.). Subsequently, based on the adjusted new parameters, the system goes back to the early steps of the segmentation process (such as S101 or S103) to re-segment the original document to be processed or the problematic area.

[0095] In this embodiment, by establishing a quantitative evaluation model integrating semantics, correlation and structure, the subjective concept of fragment quality is converted into an objective and comparable comprehensive quality score (Q). Combined with a clear quality threshold, the system can automatically and strictly make a decision of acceptance, warning or rejection for each fragment, which ensures the overall quality bottom line of the final output result and realizes standardized measurement and controllable output of fragment quality. Furthermore, when a fragment is determined to be unqualified (Q<Th_low), the system does not simply discard or mark it, but triggers a closed loop of analysis-adjustment-reprocessing, which enables the system to learn from errors and improve itself. Through iterative optimization, the system parameters can adapt to the characteristics of specific documents, so that higher-quality fragments can be produced gradually in continuous processing. Furthermore, since all output fragments have passed strict multi-dimensional quality inspection, the input received by downstream AI models (such as LLM), search engines or knowledge base construction systems is of high quality and high integrity. This greatly reduces model misunderstanding and retrieval failure caused by input fragmentation, information missing or structural confusion, thereby directly improving the processing effect, accuracy and reliability of downstream tasks.

[0096] In order to improve the flexibility and adaptability of document sharding. In some embodiments of the present invention, as Figure 7 shown, before step S101, the method comprises: S701: identifying a document type of the document to be processed.

[0097] It should be noted that: the system extracts multi-dimensional features from the document, including structural features, lexical and syntactic features, metadata and format features. The extracted multi-dimensional features are input into a pre-trained document classification model (e.g., a machine learning-based text classifier). The document classification model outputs the probability that the document belongs to each document type. Common document types include technical / academic documents (e.g., papers, technical reports), news / narrative documents (e.g., news reports, blog posts), legal / regulatory documents (e.g., contracts, clauses), code / script documents, and business / offical documents (e.g., reports, letters). Strategy association: according to the identified document type (or main document type).

[0098] S702: preprocessing the to-be-processed document according to the document type.

[0099] It should be noted that: preprocessing may include unified encoding and garbled character repair processing, format stripping and structured extraction processing, content normalization processing, special content processing, noise removal processing and the like. Wherein, the unified encoding and garbled character repair processing includes detecting an original character encoding of the to-be-processed document, and uniformly converting the original character encoding into an internal standard encoding of the system (e.g., UTF-8). Meanwhile, an attempt is made to repair or remove unrecognizable garbled characters. The format stripping and structured extraction processing includes, for rich text formats (e.g., PDF, DOCX), extracting a pure text stream by using a parsing library, and retaining explicit structural markers as much as possible, such as heading levels, list item numbers, and cell relationships of tables. The content normalization processing includes removing or standardizing irrelevant typesetting information, such as headers, footers, page numbers and watermarks, and performing preliminary segmentation based on simple line breaks and periods to provide basic units for subsequent more fine-grained semantic boundary detection. The special content processing includes identifying and protecting code blocks, so as to avoid misjudging indentation and line breaks in codes as natural language paragraphs. A natural language narrative part is distinguished from formulas, tables and lists therein, and marked separately. The noise removal processing includes stop word filtering and noise removal; for some tasks focusing on content analysis, common stop words that contribute little to semantics (such as "de", "le", "the", "is") can be selectively filtered out. Of course, meaningless advertisement texts and template words can also be removed.

[0100] In this embodiment, by accurately identifying document types and preprocessing the documents to be processed according to their types, the boundary misjudgment caused by input noise in S101 is greatly reduced, ensuring the quality, stability, and accuracy of semantic boundary recognition. Furthermore, real-world document formats vary widely; through type recognition and targeted preprocessing, the system possesses strong format compatibility and content understanding capabilities. Whether it's an article crawled from a webpage, a scanned PDF report, or a tutorial containing code snippets, the system can convert it into a unified representation that it can process internally. This reliably extends intelligent segmentation capabilities to a wide range of practical application scenarios, enhancing the practical value of the solution and significantly improving the system's generalization ability and robustness in handling complex documents. Additionally, at least one of the preset weights in step S103, the context-preserving processing strategy in step S105, and the quality evaluation criteria in step S106 is adaptively configured based on the document type.

[0101] To better implement the semantic boundary-based document fragmentation method in this invention embodiment, based on the semantic boundary-based document fragmentation method, correspondingly, as follows: Figure 8 As shown, this embodiment of the invention also provides a document fragmentation system 800 based on semantic boundaries, the document fragmentation system 800 based on semantic boundaries includes: The input processing module 801 is used to identify multi-level semantic boundaries in the document to be processed; The intelligent segmentation module 802 is used to generate multiple candidate segmentation points based on the semantic boundary, calculate the semantic weight of each candidate segmentation point based on the preset weight corresponding to the boundary type, select a target segmentation point from the multiple candidate segmentation points based on the semantic weight, and perform a segmentation operation at the target segmentation point, and perform context-preserving processing on the text before and after the target segmentation point to generate independent document segments. The quality assessment module 803 is used to assess the quality of the document segments. If the segment is determined to be unqualified based on the quality assessment result, the segmentation parameters are adjusted and the segmentation process is re-performed based on the adjusted parameters.

[0102] The semantic boundary-based document segmentation system 800 provided in the above embodiments can implement the technical solutions described in the above semantic boundary-based document segmentation method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above semantic boundary-based document segmentation method embodiments, which will not be repeated here.

[0103] like Figure 8 As shown, the core workflow of intelligent file segmentation is as follows: large document input → document type identification → content preprocessing → semantic boundary identification → boundary weight calculation → candidate segmentation point generation → optimal segmentation point selection → context preservation processing → file generation → file quality assessment → quality assessment result. If the quality assessment result indicates that the file quality is acceptable, the file segmentation result is output, and the file segmentation is complete. If the quality assessment result indicates that the file quality is unacceptable, the parameters are adjusted and the boundary weights are recalculated.

[0104] The semantic boundary recognition algorithm process is as follows: document content → multi-level boundary detection. Multi-level boundary detection includes paragraph boundary detection, sentence boundary detection, lexical boundary detection, and character boundary detection. Paragraph boundary detection includes paragraph identifier recognition, paragraph structure analysis, and paragraph topic recognition. Sentence boundary detection includes period detection, question mark detection, exclamation mark detection, and sentence structure analysis. Lexical boundary detection includes space detection, punctuation detection, and lexical integrity checking. Character boundary detection includes character type recognition and character encoding detection. After multi-level boundary detection, boundary weights are calculated based on the paragraph boundary detection results, sentence boundary detection results, lexical boundary detection results, and character boundary detection results → boundary priority ranking is performed.

[0105] The target segmentation point selection algorithm flow is as follows: candidate segmentation point set → semantic integrity evaluation → context association evaluation → segmentation point weight calculation → segmentation point sorting → segmentation point quantity determination. If there are multiple segmentation points, the optimal candidate segmentation point is selected, and its quality is verified → quality verification result judgment. If there is only a single segmentation point, the quality of this single segmentation point is directly verified → quality verification result judgment. If the quality verification result indicates that the segmentation point fails quality verification, the second-best candidate segmentation point is selected, and quality verification is performed again. If the re-verification result fails, the segmentation parameters are adjusted, a new candidate segmentation point set is generated, and semantic integrity evaluation is performed again. If the quality verification result indicates that the segmentation point passes quality verification, or if the re-verification result is successful, the target segmentation point is confirmed, and the segmentation operation is performed.

[0106] The context preservation processing mechanism is as follows: Segmentation point determination → Context scope analysis → Key information extraction after semantic analysis of the preceding text, and identification of related information after semantic analysis of the following text → Context preservation strategy selection → Preservation strategy type determination: if the preservation strategy is full preservation, the complete context is retained; if the preservation strategy is partial preservation, key context is retained; if the preservation strategy is minimal preservation, necessary context is retained → Context information embedding → Segment content generation → Semantic integrity verification → Verification result judgment. If the semantic integrity verification result is unsuccessful, the preservation strategy is adjusted, the preservation strategy type is re-determined, and context information embedding is performed again. If the semantic integrity verification result is successful, the segment is output, and segmentation is complete.

[0107] The segmentation quality assessment mechanism proceeds as follows: Semantic integrity assessment, contextual relevance assessment, and structural integrity assessment are performed on the segmented content. Semantic integrity assessment includes semantic unit integrity checks, semantic logical consistency checks, and semantic expression integrity checks. Contextual relevance assessment includes assessment of preceding and following contextual relevance, and overall coherence. Structural integrity assessment includes syntactic structure integrity assessment, punctuation integrity assessment, and format consistency checks. A quality score is calculated based on a comprehensive quality score derived from the semantic unit integrity check, semantic logical consistency check, and semantic expression integrity check. If the quality score is 80 or higher, the segmented content is considered high-quality and recommended for use. If the quality score is 60 or higher and 79 or lower, the segmented content is considered medium-quality and optimization is recommended. If the quality score is lower than 60, the segmented content is considered low-quality and requires re-segmentation.

[0108] like Figure 9 As shown, the present invention also provides an electronic device 900. The electronic device 900 includes a processor 901, a memory 902, and a display 903. Figure 9 Only some components of the electronic device 900 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0109] In some embodiments, processor 901 may be a central processing unit (CPU), microprocessor, or other data processing chip, used to run program code stored in memory 902 or process data, such as the semantic boundary-based document fragmentation method of the present invention.

[0110] In some embodiments, processor 901 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 901 may be local or remote. In some embodiments, processor 901 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, intranet, multi-cloud, etc., or any combination thereof.

[0111] In some embodiments, memory 902 may be an internal storage unit of electronic device 900, such as a hard disk or memory of electronic device 900. In other embodiments, memory 902 may also be an external storage device of electronic device 900, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 900.

[0112] Furthermore, the memory 902 may include both internal storage units of the electronic device 900 and external storage devices. The memory 902 is used to store application software and various types of data installed on the electronic device 900.

[0113] In some embodiments, display 903 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 903 is used to display information from electronic device 900 and to display a visual user interface. Components 901-903 of electronic device 900 communicate with each other via a system bus.

[0114] In one embodiment, when processor 901 executes a semantic boundary-based document fragmentation procedure in memory 902, the following steps may be performed: Identify multi-level semantic boundaries in the document to be processed; Multiple candidate segmentation points are generated based on the semantic boundaries; Based on the preset weights corresponding to the boundary types, the semantic weights of each candidate segmentation point are calculated. Based on the semantic weights, a target segmentation point is selected from the plurality of candidate segmentation points; The splitting operation is performed at the target splitting point, and the text before and after the target splitting point is processed to preserve the context, so as to generate independent document fragments; The document fragments are evaluated for quality. If the fragments are deemed unqualified based on the evaluation results, the fragmentation parameters are adjusted and the fragmentation process is repeated based on the adjusted parameters.

[0115] It should be understood that when the processor 901 executes the semantic boundary-based document fragmentation program in the memory 902, in addition to the functions mentioned above, it can also perform other functions, as detailed in the description of the corresponding method embodiments above.

[0116] Furthermore, this embodiment of the invention does not specifically limit the type of electronic device 900 mentioned. Electronic device 900 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the invention, electronic device 900 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).

[0117] Accordingly, embodiments of this application also provide a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the semantic boundary-based document fragmentation method provided in the above-described method embodiments.

[0118] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0119] The foregoing has provided a detailed description of the document fragmentation method, system, device, and medium based on semantic boundaries provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.< / h1>

Claims

1. A document fragmentation method based on semantic boundaries, characterized in that, include: Identify multi-level semantic boundaries in the document to be processed; the semantic boundary is the natural dividing line between semantic units, the semantic unit is the smallest language segment in the document text that carries a relatively independent and complete meaning, and the multi-level semantic boundaries are defined by dividing and identifying semantic boundaries of different granularities according to the semantic integrity of the document text to be processed. Multiple candidate segmentation points are generated based on the semantic boundaries; Based on the preset weights corresponding to the boundary types, the semantic weights of each candidate segmentation point are calculated. Based on the semantic weights, a target segmentation point is selected from the plurality of candidate segmentation points; A segmentation operation is performed at the target segmentation point, and context-preserving processing is applied to the text before and after the target segmentation point to generate independent document fragments. The context-preserving processing includes performing semantic analysis of the preceding and following text within a certain range before and after the segmentation point, respectively, and seamlessly embedding the selected context information into the generated fragment text content based on the analysis results. The document fragments are evaluated for quality. If the fragments are deemed unqualified based on the evaluation results, the fragmentation parameters are adjusted and the fragmentation process is repeated based on the adjusted parameters.

2. The method according to claim 1, characterized in that, The identification of multi-level semantic boundaries in the document to be processed includes: Paragraph boundaries are determined by detecting paragraph identifiers, analyzing paragraph structure, and identifying paragraph topics. Sentence boundaries are determined through sentence structure analysis and punctuation detection at the end of sentences. Word boundaries are determined through word separator detection and word integrity checks. Character boundaries are determined through character type recognition and character encoding detection; The semantic boundaries include the paragraph boundaries, sentence boundaries, word boundaries, and character boundaries.

3. The method according to claim 1, characterized in that, The semantic weights of each candidate segmentation point are calculated based on preset weights corresponding to boundary types, including: Determine the corresponding preset weights based on the semantic boundary types of the candidate segmentation points; Obtain the recognition confidence of the semantic boundary type; Based on the preset weights and the recognition confidence level, the semantic weights are calculated according to the adjustment factor; the adjustment factor is determined according to the document type of the document to be processed.

4. The method according to claim 1, characterized in that, The step of selecting a target segmentation point from the plurality of candidate segmentation points based on the semantic weights includes: The optimal candidate segmentation point is subjected to quality verification; the optimal candidate segmentation point is the candidate segmentation point with the largest semantic weight, and the quality verification includes semantic integrity verification and contextual relevance verification after simulated segmentation; If the quality verification passes, the optimal candidate segmentation point is determined as the target segmentation point; If the quality verification fails, a suboptimal candidate segmentation point is selected for verification, or the segmentation parameters are adjusted and a new candidate segmentation point is regenerated until the target segmentation point is selected; the suboptimal candidate segmentation point is the candidate segmentation point with the second largest semantic weight.

5. The method according to claim 1, characterized in that, The context-preserving processing of the text before and after the target segmentation point includes: Perform contextual semantic analysis on the text located within a first preset range before the target segmentation point to extract key information; Perform contextual semantic analysis on the text located within a second preset range after the target segmentation point to identify related information; Based on the analysis results of the key information and related information, one of the following is selected as the target context preservation strategy: complete retention strategy, key retention strategy, and necessary retention strategy; the context preservation range of the complete retention strategy is greater than that of the key retention strategy, and the context preservation range of the key retention strategy is greater than that of the necessary retention strategy. According to the target context preservation strategy, the corresponding range of context information is embedded into the segmented text generated by segmenting before and after the target segmentation point to obtain the document segment.

6. The method according to claim 1, characterized in that, The quality assessment of the document fragments includes: The document fragments are subjected to semantic integrity assessment to obtain a first score; A second score is obtained by evaluating the contextual association of the document fragments; A third score is obtained by performing a structural integrity assessment on the document fragments; The overall quality score of the document segment is calculated by weighting the first score, the second score, and the third score. The overall quality score is compared with a first quality threshold and a second quality threshold; the first quality threshold is greater than the second quality threshold. If the overall quality score is greater than or equal to the first quality threshold, the document segmentation is determined to be qualified and the segmentation result is output. Then, the next target segmentation point is selected for segmentation. If the overall quality score is less than the second quality threshold, the document segmentation is determined to be unqualified, the segmentation parameters are adjusted, and the segmentation process is re-performed based on the adjusted parameters.

7. The method according to any one of claims 1 to 6, characterized in that, Before identifying the multi-level semantic boundaries in the document to be processed, the following steps are included: Identify the document type of the document to be processed; The document to be processed is preprocessed according to the document type.

8. A document fragmentation system based on semantic boundaries, characterized in that, include: The input processing module is used to identify multi-level semantic boundaries in the document to be processed. The semantic boundary is a natural dividing line between semantic units. The semantic unit is the smallest language segment in the document text that carries a relatively independent and complete meaning. The multi-level semantic boundaries are defined by dividing and identifying semantic boundaries of different granularities according to the semantic integrity of the document text to be processed. The intelligent segmentation module is used to generate multiple candidate segmentation points based on the semantic boundary, calculate the semantic weight of each candidate segmentation point based on the preset weight corresponding to the boundary type, and select the target segmentation point from the multiple candidate segmentation points based on the semantic weight. And perform a segmentation operation at the target segmentation point, and perform context preservation processing on the text before and after the target segmentation point to generate independent document fragments; the context preservation processing includes performing preceding semantic analysis and following semantic analysis on a certain range of text before and after the segmentation point, and seamlessly embedding the selected context information into the generated fragment text content based on the analysis results; The quality assessment module is used to assess the quality of the document segments. If the segment is determined to be unqualified based on the quality assessment results, the segmentation parameters are adjusted and the segmentation process is repeated based on the adjusted parameters.

9. An electronic device, characterized in that, Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the semantic boundary-based document fragmentation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the semantic boundary-based document fragmentation method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fragmentation method, system and device based on information entropy and Transform

    CN120449880A

  • Text processing method and device, computer equipment and storage medium

    CN120975097A