Data set construction method and system for super-long text processing
Through the data set construction methods of format conversion, cleaning, symbol substitution and dynamic chunking, the problems of information loss and semantic barriers in ultra-long text processing are solved, and efficient data segmentation and semantic retention are achieved, which is suitable for text processing of large language models.
Patent Information
- Application Number
- CN202510654062.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-02
AI Technical Summary
When processing ultra-long text, the prior art is limited by the number of input tokens of large language models, resulting in information loss, semantic understanding obstacles and inefficiency.
A data set construction method is adopted, including format conversion, cleaning, symbol substitution, dynamic threshold calculation model chunking, etc. Combined with the large model input token limit, text type complexity and average paragraph length, the ultra-long text is divided into data blocks that meet the input limit, and the feature vectors when document fusion is retained.
It improves processing efficiency, reduces the number of tokens, supports subsequent restore operations, adapts to different scenarios, retains semantic coherence during document fusion, and is suitable for large model training and text processing tools.
Smart Images

Figure CN120579128A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data segmentation, and in particular to a data set construction method and system for processing extremely long texts. Background Art
[0002] With the rapid development of natural language processing technology, large language models represented by Transformer (such as GPT and BERT) have demonstrated outstanding capabilities in tasks such as long text understanding, generation, and classification. However, these large models are often constrained by the number of input tokens (such as 512, 1024, or 4096 tokens), making it difficult to effectively process very long texts (such as PDF documents, Word documents, or other long articles).
[0003] Traditional technologies usually use the following methods to process very long texts:
[0004] 1. Simple truncation: Directly truncate the text to within the token limit, resulting in content loss and possible semantic disruption.
[0005] 2. Fixed segmentation: The text is segmented by fixed length without considering semantic integrity. Paragraphs or sentences may be split in the middle, causing semantic comprehension difficulties.
[0006] 3. Manual segmentation: Relying on manual segmentation of long texts is inefficient and difficult to meet the needs of large-scale processing. Summary of the Invention
[0007] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a dataset construction method and system for processing very long texts. The present invention solves the technical problems in the existing technology that large language models have limitations on the number of input tokens, resulting in information loss, semantic understanding barriers and low efficiency when processing very long texts.
[0008] To achieve the above object, the present invention provides the following solutions:
[0009] A dataset construction method for processing very long texts, comprising:
[0010] Obtain the file to be input and convert the file into txt format to obtain a format file;
[0011] Performing a first cleaning on the format file to obtain a first cleaned file;
[0012] Counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file;
[0013] performing a second cleaning on the symbol file to obtain a second cleaned file;
[0014] Based on the dynamic threshold calculation model, combined with the large model input token limit, text type complexity factor and average paragraph length, the second cleaned file is divided into blocks to obtain a block data set;
[0015] An output data set is obtained according to the block data set.
[0016] Preferably, the first cleaning of the format file to obtain a first cleaned file includes:
[0017] Determining meaningless symbols and placeholders in the format file;
[0018] The meaningless symbols and placeholders are removed to obtain a first cleaned file.
[0019] Preferably, the step of counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file includes:
[0020] Use the collections.Counter function to quickly count word frequencies, and use most_common(top_n) to obtain the top_n high-frequency words. Sort the high-frequency words by frequency to obtain a high-frequency word set.
[0021] Comparing each word in the high-frequency word set with a preset word frequency threshold and a word length threshold to obtain a comparison result;
[0022] According to the comparison result, the corresponding vocabulary is mapped to a preset symbol dictionary to determine the marking symbol corresponding to the current vocabulary;
[0023] The corresponding words in the high-frequency vocabulary set are replaced according to the marking symbols to obtain a symbol file.
[0024] Preferably, the dynamic threshold calculation model includes:
[0025]
[0026] Among them, L0 is the dynamic threshold, γ is the safety margin ratio, α is the paragraph length weight, β is the element complexity suppression factor, E base is the benchmark element complexity, L LLM Input token limit for large models, E size is the text type complexity factor, L avg is the average paragraph length, L base The baseline paragraph length.
[0027] Preferably, it also includes:
[0028] retaining the overlapping content between adjacent text segments when segmenting the second cleaned file;
[0029] Generate feature vectors for the overlapping parts of the block content and use BERT word embedding technology to extract the corresponding feature vector group.
[0030] A dataset construction system for processing very long texts, comprising:
[0031] The format conversion module is used to obtain the file to be input and convert the file into txt format to obtain a format file;
[0032] A first cleaning module, configured to perform a first cleaning on the format file to obtain a first cleaned file;
[0033] A replacement module, for counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file;
[0034] A second cleaning module, configured to perform a second cleaning on the symbol file to obtain a second cleaned file;
[0035] A segmentation module is used to segment the second cleaned file into blocks based on a dynamic threshold calculation model, combined with the large model input token limit, text type complexity factor and average paragraph length, to obtain a block data set;
[0036] An output module is used to obtain an output data set according to the block data set.
[0037] The present invention discloses the following technical effects:
[0038] The present invention provides a data set construction method and system for processing ultra-long texts, the method comprising: obtaining a file to be input and converting the file into a txt format to obtain a format file; performing a first cleaning on the format file to obtain a first cleaning file; counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file; performing a second cleaning on the symbol file to obtain a second cleaning file; based on a dynamic threshold calculation model, combined with a large model input token restriction, a text type complexity factor, and an average paragraph length, the second cleaning file is divided into blocks to obtain a block data set; and an output data set is obtained based on the block data set. The high-frequency word symbol mapping mechanism of the present invention improves processing efficiency, reduces the number of tokens, and supports subsequent restoration operations; provides a threshold condition (word frequency threshold K x and vocabulary length threshold K z ), which can flexibly adapt to different scenarios; it retains the feature vectors required for document fusion and can be encapsulated as related text processing tools, suitable for scenarios such as Agent construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 A flowchart of a method for constructing a dataset for processing very long text provided by an embodiment of the present invention;
[0041] Figure 2 A detailed diagram of a data set construction method for processing very long texts provided by an embodiment of the present invention;
[0042] Figure 3 A workflow diagram of a paper review tool provided by an embodiment of the present invention;
[0043] Figure 4 A schematic diagram of the working of a document extractor provided in an embodiment of the present invention;
[0044] Figure 5 A schematic diagram of block operation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] like Figure 1-2 As shown, the present invention provides a data set construction method for ultra-long text processing, comprising:
[0048] Step 100: Obtain a file to be input and convert the file into a txt format to obtain a formatted file;
[0049] Specifically, the text content in super-long files (such as PDF and Word documents) is extracted into txt format through text function modules such as PyPDF2 and python-docx in the Python library.
[0050] The extracted text is stored in character array format for easy subsequent processing.
[0051] Step 200: performing a first cleaning on the format file to obtain a first cleaned file;
[0052] Step 300: Count the high-frequency words in the cleaning file and replace the high-frequency words with symbols to obtain a symbol file;
[0053] Step 400: performing a second cleaning on the symbol file to obtain a second cleaned file;
[0054] Step 500: Based on the dynamic threshold calculation model, combined with the large model input token restriction, text type complexity factor and average paragraph length, the second cleaned file is divided into blocks to obtain a block data set;
[0055] Step 600: Obtain an output data set according to the block data set.
[0056] Furthermore, the first cleaning of the format file to obtain a first cleaned file includes:
[0057] Determining meaningless symbols and placeholders in the format file;
[0058] The meaningless symbols and placeholders are removed to obtain a first cleaned file.
[0059] Furthermore, the high-frequency words in the cleaning file are counted and the high-frequency words are replaced with symbols to obtain a symbol file, including:
[0060] Use the collections.Counter function to quickly count word frequencies, and use most_common(top_n) to obtain the top_n high-frequency words. Sort the high-frequency words by frequency to obtain a high-frequency word set.
[0061] Comparing each word in the high-frequency word set with a preset word frequency threshold and a word length threshold to obtain a comparison result;
[0062] According to the comparison result, the corresponding vocabulary is mapped to a preset symbol dictionary to determine the marking symbol corresponding to the current vocabulary;
[0063] The corresponding words in the high-frequency vocabulary set are replaced according to the marking symbols to obtain a symbol file.
[0064] Specifically, clear meaningless symbols and placeholders (such as spaces, *, etc.).
[0065] Count the high-frequency words in the txt text and replace them with short token symbols, specifically:
[0066] (1) Use the collections.Counter function to quickly count word frequencies, and use most_common(top_n) to obtain the top_n high-frequency words. Sort the high-frequency words according to their frequencies.
[0067] (2) Set the high-frequency word frequency threshold K x and the word length threshold K z . When the frequency of a word > K x and the length of the word > K z , map the relevant words to the symbol dictionary M (for example, the symbol for high-frequency word 1 is M1, the symbol for high-frequency word 2 is M2,...) for restoration when outputting subsequent text.
[0068] (3) Further clean the text based on the self-built stop word list and the Harbin Institute of Technology stop word list, removing common stop words (such as "le", "guo", "ba", "a", "ma", etc.).
[0069] (4) Process noise words (referring to content that cannot provide effective information for text semantics, including garbled text, meaningless characters, etc.). Specifically:
[0070] 1) For unresolvable garbled characters, directly remove them or replace them with placeholders;
[0071] 2) For meaningless characters, use regular expressions to remove special symbols.
[0072] Furthermore, perform intelligent chunking on the ultra-long text according to the token limit. The segmented text segments (txt_1, txt_2,..., txt_n) meet the input limit of the large model; set the token threshold (such as 4096 tokens), calculate the number of tokens according to the character or word length, and chunk the text in units of paragraphs or sentences. At the same time, ensure that there is a certain overlap between adjacent chunks for subsequent splicing.
[0073] 1. Use hard decoding chunking. Specifically, set the token threshold to L0 = 63k (the maximum supported tokens input for deepseek is 64k), and the element size threshold is 500;
[0074] 2. Confirm the token length L1 of the cleaned text txt through the defined python code. If L1 < L0, there is no need to segment, otherwise execute the segmentation program according to the standard length of L0.
[0075] Further, the dynamic threshold calculation model includes:
[0076]
[0077] Among them, L0 is the dynamic threshold, γ is the safety margin ratio, α is the paragraph length weight, β is the element complexity suppression factor, E base is the benchmark element complexity, L LLM Input token limit for large models, E size is the text type complexity factor, L avg is the average paragraph length, L base The baseline paragraph length.
[0078] The expression of text type complexity factor is:
[0079]
[0080] The expression for average paragraph length is:
[0081]
[0082] More specifically, the baseline paragraph length is a constant, generally taken as 500; the baseline element complexity is a constant, generally taken as 5; the paragraph length weight is 0.2-0.8, which is used to control the impact of paragraph length; the element complexity suppression factor is 0.05-02, which is used to suppress the threshold expansion of structured text; the safety margin ratio is 0.8, which simulates the dynamic token threshold exceeding the large model input token limit; the above dynamic threshold calculation model also has a constraint condition, namely L0∈[0.5L LLM , 0.95L LLM ].
[0083] Furthermore, it also includes:
[0084] retaining the overlapping content between adjacent text segments when segmenting the second cleaned file;
[0085] Generate feature vectors for the overlapping parts of the block content and use BERT word embedding technology to extract the corresponding feature vector group.
[0086] Specifically, (1) to ensure semantic coherence of the segmented text segments during aggregation, a certain proportion of content overlap is required between the texts before and after segmentation to avoid contextual discontinuity. Specifically, during segmentation, the end of each segment is repeatedly added to the beginning of the next segment.
[0087] (2) The overlapping parts are represented by vectors, and the feature vector group {Xn|X1, X2, ..., Xn} is extracted through BERT word embedding technology for subsequent calculation during text aggregation.
[0088] The input cleaning text txt is segmented into txt_1, txt_2, ..., txt_n.
[0089] Furthermore, the processed documents are output as standardized txt files, or other formats (such as JSON) according to the application scenario. This is used to build text datasets required for large model training or fine-tuning.
[0090] More specifically, the present invention also provides another embodiment for building a tool for paper review using the DeepSeek large model.
[0091] Tool purpose: To assist in the review of master's theses.
[0092] Difficulty to be solved: Remove the long text input restriction for large models.
[0093] like Figure 3-5 As shown, the workflow first processes the uploaded paper attachments into a character array in TXT format using the document extractor. Secondly, during the code execution phase, the TXT is segmented into 15 blocks using parameters such as the token threshold. If the actual block content is smaller than this many, subsequent blocks are padded with blank characters. The segmented text is then fed sequentially (via feature vector calculation) into Segmented Text Processor 1 and Segmented Text Processor 2, which process the text separately and input the processed results into the paper analysis module. Finally, the paper analysis module uses the DeepSeek large model to perform paper analysis.
[0094] The entire process divides the extremely long text into small blocks, sends them to different processors respectively, and finally aggregates them, successfully solving the limitation of DeepSeek input tokens.
[0095] This embodiment also provides a dataset construction system for processing very long texts, including:
[0096] The format conversion module is used to obtain the file to be input and convert the file into txt format to obtain a format file;
[0097] A first cleaning module, configured to perform a first cleaning on the format file to obtain a first cleaned file;
[0098] A replacement module, for counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file;
[0099] A second cleaning module, configured to perform a second cleaning on the symbol file to obtain a second cleaned file;
[0100] A segmentation module is used to segment the second cleaned file into blocks based on a dynamic threshold calculation model, combined with the large model input token limit, text type complexity factor and average paragraph length, to obtain a block data set;
[0101] An output module is used to obtain an output data set according to the block data set.
[0102] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0103] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for constructing a dataset for processing very long text, characterized in that: include: Obtain the file to be input and convert the file into txt format to obtain a format file; Performing a first cleaning on the format file to obtain a first cleaned file; Counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file; performing a second cleaning on the symbol file to obtain a second cleaned file; Based on the dynamic threshold calculation model, combined with the large model input token limit, text type complexity factor and average paragraph length, the second cleaned file is divided into blocks to obtain a block data set; An output data set is obtained according to the block data set.
2. The method for constructing a dataset for processing very long text according to claim 1, characterized in that: The first cleaning of the format file to obtain a first cleaned file includes: Determining meaningless symbols and placeholders in the format file; The meaningless symbols and placeholders are removed to obtain a first cleaned file.
3. The method for constructing a dataset for processing very long text according to claim 1, characterized in that: The method of counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file includes: Use the collections.Counter function to quickly count word frequencies, and use most_common(top_n) to obtain the top_n high-frequency words. Sort the high-frequency words by frequency to obtain a high-frequency word set. Comparing each word in the high-frequency word set with a preset word frequency threshold and a word length threshold to obtain a comparison result; According to the comparison result, the corresponding vocabulary is mapped to a preset symbol dictionary to determine the marking symbol corresponding to the current vocabulary; The corresponding words in the high-frequency vocabulary set are replaced according to the marking symbols to obtain a symbol file.
4. The method for constructing a dataset for processing very long text according to claim 1, wherein: The dynamic threshold calculation model includes: Among them, L0 is the dynamic threshold, γ is the safety margin ratio, α is the paragraph length weight, β is the element complexity suppression factor, E base is the benchmark element complexity, L LLM Input token limit for large models, E size is the text type complexity factor, L avg is the average paragraph length, L base The baseline paragraph length.
5. The method for constructing a dataset for processing very long text according to claim 1, wherein: Also includes: retaining the overlapping content between adjacent text segments when segmenting the second cleaned file; Generate feature vectors for the overlapping parts of the block content and use BERT word embedding technology to extract the corresponding feature vector group.
6. A dataset construction system for processing very long texts, characterized in that: include: The format conversion module is used to obtain the file to be input and convert the file into txt format to obtain a format file; A first cleaning module, configured to perform a first cleaning on the format file to obtain a first cleaned file; A replacement module, for counting high-frequency words in the cleaning file and performing symbol replacement on the high-frequency words to obtain a symbol file; A second cleaning module, configured to perform a second cleaning on the symbol file to obtain a second cleaned file; A segmentation module is used to segment the second cleaned file into blocks based on a dynamic threshold calculation model, combined with the large model input token limit, text type complexity factor and average paragraph length, to obtain a block data set; An output module is used to obtain an output data set according to the block data set.