A Method and Device for Segmenting Ultra-Long Texts of Large Models

Through a hierarchical genetic algorithm, the problem of inaccurate identification of paragraph boundaries in ultra-long text is solved through hierarchical genetic algorithm, and the problem of inaccurate identification of paragraph boundaries in ultra-long text is solved, more accurate text division and more efficient semantic coherence are achieved, and the generation effect of the big model is improved.

CN120087358BActive Publication Date: 2025-07-22WENZHOU UNIV OUJIANG COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510559046.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-22
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

When the existing text segmentation method processes ultra-long text, it is difficult to accurately identify paragraph boundaries while ensuring sentence integrity, resulting in poor semantic coherence and affecting the generation effect and context consistency of the big model.

Method used

Hierarchical genetic algorithm is used to combine sentence coherence and information integrity indicators, and the text paragraph boundaries are determined through iterative optimization of genetic algorithms, and the language model is used to calculate the fitness value and perform cross-mutation updates to construct a global-level population to determine the final slicing point.

Benefits of technology

It realizes more accurate and efficient division of text structure, improves semantic coherence and reduces processing complexity, and improves the generation accuracy and efficiency of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087358B_ABST
    Figure CN120087358B_ABST
Patent Text Reader

Abstract

The present invention provides a method for segmenting extremely long texts of large models, which includes obtaining an initial text data set, dividing it into fixed-length sub-segments and initializing a sub-population; calculating the sentence coherence and information integrity in the sub-population using a language model, and evaluating individuals as a fitness function; setting relevant parameters of the genetic algorithm, and updating the population using the genetic algorithm iteration formula to obtain a locally optimal sub-population scheme; subsequently merging all the locally optimal sub-populations to construct a global-level population; and optimizing this global population through the genetic algorithm to determine the final text paragraph boundaries. Implementing the present invention realizes a more accurate and efficient division of the text structure, improves semantic coherence and reduces processing complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data models, and in particular to a method and device for segmenting ultra-long texts of large models. Background Art

[0002] In the context of the rapid development of current large model-driven intelligent systems, text segmentation, as a key pre-step in large model retrieval and generation systems, has received increasing attention. Especially when dealing with ultra-long text inputs, how to reasonably and clearly structure the segmentation of the original text not only directly affects the effects of downstream embedding construction and semantic retrieval, but also relates to the accuracy and context consistency of the content generated by large models. Although the traditional fixed window sliding method is simple and easy to use, it ignores the semantic boundaries and paragraph structures between texts, often resulting in unreasonable selection of segmentation points, destroying semantic coherence, and affecting the performance of subsequent tasks.

[0003] In recent years, researchers have tried to combine statistical methods with deep language models to improve the intelligence level of text segmentation. However, in the face of different corpus scenarios, language styles, and context complexities, existing methods still pose great challenges in terms of processing flexibility, structure preservation ability, and global optimization effects. Especially in texts with strong structures and strict semantics such as literature and law, how to accurately identify paragraph boundaries while ensuring sentence integrity has become an important bottleneck restricting the improvement of large model retrieval accuracy. The mainstream text segmentation methods mainly include rule-based methods, statistic-based methods, and deep learning-based methods.

[0004] Rule-based methods rely on manually set dictionaries and segmentation rules. Although they are simple to implement and have high efficiency, they have weak processing capabilities for new words and ambiguous words and poor generality; statistic-based methods train models such as word frequency and mutual information through a large amount of corpus, and have a certain degree of self-adaptability, but their effects are limited when dealing with low-frequency words; while deep learning-based segmentation methods, such as Bi-LSTM, BERT, etc., can achieve better results in context understanding, but their training costs are high, they rely on a large amount of labeled data, and the model inference speed is slow.

[0005] In this context, swarm intelligence algorithms have been introduced into the text segmentation task and shown certain advantages. Swarm intelligence algorithms simulate the group cooperation behaviors in nature, such as ant colony optimization, particle swarm optimization, etc., and have the characteristics of strong global search ability, flexible parameter adjustment, and good adaptability to unstructured problems. Compared with traditional methods, swarm intelligence algorithms show stronger robustness and generalization ability in dealing with multi-solution and dynamically changing text segmentation tasks, providing a new direction for text segmentation research.

[0006] As an optimization method that simulates natural selection and genetic mechanisms, the Genetic Algorithm (GA) has remarkable global search capabilities and can effectively avoid falling into local optima in complex, multi-peak search spaces. Compared with traditional gradient optimization methods, it has no requirements for the continuity or differentiability of the objective function, so it has stronger adaptability and can be widely applied to problems with complex structures and unclear constraints. However, in large-scale text segmentation tasks, the genetic algorithm also has some deficiencies, such as a slow convergence rate, especially the problem of "premature convergence" is likely to occur when approaching the optimal solution. In this case, it is very difficult to maintain the balance between exploration and exploitation.

[0007] Therefore, it is necessary to provide a new text segmentation method to achieve a more accurate and efficient division of the text structure, improve semantic coherence and reduce processing complexity. Summary of the Invention

[0008] The technical problem to be solved by the embodiments of the present invention is to provide a method and device for segmenting ultra-long texts of large models, which can achieve a more accurate and efficient division of the text structure, improve semantic coherence and reduce processing complexity.

[0009] To solve the above technical problem, the embodiments of the present invention provide a method for segmenting ultra-long texts of large models, and the method includes the following steps:

[0010] S1. Obtain an initial text data set, divide it into sub-segments of a fixed length, and initialize the sub-population; wherein, the initialized sub-population includes the number of sub-populations, the total number of iterations, the texts divided in each sub-population, and the number and positions of text segmentations for each text.

[0011] S2. Determine the current iteration number and the corresponding sub-population, and calculate the fitness values of the individuals in the currently determined sub-population; wherein, the fitness value is composed of an adjacent sentence coherence index and an information integrity index:

[0012] S3. Set the relevant parameters of the genetic algorithm, and update the current sub-population through the selection crossover mutation formula to obtain a new sub-population.

[0013] S4. Calculate the fitness values of the obtained new sub-population, sort them from largest to smallest, and select the individuals that meet the predetermined conditions as the next-generation sub-population.

[0014] S5. If the current iteration number is equal to the total number of iterations, end the loop and output the optimal solutions of all sub-populations; otherwise, if the current iteration number is less than the total number of iterations, after adding 1 to the current iteration number, return to step S2.

[0015] S6. Use all sub-populations that output the optimal solution as new individuals to construct a global-level population;

[0016] S7. Through the genetic algorithm, optimize the global-level population until the loop ends to determine the final text paragraph boundary.

[0017] Among them, the sub-population in step S1 consists of N individuals Each individual contains the positions of multiple text segmentation points, numbered from the 1st to the D th segmentation point. Among them,

[0018] i = 1, 2, ..., N; j = 1, 2, ..., D ; N is the number of training sample individuals; D is the number of segments for each small text segmentation; represents the population obtained in the t th iteration; represents the position of the t th text segmentation point of the i th individual in the j th iteration; t represents the total number of iterations, and its value range is [1, 1000].

[0019] Among them, the fitness value of the individual in step S2 is calculated through formulas (1) to (3); among them,

[0020] (1);

[0021] (2);

[0022] (3);

[0023] Among them, represents the fitness value of the th individual in the population; and respectively represent the sentence coherence index and information integrity index of the th individual; and are the weights of the sentence coherence index and information integrity index respectively; represents the trained language model; and are the two sentences before and after the th segmentation point respectively; is the current paragraph; The function is used to calculate and the cosine similarity between them.

[0024] Among them, in the step S3, the current sub-population is updated through the crossover mutation formulas (5) to (7) to obtain a new sub-population; among them,

[0025] (5);

[0026] (6);

[0027] (7);

[0028] Among them, and respectively represent two randomly selected individuals at the th iteration; is the subscript of the crossover point; represents the value of the first dimension of the th individual at the th iteration; is a random number between [0, 1]; and represent the crossover probability and the mutation probability, both of which are related parameters of the genetic algorithm; is a Gaussian distribution random number with a mean of 0 and a standard deviation of ;

[0029] Among them, the individuals meeting the predetermined conditions in the step S4 are specifically the first N individuals with large fitness.

[0030] Among them, the specific steps of the step S7 include:

[0031] First, construct the global fitness function (8) and calculate the fitness value of the global-level population;

[0032] (8)

[0033] Among them, represents the fitness value of the th individual in the global population; and respectively represent the keyword sets of the th paragraph and the th paragraph;

[0034] Secondly, sort the fitness values of the individuals in the global population, select the top N individuals with the highest fitness values as the population for the next generation, and run the algorithm until the maximum number of iterations is reached and output the optimal solution.

[0035] An embodiment of the present invention also provides a large model ultra-long text segmentation device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the aforementioned large model ultra-long text segmentation method are implemented.

[0036] Implementing the embodiments of the present invention has the following beneficial effects:

[0037] In view of the characteristics of large-scale text datasets, the present invention designs a hierarchical genetic algorithm and provides different methods for calculating fitness functions, thereby effectively segmenting large-scale texts, making the generation results of large models more accurate, achieving a more accurate and efficient division of text structures, improving semantic coherence, and reducing processing complexity. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, obtaining other drawings without creative efforts still belongs to the scope of the present invention.

[0039] Figure 1 It is a flowchart of a large model ultra-long text segmentation method provided by an embodiment of the present invention. Detailed Embodiments

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings.

[0041] As Figure 1 shown, in an embodiment of the present invention, a large model ultra-long text segmentation method is proposed. The method includes the following steps:

[0042] Step S1: Obtain an initial text dataset, divide it into fixed-length sub-segments, and initialize the sub-population; wherein, the initialization of the sub-population includes the number of sub-populations, the total number of iterations, the texts divided in each sub-population, and the number and positions of text segments in each text.

[0043] S2: Determine the current iteration number and the corresponding sub-population, and calculate the fitness values of the individuals in the currently determined sub-population; wherein, the fitness value is composed of an adjacent sentence coherence index and an information integrity index:

[0044] S3. Set the relevant parameters of the genetic algorithm, and update the current sub-population by selecting the crossover and mutation formulas to obtain a new sub-population;

[0045] S4. Calculate the fitness values of the obtained new sub-population, sort them from largest to smallest, and select the individuals that meet the predetermined conditions as the next-generation sub-population;

[0046] S5. If the current iteration number is equal to the total number of iterations, end the loop and output the optimal solutions of all sub-populations; otherwise, if the current iteration number is less than the total number of iterations, after adding 1 to the current iteration number, return to step S2;

[0047] S6. Use all the sub-populations that output the optimal solutions as new individuals to construct a global-level population;

[0048] S7. Optimize the global-level population through the genetic algorithm until the loop ends to determine the final text paragraph boundaries.

[0049] Specifically, in step S1, according to the initial text training set, it is divided into multiple small text segments according to the fixed sequence length K, and after constructing a sub-population with each small text segment, the sub-population is initialized; among them, the sub-population is expressed by the formula =( ,…, ); i =1,2, ..., N; j =1,2, ..., D; k is the serial number of the text segment; N is the number of training sample individuals; D is the number of cuts for each small text; represents the population obtained at the t th iteration; represents at t the i th individual at the j th text cut point position; t represents the total number of iterations, and its value range is [1, 1000].

[0050] In step S2, determine the current iteration number and the corresponding sub-population, and calculate the fitness value of the individual through formulas (1) to (3); among them,

[0051] (1);

[0052] (2);

[0053] (3);

[0054] Among them, represents the fitness value of the -th individual in the population; and respectively represent the sentence coherence index and information integrity index of the -th individual; and are the weights of the sentence coherence index and information integrity index respectively; represents the pre-trained language model; and are the two sentences before and after the -th division point respectively; is the current paragraph; The function is used to calculate the and cosine similarity between them.

[0055] In step S3, set the crossover probability and mutation probability of the genetic algorithm, and update the current sub-population through the crossover and mutation formulas (5) - (7) to obtain a new sub-population; among them,

[0056] (5);

[0057] (6);

[0058] (7);

[0059] Among them, and respectively represent two randomly selected individuals at the -th iteration; is the subscript of the crossover point; represents the value of the first dimension of the -th individual at the -th iteration; is a random number between [0, 1]; and represent the crossover probability and mutation probability respectively, and both are related parameters of the genetic algorithm; is a Gaussian distribution random number with a mean of 0 and a standard deviation of ;

[0060] In step S4, for the population obtained in step S3, calculate the fitness value according to formulas (1), (2) and (3), sort them from largest to smallest, and select the top NIndividuals with high fitness are used as the next-generation population. That is, the individuals meeting the predetermined conditions are specifically the top N individuals with high fitness.

[0061] In step S5, if the algorithm runs to the maximum number of iterations (i.e., the total number of iterations), the loop ends and all the optimal solutions of the sub-populations are output. Otherwise, the current iteration number is incremented by 1, and the process returns to step S2.

[0062] In step S6, all the sub-populations obtained after the algorithm in step S5 runs to completion are used as new individuals to construct the global-level population GX.

[0063] In step S7, first, a global fitness function (8) is constructed to calculate the fitness values of the global-level population;

[0064] (8)

[0065] where represents the fitness value of the -th individual in the global population; and respectively represent the keyword sets of the -th paragraph and the -th paragraph;

[0066] Secondly, the fitness values of the global population individuals are sorted, and the top N individuals with the highest fitness are selected as the next-generation population. The algorithm runs to the maximum number of iterations and outputs the optimal solution.

[0067] Corresponding to a large model ultra-long text segmentation method provided by an embodiment of the present invention, an embodiment of the present invention also provides a large model ultra-long text segmentation device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the large model ultra-long text segmentation method provided by the embodiment of the present invention. For specific details, please refer to the relevant content above and will not be elaborated here.

[0068] Implementing the embodiments of the present invention has the following beneficial effects:

[0069] Aiming at the characteristics of large-scale text data sets, the present invention designs a hierarchical genetic algorithm and provides different fitness function calculation methods, thereby effectively segmenting large-scale texts, making the generation results of large models more accurate, achieving a more accurate and efficient division of text structures, improving semantic coherence, and reducing processing complexity.

[0070] Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as ROM / RAM, disk, optical disc, etc.

[0071] The above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A method for segmenting ultra-long texts of large models, characterized in that, The method includes the following steps: S1. Obtain an initial text data set, divide it into sub-segments of a fixed length, and initialize the sub-populations; wherein, the initialization of the sub-populations includes the number of sub-populations constructed from each sub-segment, the total number of iterations, the texts divided in each sub-population, and the number and positions of the cuts for each text; S2. Determine the current iteration number and the corresponding sub-population, and calculate the fitness values of the individuals in the currently determined sub-population; wherein, the fitness value is composed of an adjacent sentence coherence index and an information integrity index: S3. Set the relevant parameters of the genetic algorithm, and update the current sub-population through the selection, crossover, and mutation formula to obtain a new sub-population; S4. Calculate the fitness values of the obtained new sub-population, sort them from largest to smallest, and select the individuals that meet the predetermined conditions as the next-generation sub-population; S5. If the current iteration number is equal to the total number of iterations, end the loop and output the optimal solutions of all sub-populations; otherwise, if the current iteration number is less than the total number of iterations, after adding 1 to the current iteration number, return to step S2; S6. Use all the sub-populations that output the optimal solutions as new individuals to construct a global-level population; S7. Optimize the global-level population through the genetic algorithm until the loop ends to determine the final text paragraph boundaries; The sub-population in the step S1 consists of N individuals, and each individual contains the positions of multiple text segmentation points, which are numbered sequentially from the 1st to the D th segmentation point, D where is the number of segmentation points for each small text. The fitness value of the individual in step S2 is calculated by formulas (1) to (3); wherein, (1); (2); (3); Among them, represents the fitness value of the th individual in the population; and respectively represent the sentence coherence index and the information integrity index of the th individual; and are the weights of the sentence coherence index and the information integrity index respectively; represents the trained language model; and are the two sentences before and after the th division point respectively; is the current paragraph; The function is used to calculate the cosine similarity between and .

2. The method for splitting ultra-long texts of a large model according to claim 1, wherein The sub-population in the step S1 consists of N individuals , where i = 1, 2, ..., N;j = 1, 2, ..., D ; N is the number of training sample individuals; represents the population obtained in the t th iteration; represents the position of the t th text segmentation point of the i nd individual in the j th iteration; t represents the total number of iterations, and its value range is [1, 1000].

3. A large model ultra-long text segmentation device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the large model ultra-long text segmentation method described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Network text segmenting method based on genetic algorithm

    CN101710333A

  • Long text segmentation method and device, storage medium and electronic device

    CN113076720A