Text segmentation method and device, electronic equipment and computer readable storage medium
The method of sliding the sliding window and performing semantic segmentation and clustering solves the problem of semantic segmentation in text segmentation, and improves the accuracy and reliability of the Q&A output of the RAG system.
Patent Information
- Application Number
- CN202510210103.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has semantic splitting problems in the text segmentation process, which affects the accuracy and reliability of the Q&A output of the RAG system.
By sliding the preset sliding window, the target sentence is semantically divided, and the semantic segments obtained by the segmentation are clustered with the semantic segments obtained by the last segmentation to ensure that the semantic segments after the segmentation have semantic coherence.
It improves the accuracy and reliability of the Q&A output feedback from the RAG system, and ensures that the semantic segments after text segmentation have good semantic coherence.
Smart Images

Figure CN120046618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a text segmentation method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the breakthrough progress of large models in the field of natural language processing, question-and-answer systems based on large models have overcome the defects of traditional question-and-answer systems, such as being unable to effectively handle diverse user questions and mechanical responses. Because they can understand more complex language structures and capture richer context information, they can provide more accurate, fluent, and personalized answers to users, and interact with users more humanely and intelligently.
[0003] Although large models have improved the response quality of question-and-answer systems, due to the limited coverage of specific domain knowledge in their training process, there are defects that the generated answers may seem reasonable and logically coherent, but are actually inaccurate or contrary to facts. For this reason, based on the RAG (Retrieval-Augmented Generation) method, by combining external knowledge retrieval with generative large models, it provides strong support for question-and-answer systems based on large models. The RAG method first retrieves relevant information from the database before generating an answer, and uses the retrieved knowledge as context input to the model, thereby enhancing the knowledge coverage of the large model and improving its accuracy and reliability.
[0004] And the information in the database usually comes from external documents. It is necessary to first segment the external documents into multiple fragments, and then convert each fragment into a vector and store it in the database. Whether the external document segmentation is reasonable will affect the integrity of the information in the context retrieved by RAG, and further affect the accuracy and reliability of the question-and-answer output feedback by RAG. Summary of the Invention
[0005] The purpose of the present invention is to provide a text segmentation method, apparatus, electronic device, and computer-readable storage medium, which can reasonably segment the text according to semantics, ensure the semantic coherence of the segmented semantic segments, and further improve the accuracy and reliability of the question-and-answer output feedback by RAG.
[0006] The embodiments of the present invention can be implemented as follows:
[0007] In a first aspect, the present invention provides a text segmentation method, the method comprising:
[0008] Obtaining a plurality of sentences obtained by performing sentence-level segmentation on a target text;
[0009] Starting from the first sentence and ending with the last sentence among the multiple sentences, slide a preset sliding window multiple times. After each slide, segment the multiple target sentences within the preset sliding window semantically to obtain at least one semantic segment;
[0010] For each semantic segment obtained by segmentation, cluster the semantic segment obtained by this segmentation and the semantic segment obtained by the previous segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window belongs after the previous slide and this slide, until the semantic segment to which each sentence among the multiple sentences belongs is obtained.
[0011] In an alternative embodiment, the step of segmenting the multiple target sentences within the preset sliding window semantically to obtain at least one semantic segment includes:
[0012] Input the multiple target sentences into a pre-trained text segmentation model to obtain the semantic continuity of sentence pairs formed by two adjacent target sentences;
[0013] For any sentence pair, if the semantic continuity of the sentence pair is greater than a preset value, attribute the two target sentences in the sentence pair to the same semantic segment; otherwise, attribute the two target sentences in the sentence pair to different semantic segments;
[0014] Successively take each sentence pair as the target sentence pair to obtain at least one semantic segment after semantic segmentation of the multiple target sentences.
[0015] In an alternative embodiment, the first target sentence within the preset sliding window after this slide is the same as the last target sentence within the preset sliding window after the previous slide, and sentence pairs are formed by two adjacent target sentences within the preset sliding window;
[0016] The step of clustering the semantic segment obtained by this segmentation and the semantic segment obtained by the previous segmentation for each semantic segment obtained by segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window belongs after the previous slide and this slide includes:
[0017] Obtain the semantic continuity of the first sentence pair within the preset sliding window after this slide;
[0018] Calculate the semantic similarity between the two target sentences in the first sentence pair;
[0019] Calculate the semantic association score of the first sentence pair according to the semantic continuity and semantic similarity between the two target sentences in the first sentence pair;
[0020] If the semantic association score is greater than a preset score, merge the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from the current segmentation into one semantic segment; otherwise, use the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from the current segmentation as different semantic segments.
[0021] In an alternative embodiment, after the step of clustering the semantic segments obtained from each segmentation, i.e., the semantic segments obtained from the current segmentation and the semantic segments obtained from the previous segmentation, to obtain the target semantic segments to which each target sentence within the preset sliding window belongs after the previous slide and after the current slide, the method further includes:
[0022] Obtain a preset constraint condition, where the preset constraint condition is used to limit the semantic segments input to the large model in a Retrieval-Augmented Generation (RAG) system;
[0023] According to the preset constraint condition, process each of the target semantic segments to obtain a corresponding reference semantic segment for each of the target semantic segments, and store each of the reference semantic segments in a vector database in a vectorized form, so that when the large model receives input data, retrieve semantic segment vectors related to the input data from the vector database, and provide feedback based on the input data and the retrieved semantic segment vectors.
[0024] In an alternative embodiment, the preset constraint condition is a preset length, and the sum of the lengths of all sentences in the target semantic segment is greater than the preset length; the step of processing each of the target semantic segments according to the preset constraint condition to obtain a corresponding reference semantic segment for each of the target semantic segments includes:
[0025] For any one of the target semantic segments, if the target semantic segment does not meet the preset constraint condition, then re-segment the target semantic segment according to the preset constraint condition to obtain the reference semantic segment corresponding to the target semantic segment;
[0026] If the target semantic segment meets the preset constraint condition, then use the target semantic segment as its corresponding reference semantic segment.
[0027] In an alternative embodiment, the step of re-segmenting the target semantic segment according to the preset constraint condition to obtain the reference semantic segment corresponding to the target semantic segment includes:
[0028] Use the first sentence in the target semantic segment as a reference sentence;
[0029] Starting from the reference sentence, determine a target sentence sequence from the target semantic segment, where the sum of the lengths of all sentences in the target sentence sequence is the largest and less than the preset length;
[0030] Group all the sentences in the target sentence sequence into a reference semantic segment;
[0031] If the last sentence in the target sentence sequence is not the last sentence of the target semantic segment, re-determine a reference sentence from the target semantic segment, and return the step of determining the target sentence sequence from the target semantic segment starting from the reference sentence until all the reference semantic segments corresponding to the target semantic segment are obtained.
[0032] In an optional implementation manner, the step of re-determining a reference sentence from the target semantic segment includes:
[0033] Determine a first associated sentence from the target sentence sequence;
[0034] Determine a second associated sentence from the sentences in the target semantic segment that are after the target sentence sequence;
[0035] If the sum of the lengths of the first associated sentence and the second associated sentence is greater than the preset length, determine the reference sentence according to the second associated sentence, otherwise determine the reference sentence according to the first associated sentence.
[0036] In a second aspect, the present invention provides a text segmentation device, and the device includes:
[0037] An acquisition module, configured to acquire a plurality of sentences obtained by performing sentence-level segmentation on a target text;
[0038] A segmentation module, configured to use the first sentence in the plurality of sentences as a starting point and the last sentence as an ending point, slide a preset sliding window multiple times, and after each slide, segment a plurality of target sentences within the preset sliding window according to semantics to obtain at least one semantic segment;
[0039] A clustering module, configured to, for each semantic segment obtained by each segmentation, cluster the semantic segment obtained by this segmentation and the semantic segment obtained by the previous segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window belongs after the previous slide and after this slide, until the semantic segment to which each sentence in the plurality of sentences belongs is obtained.
[0040] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the memory is used to store a program, and the processor is configured to, when executing the program, implement the text segmentation method as described in any one of the foregoing implementation manners.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the text segmentation method as described in any one of the foregoing implementation manners is implemented.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] For multiple sentences of the present invention, a preset sliding window is slid, and the target sentences within the preset sliding window are segmented semantically in sequence to ensure the semantic coherence of the target sentences within the preset sliding window after segmentation. Moreover, the semantic segments obtained from the current segmentation and the semantic segments obtained from the previous segmentation are clustered to ensure the semantic coherence of the sentences between the preset sliding windows after two adjacent slides. Finally, reasonable segmentation of the text is achieved, ensuring the semantic coherence of the segmented semantic segments, thereby improving the accuracy and reliability of the Q&A output of the RAG feedback. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 It is a framework diagram of the RAG system provided in this embodiment.
[0046] Figure 2 It is a flow example diagram of the text segmentation method provided in this embodiment.
[0047] Figure 3 It is a process example diagram of semantic segmentation using a preset sliding window provided in this embodiment.
[0048] Figure 4 It is a block example diagram of the text segmentation device provided in this embodiment.
[0049] Figure 5 It is a block example diagram of the electronic device provided in this embodiment.
[0050] Icons: 10 - Electronic device; 11 - Processor; 12 - Memory; 13 - Bus; 100 - Text segmentation device; 110 - Acquisition module; 120 - Segmentation module; 130 - Clustering module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0052] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0053] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0054] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0055] In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0056] It should be noted that, without conflict, the features in the embodiments of the present invention can be combined with each other.
[0057] In the existing RAG system, before generating an answer, relevant information is first retrieved from external documents, and the retrieved knowledge is used as context to input into the large model, thereby enhancing the knowledge coverage of the large model and improving its accuracy and reliability. Please refer to Figure 1 , Figure 1 which is the framework diagram of the RAG system provided for this embodiment. Figure 1 In
[0058] (1) Perform text segmentation on the document to obtain segmented text segments;
[0059] (2) Vectorize and store the segmented text segments to generate a vector database;
[0060] (3) When processing a specified question, perform vector retrieval on the vector database, retrieve vectors associated with the specified question from the vector database, restore the retrieved vectors into corresponding text segments as the context of the specified question, and construct a prompt for the large model in the RAG system based on the specified question and its context.
[0061] (4) Generate question-and-answer outputs based on the prompt words.
[0062] In the RAG framework, the segmentation strategies adopted for document segmentation mostly use a fixed segmentation granularity. The segmentation granularity is usually associated with the input length limit of the large model. Although this fixed segmentation granularity can meet the input requirements of the large model, it will inevitably have the problem of semantic fragmentation. Moreover, for different large models, the input length limits are also different, and the fixed-granularity segmentation method cannot flexibly adapt to the requirements of different scenarios.
[0063] In view of this, the present embodiment provides a text segmentation method, apparatus, electronic device, and computer-readable storage medium, which can reasonably segment the text according to semantics and ensure the semantic coherence of the segmented semantic segments. The following will describe it in detail.
[0064] Please refer to Figure 2 , Figure 2 which is a flowchart example of the text segmentation method provided in this embodiment. The method includes the following steps:
[0065] Step S101: Obtain multiple sentences obtained by performing sentence-level segmentation on the target text.
[0066] In this embodiment, the target text can be text extracted from one or more documents. The documents can be word documents, txt documents, pdf documents, etc., or text extracted after performing optical character recognition on pictures.
[0067] In this embodiment, a pre-trained model can be used to perform sentence-level segmentation on the target text, or the target text can be segmented at the sentence level according to punctuation marks and line break symbols in the target text.
[0068] Step S102: Use the first sentence among the multiple sentences as the starting point and the last sentence as the ending point, slide a preset sliding window multiple times, and after each slide, segment the multiple target sentences within the preset sliding window according to semantics to obtain at least one semantic segment.
[0069] In this embodiment, the size of the preset sliding window affects the segmentation efficiency and accuracy. The window size represents the number of sentences that need to be semantically segmented each time, and can be adjusted according to the business application scenario. For example, the preset sliding window size is 2.
[0070] In this embodiment, semantic segmentation can be performed using a pre-trained language model, including but not limited to the BERT-NSP model (a deep learning model that uses the Transformer architecture to process natural language processing), the BGE model (a text embedding model designed to convert text into low-dimensional dense vectors for efficient semantic analysis and retrieval), and the SimCSE (Similarity Consensus Semantic Encoding) model (a model for generating sentence embeddings), etc.
[0071] Step S103: For each semantic segment obtained by segmentation, cluster the semantic segment obtained in this segmentation and the semantic segment obtained in the previous segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window after the previous slide and after the current slide belongs, until the semantic segment to which each sentence among the multiple sentences belongs is obtained.
[0072] In this embodiment, to ensure the semantic coherence between the semantic segments obtained by two segmentations, if the last semantic segment obtained by the previous segmentation and the first semantic segment obtained by the current segmentation are semantically coherent, the two semantic segments need to be merged into one semantic segment; otherwise, they are regarded as two independent semantic segments. For example, if the last semantic segment 1 obtained by the previous segmentation includes sentences a and b, and the first semantic segment 2 obtained by the current segmentation includes sentences b, c, and d, then the semantic segment 1 and the semantic segment 2 are semantically coherent, and the semantic segment 1 and the semantic segment 2 can be merged into one semantic segment.
[0073] In this embodiment, as a specific implementation method, it is possible to perform clustering once after one slide, or to perform successive clustering on the semantic segments obtained by semantic segmentation after each slide after multiple slides. Eventually, the semantic segment to which each sentence among the multiple sentences belongs can be obtained.
[0074] The above method provided in this embodiment performs semantic segmentation on the target sentences within the preset sliding window in sequence, ensuring the semantic coherence of the target sentences within the preset sliding window after segmentation, and clustering the semantic segment obtained in this segmentation and the semantic segment obtained in the previous segmentation to ensure the semantic coherence of the sentences between the preset sliding windows after two adjacent slides, thus realizing reasonable segmentation of the text.
[0075] In an alternative embodiment, to perform semantic segmentation on multiple target sentences, it is possible to determine whether they are semantically continuous between two adjacent target sentences for segmentation. One implementation method is as follows:
[0076] First, input the multiple target sentences into a pre-trained text segmentation model to obtain the semantic continuity degree of the sentence pairs composed of two adjacent target sentences.
[0077] In this embodiment, the text segmentation model can be a model based on BERT-NSP, BGE, and SimCSE (Similarity Consensus Semantic Encoding). For the text segmentation model based on BERT-NSP, during the construction of the training dataset, sentences can be misaligned annotated according to the document content, that is, when processing the document, identify and annotate those sentences that are inconsistent with the overall content of the document or contain errors, so as to automatically correct errors in the document based on the misaligned annotation, improve the accuracy of extracting information from the document, and accurately evaluate the overall quality and consistency of the document, etc.
[0078] In addition, in the case where the samples between paragraphs in the dataset are scarce, an oversampling strategy can also be adopted, that is, by increasing the number of samples of the minority class to balance the number of samples between different classes. Common oversampling strategies include but are not limited to random oversampling and interpolating between samples of the minority class to generate new synthetic samples, so as to ensure the sample balance of the training data and thus improve the generalization ability of the model.
[0079] Secondly, for any target sentence pair, if the semantic continuity degree of the target sentence pair is greater than the preset value, the two target sentences in the target sentence pair belong to the same semantic segment; otherwise, the two target sentences in the target sentence pair belong to different semantic segments.
[0080] In this embodiment, as an implementation manner, the semantic continuity degree obtained by the text segmentation model can be the probability value of semantic continuity between two target sentences. If it is greater than the preset value, it is considered that these two target sentences belong to the same semantic segment; otherwise, it is considered that the two target sentences belong to different semantic segments. As another implementation manner, the semantic continuity degree obtained by the text segmentation model can be a string composed of 0 and / or 1, where 0 indicates that the semantic between the corresponding two target sentences is discontinuous, and 1 indicates that the semantic between the corresponding two target sentences is continuous. For example, there are multiple target sentences: sentence a, sentence b, sentence c, and sentence d. Sentence a and sentence b form sentence pair 1, sentence b and sentence c form sentence pair 2, and sentence c and sentence d form sentence pair 3. The string output by the text segmentation model is 011. Then the "0" in it represents the semantic continuity degree of sentence pair 1, indicating that the semantic between sentence a and sentence b is discontinuous. The first "1" represents the semantic continuity degree of sentence pair 2, indicating that the semantic between sentence b and sentence c is continuous. The second "1" represents the semantic continuity degree of sentence pair 3, that is, the semantic between sentence c and sentence d is continuous.
[0081] Finally, each sentence pair is sequentially used as the target sentence pair to obtain at least one semantic segment after semantic segmentation of multiple target sentences.
[0082] In this embodiment, sentences with continuous semantics among multiple target sentences are divided into the same semantic segment, and sentences with discontinuous semantics are divided into different semantic segments, obtaining at least one semantic segment after semantic segmentation of the multiple target sentences.
[0083] In an alternative embodiment, to ensure semantic coherence between adjacent segmented semantic segments, the last target sentence within the preset sliding window during the previous slide is used as the first target sentence within the preset sliding window after this slide, that is, the first target sentence within the preset sliding window after this slide is the same as the last target sentence within the preset sliding window during the previous slide. Based on this, an implementation method for clustering the semantic segment obtained from this segmentation and the semantic segment obtained from the previous segmentation is as follows:
[0084] First, obtain the semantic continuity degree of the first sentence pair within the preset sliding window after this slide.
[0085] In this embodiment, the semantic continuity degree of the first sentence pair can be obtained by inputting the two sentences in the first sentence pair into the text segmentation model described above.
[0086] Second, calculate the semantic similarity between the two target sentences in the first sentence pair.
[0087] In this embodiment, one implementation method is: the semantic similarity can be obtained by inputting the two target sentences into models such as BGE, M3E (Multi-Modal Multi-Task Embedding); another implementation method can be: by separately extracting the feature vectors of the two target sentences, then calculating the cosine similarity of the two feature vectors, and taking the calculated cosine similarity as the semantic similarity.
[0088] Third, calculate the semantic association score of the first sentence pair according to the semantic continuity degree and semantic similarity between the two target sentences in the first sentence pair.
[0089] In this embodiment, a continuity weight and a similarity weight can be set for the semantic continuity degree and the semantic similarity respectively, and the semantic association score = semantic continuity degree × continuity weight + semantic similarity × similarity weight. The semantic continuity degree and the semantic similarity can be set according to the actual application scenario and the reliability of their respective corresponding models or the matching degree with the actual application scenario. For example, the continuity weight is set to 0.7 and the similarity weight is set to 0.3.
[0090] Finally, if the semantic association score is greater than the preset score, the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from this segmentation are merged into one semantic segment; otherwise, the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from this segmentation are used as different semantic segments.
[0091] In this embodiment, the preset score can be set according to actual needs. For example, the preset score is 0.5.
[0092] It should be noted that, when the semantic segmentation efficiency permits, another implementation method of semantic segmentation of multiple target sentences within the preset sliding window according to the semantic association score after each sliding is also possible:
[0093] Obtain the semantic continuity of sentence pairs formed by two adjacent target sentences;
[0094] Calculate the semantic similarity of sentence pairs formed by two adjacent target sentences;
[0095] Calculate the semantic association score of sentence pairs formed by two adjacent target sentences according to the semantic continuity and semantic similarity between the two target sentences in the sentence pairs formed by two adjacent target sentences;
[0096] If the semantic association score is greater than the preset score, the target sentences in the corresponding sentence pair belong to the same semantic segment; otherwise, they belong to different semantic segments.
[0097] Furthermore, in order to improve the efficiency of semantic segmentation, it is also possible to only calculate the semantic similarity of the to-be-confirmed semantic pairs with semantic continuity less than or equal to the preset value among multiple target sentences, calculate the semantic association score of the to-be-confirmed semantic pairs according to the semantic continuity and semantic similarity of the to-be-confirmed semantic pairs. If the semantic association score is greater than the preset score, the target sentences in the to-be-confirmed semantic pair belong to the same semantic segment; otherwise, they belong to different semantic segments. Since it only targets the to-be-confirmed semantic pairs with semantic continuity less than or equal to the preset value and determines whether the target sentences included belong to the same semantic segment by calculating the semantic association score, the amount of calculation is reduced to a certain extent. Compared with the method of calculating the semantic association score of each sentence pair in the target sentences, the efficiency of semantic segmentation is improved to a certain extent.
[0098] To more intuitively show the process of semantic segmentation of multiple sentences during the sliding of the preset sliding window, please refer to Figure 3 , Figure 3 which is a process example diagram of semantic segmentation using the preset sliding window provided in this embodiment. Figure 3Among them, for sentences 1, 2, and 3 within the preset sliding window, semantic segmentation is performed to obtain: semantic segments [sentence 1, sentence 2] and semantic segment [sentence 3]. The preset sliding window is slid, and for sentences 3, 4, and 5 within the preset sliding window, semantic segmentation is performed to obtain: semantic segment [sentence 3, sentence 4, sentence 5]. The semantic segment [sentence 3, sentence 4, sentence 5] obtained this time, the semantic segments [sentence 1, sentence 2] and semantic segment [sentence 3] obtained last time are clustered to obtain semantic segments [sentence 1, sentence 2] and semantic segment [sentence 3, sentence 4, sentence 5].
[0099] To demonstrate the effect of the semantic segmentation method provided in this embodiment, this embodiment performs segmentation on 19 segments of sentences with manual annotation using three segmentation methods and evaluates and compares the segmentation results of the three segmentation methods. Method 1: Fixed-word segmentation; Method 2: Segmentation using the preset sliding window and the BERT-NSP model in this embodiment; Method 3: Segmentation using the preset sliding window and the BERT-NSP+M3E model in this embodiment.
[0100] The segmentation result of Method 1: 18 segments, the number of correctly segmented segments is 2; P = 11.11%, R = 94.73%, F1 = 19.87%;
[0101] The segmentation result of Method 2: 16 segments, the number of correctly segmented segments is 13, P = 81.25%, R = 84.21%, F1 = 82.70%;
[0102] The segmentation result of Method 3: 18 segments, the number of correctly segmented segments is 16, P = 88.88%, R = 94.73%, F1 = 91.71%;
[0103] Among them, P represents the precision rate, and the calculation method is: the number of correctly segmented segments / the total number of segmented segments; R represents the recall rate, and the calculation method is: the number of correctly segmented segments / the total number of segments that should be segmented; F1 represents the segmentation result score, and the calculation method is: 2×precision rate×recall rate / (precision rate + recall rate).
[0104] In this embodiment, in order to enable the semantic segments after segmentation to flexibly adapt to the input requirements of different large models, after obtaining the target semantic segments to which each target sentence within the preset sliding window belongs after the last slide and the current slide, it is also necessary to process the target semantic segments according to the input requirements of the required large model. This embodiment provides a processing method:
[0105] First, obtain the preset restriction conditions, and the preset restriction conditions are used to restrict the semantic segments input to the large model in the retrieval augmented generation (RAG) system.
[0106] In this embodiment, the preset restriction condition can be the maximum length of a single semantic segment input to the large model. For example, the maximum length of each semantic segment does not exceed 500 words. It can also be the overlap degree between two adjacent semantic segments input to the large model. For example, the first semantic segment is from sentence 1 to sentence 10, and the second semantic segment is from sentence 6 to sentence 15.
[0107] It should be noted that the preset restriction condition can also combine the longest length and the overlap degree, that is, both the length of a single semantic segment input to the large model is restricted, and at the same time, the overlap degree between two adjacent semantic segments is restricted.
[0108] Secondly, according to the preset restriction condition, each target semantic segment is processed to obtain a reference semantic segment corresponding to each target semantic segment, and each reference semantic segment is stored in the vector database in a vectorized form, so that when the large model receives input data, a semantic segment vector related to the input data can be retrieved from the vector database, and feedback is performed based on the input data and the retrieved semantic segment vector.
[0109] In this embodiment, the reference semantic segment is a semantic segment that meets the preset restriction condition. A target semantic segment can correspond to one reference semantic segment, and a target semantic segment can also correspond to multiple reference semantic segments. When a target semantic segment corresponds to multiple reference semantic segments, in order to ensure the semantic coherence between the multiple reference semantic segments, there can be partial overlap between the multiple reference semantic segments.
[0110] In an alternative embodiment, for a large model with limited input length, in order to meet the length requirements of the large model while maintaining the semantic coherence between sentences as much as possible, this embodiment provides an implementation method for processing the target semantic segment to obtain a reference semantic segment that meets the input length requirements of the large model:
[0111] First, for any target semantic segment, if the target semantic segment does not meet the preset restriction condition, the target semantic segment is split again according to the preset restriction condition to obtain a reference semantic segment corresponding to the target semantic segment;
[0112] Secondly, if the target semantic segment meets the preset restriction condition, the target semantic segment is used as its corresponding reference semantic segment.
[0113] In this embodiment, the preset restriction condition is a preset length. If the target semantic segment meets the preset restriction condition, it means that the sum of the lengths of all sentences in the target semantic segment is less than or equal to the preset length, then the target semantic segment does not need to be processed, and the target semantic segment is the reference semantic segment. Otherwise, it means that the sum of the lengths of all sentences in the target semantic segment is greater than the preset length, then the target semantic segment needs to be split again so that the split semantic segment meets the preset restriction condition. An implementation method of the re-splitting is:
[0114] (1) Use the first sentence in the target semantic segment as the reference sentence;
[0115] (2) Starting from the reference sentence, determine the target sentence sequence from the target semantic segment, where the sum of the lengths of all sentences in the target sentence sequence is the largest and less than the preset length;
[0116] (3) Group all sentences in the target sentence sequence into a reference semantic segment;
[0117] (4) If the last sentence in the target sentence sequence is not the last sentence in the target semantic segment, re-determine the reference sentence from the target semantic segment and return to step (2) until all reference semantic segments corresponding to the target semantic segment are obtained.
[0118] In an alternative embodiment, in order to ensure the semantic coherence between the finally obtained reference semantic segments as much as possible during re-segmentation, this embodiment provides an implementation method for determining the reference sentence:
[0119] First, determine the first associated sentence from the target sentence sequence;
[0120] Second, determine the second associated sentence from the sentences in the target semantic segment that are after the target sentence sequence;
[0121] Finally, if the sum of the lengths of the first associated sentence and the second associated sentence is greater than the preset length, determine the reference sentence according to the second associated sentence, otherwise determine the reference sentence according to the first associated sentence.
[0122] In this embodiment, the method of determining the reference sentence according to the second associated sentence can be to use the first sentence of the second associated sentence as the reference sentence, or to use the first sentence of several subsequent sentences in the first associated sentence as the reference sentence, as long as the sum of the lengths of several subsequent sentences in the first associated sentence and the second associated sentence is less than or equal to the preset length.
[0123] In this embodiment, in this embodiment, the number of the first associated sentence and the second associated sentence can both be one or more, specifically determined according to the maximum allowable overlap number during segmentation. The first associated sentence can be the sentences in the target sentence sequence that are after a preset number. As an implementation method, the preset number can be fixed and determined according to the maximum allowable overlap number during segmentation. For example, the preset number is set to 2. As another implementation method, the preset number can also be variable, as long as the sum of the lengths of the first associated sentence and the second associated sentence is less than or equal to the preset length.
[0124] This embodiment takes the number of the first associated sentences as 2 and the second associated sentence as 1 as an example for illustration. For example, the sentences included in the target semantic segment are as follows:
[0125]
Sentence 1, Sentence 2, Sentence 3, Sentence 4, Sentence 5, Sentence 6, Sentence 7
[0126] Taking Sentence 1 as the reference sentence, the sum of the lengths of Sentence 1, Sentence 2, and Sentence 3 is the largest and less than the preset length, then the target sentence sequence is: (Sentence 1, Sentence 2, Sentence 3), and Sentence 1, Sentence 2, and Sentence 3 belong to the reference semantic segment 1;
[0127] Taking Sentence 2 and Sentence 3 as the first associated sentences and Sentence 4 as the second associated sentence, if the sum of the lengths of Sentence 2, Sentence 3, and Sentence 4 is less than the preset length, then take Sentence 2 as the reference sentence;
[0128] The sum of the lengths of Sentence 2, Sentence 3, Sentence 4, and Sentence 5 is the largest and less than the preset length, then the target sentence sequence is: (Sentence 2, Sentence 3, Sentence 4, Sentence 5), and Sentence 2, Sentence 3, Sentence 4, and Sentence 5 belong to the reference semantic segment 2;
[0129] Taking Sentence 4 and Sentence 5 as the first associated sentences and Sentence 6 as the second associated sentence, the sum of the lengths of Sentence 4, Sentence 5, and Sentence 6 is greater than the preset length, then Sentence 6 can be taken as the reference sentence. If the sum of the lengths of Sentence 5 and Sentence 6 is less than the preset length, then Sentence 5 can also be taken as the reference sentence; in this example, Sentence 5 is taken as the reference sentence;
[0130] The sum of the lengths of Sentence 5, Sentence 6, and Sentence 7 is less than the preset length, then the target sentence sequence is (Sentence 5, Sentence 6, Sentence 7), and Sentence 5, Sentence 6, and Sentence 7 belong to the reference semantic segment 3;
[0131] Thus, the target semantic segment is sliced into 3 reference semantic segments again: reference semantic segment 1 to reference semantic segment 3.
[0132] In this embodiment, the preset limiting condition can be not only the preset length but also the overlap number. Taking the overlap number as 2, for the target semantic segment
Sentence 1, Sentence 2, Sentence 3, Sentence 4, Sentence 5
[0133] To execute the corresponding steps in the above embodiments and each possible implementation manner, an implementation manner of a text segmentation device 100 is given below. Please refer to Figure 4 , Figure 4 which is a block diagram of the text segmentation device provided in this embodiment. It should be noted that for the text segmentation device 100 provided by the present invention, its basic principle, the generated technical effects are the same as those of the corresponding above embodiments. For the sake of brief description, they are not mentioned in this embodiment.
[0134] The text segmentation device 100 includes an acquisition module 110, a segmentation module 120, and a clustering module 130.
[0135] The acquisition module 110 is configured to acquire a plurality of sentences obtained by performing sentence-level segmentation on a target text.
[0136] The segmentation module 120 is configured to use the first sentence among the plurality of sentences as a starting point and the last sentence as an ending point, slide a preset sliding window multiple times, and after each slide, segment a plurality of target sentences within the preset sliding window according to semantics to obtain at least one semantic segment.
[0137] The clustering module 130 is configured to, for each semantic segment obtained by each segmentation, cluster the semantic segment obtained by this segmentation and the semantic segment obtained by the previous segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window belongs after the previous slide and this slide, until the semantic segment to which each sentence in the plurality of sentences belongs is obtained.
[0138] In an alternative implementation manner, the segmentation module 120 is specifically configured to:
[0139] Input the plurality of target sentences into a pre-trained text segmentation model to obtain the semantic continuity of sentence pairs formed by two adjacent target sentences;
[0140] For any sentence pair, if the semantic continuity of the sentence pair is greater than a preset value, attribute the two target sentences in the sentence pair to the same semantic segment, otherwise, attribute the two target sentences in the sentence pair to different semantic segments;
[0141] Successively use each sentence pair as the sentence pair to obtain at least one semantic segment after semantic segmentation of the plurality of target sentences.
[0142] In an alternative implementation manner, the first target sentence within the preset sliding window after this slide is the same as the last target sentence within the preset sliding window after the previous slide, and sentence pairs are formed by two adjacent target sentences within the preset sliding window;
[0143] For each semantic segment obtained by each segmentation, the clustering module 130 is specifically configured to:
[0144] Obtain the semantic continuity of the first sentence pair within the preset sliding window after this sliding;
[0145] Calculate the semantic similarity between the two target sentences in the first sentence pair;
[0146] Calculate the semantic association score of the first sentence pair according to the semantic continuity and semantic similarity between the two target sentences in the first sentence pair;
[0147] If the semantic association score is greater than the preset score, merge the last semantic segment obtained by the previous segmentation and the first semantic segment obtained by this segmentation into one semantic segment; otherwise, use the last semantic segment obtained by the previous segmentation and the first semantic segment obtained by this segmentation as different semantic segments.
[0148] In an alternative embodiment, the clustering module 130 is further configured to:
[0149] Obtain a preset restriction condition, where the preset restriction condition is used to restrict the semantic segments input to the large model in the retrieval-augmented generation (RAG) system;
[0150] Process each target semantic segment according to the preset restriction condition to obtain a reference semantic segment corresponding to each target semantic segment, and store each reference semantic segment in a vector database in a vectorized form, so that when the large model receives input data, retrieve a semantic segment vector related to the input data from the vector database, and provide feedback based on the input data and the retrieved semantic segment vector.
[0151] In an alternative embodiment, the preset restriction condition is a preset length, and the sum of the lengths of all sentences in the target semantic segment is greater than the preset length; when the clustering module 130 is used to process each target semantic segment according to the preset restriction condition to obtain a reference semantic segment corresponding to each target semantic segment, it is specifically configured to:
[0152] For any target semantic segment, if the target semantic segment does not meet the preset restriction condition, re-segment the target semantic segment according to the preset restriction condition to obtain a reference semantic segment corresponding to the target semantic segment;
[0153] If the target semantic segment meets the preset restriction condition, use the target semantic segment as its corresponding reference semantic segment.
[0154] In an alternative embodiment, when the clustering module 130 is configured to further segment the target semantic segment according to the preset constraint conditions to obtain the reference semantic segments corresponding to the target semantic segment, it is specifically configured to:
[0155] Take the first sentence in the target semantic segment as the reference sentence;
[0156] Starting from the reference sentence, determine a target sentence sequence from the target semantic segment, where the sum of the lengths of all sentences in the target sentence sequence is the largest and less than the preset length;
[0157] Group all the sentences in the target sentence sequence into one reference semantic segment;
[0158] If the last sentence in the target sentence sequence is not the last sentence of the target semantic segment, re-determine the reference sentence from the target semantic segment, and return to the step of determining the target sentence sequence from the target semantic segment starting from the reference sentence until all the reference semantic segments corresponding to the target semantic segment are obtained.
[0159] In an alternative embodiment, when the clustering module 130 is configured to re-determine the reference sentence from the target semantic segment, it is specifically configured to:
[0160] Determine a first associated sentence from the target sentence sequence;
[0161] Determine a second associated sentence from the sentences in the target semantic segment that are after the target sentence sequence;
[0162] If the sum of the lengths of the first associated sentence and the second associated sentence is greater than the preset length, determine the reference sentence according to the second associated sentence, otherwise determine the reference sentence according to the first associated sentence.
[0163] An embodiment of the present invention further provides a block diagram of the electronic device 10. For the electronic device 10 to implement the text segmentation method of the foregoing embodiment, please refer to Figure 5 , Figure 5 which is the block diagram of the electronic device 10 provided in this embodiment. The electronic device 10 includes a processor 11, a memory 12, and a bus 13. The processor 11 and the memory 12 are connected through the bus 13.
[0164] The processor 11 may be an integrated circuit chip with signal processing capability. In the implementation process, each step of the text segmentation method of the above embodiment may be completed by an integrated logic circuit of hardware in the processor 11 or by instructions in the form of software. The above processor 11 may be a general-purpose processor, including a CPU (Central Processing Unit), a NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Logic Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0165] The memory 12 is used to store a program for implementing the text segmentation method. The program may be a software function module stored in the memory 12 in the form of software or firmware or solidified in the OS (Operating System) of the electronic device 10.
[0166] After receiving the execution instruction, the processor 11 executes the program to implement the text segmentation method of the above embodiment.
[0167] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the text segmentation method as described above is implemented.
[0168] In summary, the embodiments of the present invention provide a text segmentation method, apparatus, electronic device, and computer-readable storage medium. The method includes: obtaining a plurality of sentences obtained by performing sentence-level segmentation on a target text; starting from the first sentence among the plurality of sentences and ending with the last sentence, sliding a preset sliding window multiple times, and after each sliding, segmenting the plurality of target sentences within the preset sliding window semantically to obtain at least one semantic segment; for each semantic segment obtained by segmentation, clustering the semantic segment obtained in this segmentation and the semantic segment obtained in the previous segmentation to obtain the target semantic segment to which each target sentence within the preset sliding window belongs after the previous sliding and this sliding, until the semantic segment to which each sentence among the plurality of sentences belongs is obtained. Compared with the prior art, this embodiment has at least the following advantages: (1) For a plurality of sentences, slide the preset sliding window, and sequentially segment the target sentences within the preset sliding window semantically to ensure the semantic coherence of the target sentences after segmentation within the preset sliding window, and cluster the semantic segment obtained in this segmentation and the semantic segment obtained in the previous segmentation to ensure the semantic coherence of the sentences between the preset sliding windows after two adjacent slidings, ultimately realizing the reasonable segmentation of the text, ensuring the semantic coherence of the segmented semantic segments, and thus improving the accuracy and reliability of the Q&A output of the RAG feedback; (2) Utilize the semantic continuity and semantic similarity between the sentences of the preset sliding windows after two adjacent slidings to accurately judge the semantic coherence between them and improve the rationality of semantic segmentation; (3) Since the semantic coherence between the sentences within the semantic segment after semantic segmentation is ensured, the hardware resource requirements for the large model of RAG in the subsequent stage are reduced, making it more general and economical, achieving the purpose of cost reduction and efficiency improvement; (4) While segmenting, it can take into account the preset limiting conditions for the semantic segments used to limit the input of the large model in the RAG system, and process the segmented semantic segments according to the preset limiting conditions, so that the final segmentation result can meet the requirements of the large model in the RAG system while maximizing the semantic coherence, achieving the effect of being able to flexibly adapt to the requirements of different large models.
[0169] As described above, these are only various embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A text segmentation method, characterized in that: The method comprises: Obtain multiple sentences obtained after performing sentence-level segmentation on the target text; Sliding the preset sliding window multiple times with the first sentence of the multiple sentences as the starting point and the last sentence as the end point, and after each sliding, segmenting the multiple target sentences in the preset sliding window according to semantics to obtain at least one semantic segment; For the semantic segments obtained from each segmentation, the semantic segments obtained from this segmentation and the semantic segments obtained from the previous segmentation are clustered to obtain the target semantic segments to which each target sentence in the preset sliding window after the previous sliding and this sliding belongs, until the semantic segments to which each sentence in the multiple sentences belongs are obtained.
2. The text segmentation method according to claim 1, characterized in that: The step of segmenting the plurality of target sentences in the preset sliding window according to semantics to obtain at least one semantic segment comprises: Inputting the plurality of target sentences into a pre-trained text segmentation model to obtain semantic continuity of sentence pairs consisting of two adjacent target sentences; For any target sentence pair, if the semantic continuity of the target sentence pair is greater than a preset value, the two target sentences in the target sentence pair are attributed to the same semantic segment; otherwise, the two target sentences in the target sentence pair are attributed to different semantic segments; Each sentence pair is taken as the target sentence pair in turn to obtain at least one semantic segment after semantic segmentation of the multiple target sentences.
3. The text segmentation method according to claim 1, characterized in that: The first target sentence in the preset sliding window after the current sliding is the same as the last target sentence in the preset sliding window of the previous sliding, and the target sentences adjacent to each other in the preset sliding window constitute a sentence pair; The step of clustering the semantic segments obtained by each segmentation with the semantic segments obtained by the previous segmentation to obtain the target semantic segments to which each target sentence in the preset sliding window belongs after the previous sliding and the current sliding comprises: Obtaining the semantic continuity of the first sentence pair in the preset sliding window after the current sliding; Calculating the semantic similarity between the two target sentences in the first sentence pair; Calculating a semantic association score of the first sentence pair according to the semantic continuity and semantic similarity between the two target sentences in the first sentence pair; If the semantic association score is greater than the preset score, the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from the current segmentation are merged into one semantic segment; otherwise, the last semantic segment obtained from the previous segmentation and the first semantic segment obtained from the current segmentation are treated as different semantic segments.
4. The text segmentation method according to claim 1, characterized in that: After the step of clustering the semantic segments obtained by each segmentation with the semantic segments obtained by the previous segmentation to obtain the target semantic segments to which each target sentence in the preset sliding window belongs after the previous sliding and the current sliding, the method further comprises: Acquire a preset restriction condition, where the preset restriction condition is used to restrict the semantic segment of the large model input in the retrieval enhancement generation RAG system; According to the preset restriction conditions, each of the target semantic segments is processed to obtain a reference semantic segment corresponding to each of the target semantic segments, and each of the reference semantic segments is stored in a vectorized form in a vector database, so that when the large model receives input data, the semantic segment vector related to the input data can be retrieved from the vector database, and feedback can be provided based on the input data and the retrieved semantic segment vector.
5. The text segmentation method according to claim 4, characterized in that: The preset restriction condition is a preset length, and the sum of the lengths of all sentences in the target semantic segment is greater than the preset length; The step of processing each of the target semantic segments according to the preset restriction condition to obtain a reference semantic segment corresponding to each of the target semantic segments comprises: For any of the target semantic segments, if the target semantic segment does not satisfy the preset restriction condition, the target semantic segment is segmented again according to the preset restriction condition to obtain a reference semantic segment corresponding to the target semantic segment; If the target semantic segment meets the preset restriction condition, the target semantic segment is used as its corresponding reference semantic segment.
6. The text segmentation method according to claim 5, characterized in that: The step of segmenting the target semantic segment again according to the preset restriction condition to obtain a reference semantic segment corresponding to the target semantic segment comprises: Taking the first sentence in the target semantic segment as a reference sentence; Starting from the reference sentence, determining a target sentence sequence from the target semantic segment, wherein the sum of the lengths of all sentences in the target sentence sequence is the largest and less than the preset length; Classifying all sentences in the target sentence sequence into a reference semantic segment; If the last sentence in the target sentence sequence is not the last sentence of the target semantic segment, then the reference sentence is re-determined from the target semantic segment, and the step of starting from the reference sentence and determining the target sentence sequence from the target semantic segment is returned to until all reference semantic segments corresponding to the target semantic segment are obtained.
7. The text segmentation method according to claim 6, characterized in that: The step of re-determining the reference sentence from the target semantic segment comprises: Determining a first related sentence from the target sentence sequence; Determining a second related sentence from the sentences in the target semantic segment that follow the target sentence sequence; If the sum of the length of the first related sentence and the length of the second related sentence is greater than the preset length, the reference sentence is determined according to the second related sentence; otherwise, the reference sentence is determined according to the first related sentence.
8. A text segmentation device, characterized in that: The device comprises: An acquisition module, used to acquire multiple sentences obtained after performing sentence-level segmentation on the target text; a segmentation module, configured to slide a preset sliding window multiple times with the first sentence of the multiple sentences as a starting point and the last sentence as an end point, and after each sliding, segment the multiple target sentences in the preset sliding window according to semantics to obtain at least one semantic segment; The clustering module is used to cluster the semantic segments obtained from each segmentation with the semantic segments obtained from the previous segmentation, and obtain the target semantic segments to which each target sentence in the preset sliding window after the previous sliding and the current sliding belongs, until the semantic segments to which each sentence in the multiple sentences belongs are obtained.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a program, and the processor is used to implement the text segmentation method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the text segmentation method as described in any one of claims 1 to 7 is implemented.