Segmentation method for iterative multi-granularity document
Through the iterative multi-grained document slicing method, the deep learning model is used to uniformly process the segmentation of paragraphs, sentences and words, and solve the problem of slicing results in the existing technology that do not consider the overall semantics of the document, and improve the accuracy and consistency of slicing.
Patent Information
- Application Number
- CN202510184159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
AI Technical Summary
When slicing paragraphs, sentences and words, the existing document segmentation technology lacks a unified multi-grained slicing method, which leads to the slicing results that do not take into account the overall semantics of the document and are prone to introduce errors.
The iterative multi-grained document segmentation method is adopted to achieve the unified division of paragraphs, words and sentences by constructing a deep learning model that trains corpus and trains GPT structure. The method includes building a training corpus, training a slicing model, and segmenting the input document according to the model.
It improves the overall slicing semantics of the document and the accuracy of the slicing results, solves the problem that multi-grained slicing cannot be unified, and reduces the occurrence of slicing errors.
Smart Images

Figure CN120181069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence recognition technology, and particularly to a method for segmenting iterative multi-granularity documents. Background Art
[0002] In the scenario of automated composition grading, the accurate segmentation of paragraphs, sentences, and words will affect the accuracy of downstream tasks. For example, the segmentation accuracy will affect the statistics of measures such as word complexity and sentence complexity that measure the quality of a composition, and will also affect the performance of the model's grammar correction. Currently, the commonly used segmentation techniques for the overall paragraphs, sentences, and words in a composition are relatively rough, and only models are introduced for word-level segmentation. When segmenting paragraphs, only line break symbols are used for segmentation. When segmenting sentences, sentence end symbols are used for segmentation. This multi-stage segmentation method is separated from each other and the final result does not consider the overall segmentation effect. At the same time, the composition itself may contain punctuation and symbol errors, or the corrected composition corpus comes from pictures recognized by OCR, and OCR will introduce some recognition errors of punctuation or line break symbols. Therefore, segmentation errors may occur. In summary, in the current related document segmentation technology, when segmenting paragraphs, sentences, and words, the existing methods regard them as separate processes and fail to unify the multi-granularity segmentation together, without considering the overall semantics and segmentation results of the document. Summary of the Invention
[0003] In view of this, the present invention provides a method for segmenting iterative multi-granularity documents to solve the problem that multi-granularity segmentation cannot be unified, and improve the overall segmentation semantics and segmentation results of the document.
[0004] In a first aspect, the present invention provides a method for segmenting iterative multi-granularity documents, the method comprising:
[0005] Step 1, constructing a training corpus and performing segmentation on it at different granularities of paragraphs, words, and sentences, and the training corpus consists of an unsegmented document and a segmented document;
[0006] Step 2, training a deep learning model with a GPT structure through the training corpus to obtain a trained segmentation model;
[0007] Step 3, segmenting an input document according to the trained segmentation model and outputting a segmentation result.
[0008] Optionally, the segmented document in Step 1 includes special characters and fixed components;
[0009] The special characters include [B], [S], and [P]; [B] represents the segmentation position of a word; [S] represents the segmentation position of a sentence; [P] represents the segmentation position of a paragraph; by outputting the above special characters, the original training corpus is segmented at different granularities of paragraphs, words, and sentences;
[0010] The fixed composition includes two fixed parts: [Sentence to be segmented] and [Segmentation result]. [Sentence to be segmented] is at the beginning of the input, and [Segmentation result] is at the beginning of the segmentation result.
[0011] Optionally, when constructing the training corpus in step 1, assume that there is an intermediate segmentation state in the document between d and d final denoted by d i ; for any i < j, d j contains all the segmentation special characters in d i , and the number of segmentation special characters in d j is greater than that in d i , denoted by |d i | for the number of segmentation special characters in d i , then |d| = 0, |d| < |d i | < |d final |; the constructed training corpus also includes: the document d to be segmented i and the segmentation result d j .
[0012] Optionally, in step 3, the document d is input into the trained segmentation model, and the segmentation process is as follows:
[0013] Step 31: Construct the input including the document d to be segmented m and the segmentation result; input it into the trained segmentation model;
[0014] Step 32: Gradually predict the next token through the trained segmentation model until the output ends; for the output segmentation result, use the regular expression pattern to extract the segmentation result, denoted as d m+1 ;
[0015] Step 33: Adopt a segmentation end detection algorithm to determine whether d m+1 is used as the final segmentation result; if not, then use the segmentation result d m+1 as the document d to be segmented m , and repeat step 31; if so, execute step 34;
[0016] Step 34: Convert the special characters in the segmentation result d m+1 into the form of the segmentation result required by the downstream, and output the segmentation result.
[0017] Optionally, the segmentation end detection algorithm includes:
[0018] Step 331: Calculate the segmentation granularity of the segmentation result d m+1 , and use the number of segmentation special characters divided by the total number of characters to measure the segmentation granularity;
[0019] Step 332, compare d m+1 and d m for the change quantity of the special characters segmented therein; if the segmentation granularity > the segmentation granularity threshold and the change quantity of the segmented characters < the change quantity threshold of the segmented character quantity, it is considered that a stable segmentation result has been reached and the segmentation ends.
[0020] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute the segmentation method of the iterative multi-granularity document in the first aspect or any possible implementation manner of the first aspect.
[0021] In a third aspect, an embodiment of the present invention provides an electronic device, including: one or more processors; a memory; and one or more computer programs, where the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, and when the instructions are executed by the device, the device executes the segmentation method of the iterative multi-granularity document in the first aspect or any possible implementation manner of the first aspect.
[0022] In the technical solution provided by the present invention, the method includes constructing a training corpus and performing segmentations at different granularities of segments, words, and sentences on it. The training corpus consists of an unsegmented document and a segmented document; training a deep learning model with a GPT structure through the training corpus to obtain a trained segmentation model; segmenting an input document according to the trained segmentation model and outputting a segmentation result. This method solves the problem that multi-granularity segmentation cannot be unified and improves the overall segmentation semantics and segmentation result of the document. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is a flowchart of the segmentation method of the iterative multi-granularity document provided by the embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of the deep learning model with a GPT structure provided by the embodiment of the present invention;
[0026] Figure 3 is a schematic diagram of an electronic device provided by the embodiment of the present invention. Detailed Embodiments
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be clear that the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention are also intended to include the plural forms unless the context clearly indicates otherwise.
[0030] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, a / or b may represent: a exists alone, a and b exist simultaneously, and b exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0031] Depending on the context, the word "if" as used herein may be interpreted as "when", "while", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0032] Figure 1 It is a flowchart of the segmentation method for the iterative multi-granularity document provided by the embodiments of the present invention. As Figure 1 shown, the method includes:
[0033] Step 1: Construct a training corpus and perform segmentations at different granularities of paragraphs, words, and sentences on it. The training corpus consists of the unsegmented document and the segmented document.
[0034] In the embodiments of the present invention, the segmented document in Step 1 includes special characters and fixed components;
[0035] Special characters include [B], [S], and [P]; [B] represents the segmentation position of a word; [S] represents the segmentation position of a sentence; [P] represents the segmentation position of a paragraph; by outputting the above special characters, the original training corpus is segmented at different granularities of paragraphs, words, and sentences;
[0036] The fixed composition includes two fixed parts:
sentence to be segmented
segmentation result
sentence to be segmented
segmentation result
[0037] In the embodiment of the present invention, when constructing the training corpus in step 1, it is assumed that the document has an intermediate segmentation state between d and d final denoted by d i For any i < j, d j contains all the segmentation special characters in d i and the number of segmentation special characters contained in d j is greater than that in d i denoted by |d i |, then |d| represents the number of segmentation special characters in d i so |d| = 0, |d| < |d i | < |d final |; the constructed training corpus also includes: the document to be segmented d i and the segmentation result d j .
[0038] For example, when training "The weather is nice today. I came to the park.", the corpus input to the model is "Document to be segmented: The weather is nice today. I came to the park.\nSegmentation result: Today [B] weather [B] nice [B]. [S] I [B] came [B] to [B] the [B] park [B]."
[0039] Step 2: Train a deep learning model with the GPT structure using the training corpus to obtain a trained segmentation model.
[0040] GPT (Generative Pre-trained Transformer) is a natural language processing model based on deep learning. GPT uses the Transformer architecture, which is a neural network model particularly suitable for processing sequence data (such as natural language). It relies on the self-attention mechanism, enabling the model to be trained in parallel and better capture long-range dependencies.
[0041] During the training phase, the model learns on the constructed training corpus by predicting the next token. As Figure 2 shown, the model predicts the next token simultaneously during training. The document is first tokenized to obtain a list of words, and<start>Special string, add at the end <end>Special characters are used to form a new word list; each word in the list is first embedded and transformed into a 512-dimensional vector. Each vector passes through a GPT model to obtain the vector representation V of each word with context information. The vector of the last word is transformed into a vector of the vocabulary size through an FFN as the score logits_next of the next word. When calculating the loss, the "segmentation result" and the content before it are not used to calculate the loss, and the remaining tokens use the CrossEntropy loss.
[0042] This structure considers the context information of the sentence where the word is located when judging document word segmentation, sentence segmentation, and paragraph segmentation. At the same time, it can be segmented iteratively, improving the accuracy and efficiency.
[0043] Step 3: Segment the input document according to the trained segmentation model and output the segmentation result.
[0044] In the embodiment of the present invention, in step 3, the document d is input into the trained segmentation model, and the segmentation process is as follows:
[0045] Step 31: Construct the input including the document d to be segmented m and the segmentation result; input it into the trained segmentation model;
[0046] Step 32: Gradually predict the next token through the trained segmentation model until the output ends; for the output segmentation result, use the regular expression pattern to extract the segmentation result, denoted as d m+1 ;
[0047] Step 33: Adopt a segmentation end detection algorithm to judge whether d r+1 is used as the final segmentation result; if not, then use the segmentation result d m+1 as the document d to be segmented m , and repeat step 31; if so, execute step 34;
[0048] Step 34: Convert the special characters in the segmentation result d m+1 into the form of the segmentation result required by the downstream and output the segmentation result.
[0049] In the embodiment of the present invention, the segmentation end detection algorithm includes:
[0050] Step 331: Calculate the segmentation granularity of the segmentation result d m+1 , and use the number of special characters for segmentation divided by the total number of characters to measure the segmentation granularity;
[0051] Step 332: Compare d m+1 and d m The number of changes in special characters for mid-segmentation; if the segmentation granularity > the segmentation granularity threshold and the number of changes in segmented characters < the threshold for the change in the number of segmented characters, it is considered that a stable segmentation result has been achieved and the segmentation ends.
[0052] When existing technologies perform segmentation on documents, they use segment, sentence, and word separation-based segmentation and do not comprehensively consider the overall segmentation effect, which will introduce many segmentation errors. The present invention adopts an end-to-end method to effectively and efficiently solve the problem of unified segmentation of segments, sentences, and words in the scenario of Chinese compositions.
[0053] In the technical solution provided by the present invention, the method includes constructing a training corpus and performing segmentation on it at different granularities of segments, words, and sentences. The training corpus consists of an unsegmented document and a segmented document; training a deep learning model with a GPT structure through the training corpus to obtain a trained segmentation model; segmenting the input document according to the trained segmentation model and outputting the segmentation result. This method solves the problem that multi-granularity segmentation cannot be unified and improves the overall segmentation semantics and segmentation result of the document.
[0054] Each step of the embodiments of the present invention can be executed by an electronic device. Among them, the electronic device includes, but is not limited to, mobile phones, tablet computers, portable PCs, desktop computers, etc.
[0055] The embodiments of the present invention provide a computer-readable storage medium. The computer-readable storage medium includes a stored program. Among them, when the program runs, it controls the electronic device where the computer-readable storage medium is located to execute the embodiments of the above-mentioned iterative multi-granularity document segmentation method.
[0056] Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present invention, as Figure 3 shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, it implements the iterative multi-granularity document segmentation method in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0057] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art can understand that Figure 3 this is only an example of the electronic device 21 and does not constitute a limitation on the electronic device 21. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0058] The so-called processor 211 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0059] The memory 212 may be an internal storage unit of the electronic device 21, such as the hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 21. Further, the memory 212 may also include both the internal storage unit and the external storage device of the electronic device 21. The memory 212 is used to store computer programs and other programs and data required by the network device. The memory 212 may also be used to temporarily store data that has been output or is to be output.
[0060] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.< / end> < / start>
Claims
1. An iterative multi-granularity document segmentation method, characterized in that: The method comprises: Step 1: Construct a training corpus and segment it into segments, words, and sentences at different granularities. The training corpus consists of unsegmented documents and segmented documents. Step 2: Train the deep learning model of the GPT structure through the training corpus to obtain the trained segmentation model; Step 3: Segment the input document according to the trained segmentation model and output the segmentation results.
2. The method according to claim 1, characterized in that The segmented document in step 1 includes special characters and fixed components; Special characters include [B], [S], and [P]; [B] represents the segmentation position of a word; [S] represents the segmentation position of a sentence; [P] represents the segmentation position of a paragraph; by outputting the above special characters, the original training corpus is segmented into different granularities of paragraphs, words, and sentences; The fixed composition includes two fixed parts: [sentence to be segmented] and [segmentation result]. [Sentence to be segmented] is at the beginning of the input, and [segmentation result] is at the beginning of the segmentation result.
3. The method according to claim 1 or 2, characterized in that: When constructing the training corpus in step 1, assume that the document has an intermediate segmentation state between d and d final denoted by d i ; for any i < j, d j contains all the segmentation special characters in d i , and the number of segmentation special characters contained in d j is greater than that in d i . Denote the number of segmentation special characters in d i by |d i |, then |d| = 0, |d| < |d i | < |d final |; the constructed training corpus also includes: the document d i to be segmented and the segmentation result d j .
4. The method according to claim 1, characterized in that: In step 3, document d is input into the trained segmentation model, and the segmentation process is as follows: Step 31: construct the input including the document to be segmented d m and segmentation results; input them into the trained segmentation model; Step 32: Use the trained segmentation model to gradually predict the next token until the output is completed; use the regular expression pattern to extract the segmentation result of the output, and record it as d m+1 ; Step 33: Use the segmentation end detection algorithm to determine whether d m+1 As the final segmentation result; if not, the segmentation result d m+1 As the document to be segmented d m , repeat step 31; if yes, go to step 34; Step 34: The segmentation result d m+1 The special characters in are converted into the segmentation result format required by the downstream, and the segmentation result is output.
5. The method according to claim 4, characterized in that The segmentation end detection algorithm includes: Step 331: Calculate the segmentation result d m+1 The segmentation granularity is measured by dividing the number of segmented special characters by the total number of characters; Step 332: Compare d m+1 and d m The number of changes in special characters in the segmentation; if the segmentation granularity > segmentation granularity threshold, and the number of changes in segmented characters < the threshold for changes in the number of segmented characters, it is considered that a stable segmentation result has been achieved and the segmentation is terminated.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is run, the computer-readable storage medium is controlled to execute the iterative multi-granularity document segmentation method according to any one of claims 1 to 5.
7. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the iterative multi-granularity document segmentation method described in any one of claims 1 to 5.