Text processing method and device, medium and product

By constructing a balanced entropy index system and dynamic inter-segment overlap mechanism, the information density and semantic integrity problems in text data paragraph segmentation are solved, and more efficient paragraph segmentation and data analysis effects are achieved.

CN120337907AActive Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510820390.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The prior art fails to consider the information density distribution and semantic integrity when segmenting text data paragraphs, resulting in low validity of segmented paragraphs and cannot meet the needs of subsequent data analysis.

Method used

By constructing an equilibrium entropy index system based on fusion word-level entropy, weighted entropy and sentence-level entropy, the entropy increase before and after paragraph expansion is calculated, the paragraph segment boundary is dynamically judged, and a dynamic inter-segment overlap mechanism is adopted to ensure semantic continuity and information integrity.

Benefits of technology

It improves the accuracy and coverage of paragraph segmentation, avoids the truncation of information-intensive content, meets the paragraph generation density and coverage requirements in different application scenarios, and improves the effectiveness of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337907A_ABST
    Figure CN120337907A_ABST
Patent Text Reader

Abstract

The invention discloses a text processing method and device, a medium and a product, and relates to the technical field of natural language processing, information entropies of text paragraphs are calculated from multiple dimensions by constructing a'balanced entropy 'index system based on fusion word level entropy, weighted entropy and sentence level entropy, and paragraph segmentation boundaries are determined by comparing entropy increases before and after paragraph expansion; the method is advantaged in that information dense content is prevented from being cut off, paragraph integrity is improved, paragraph generation density and coverage rate requirements in different application scenes are satisfied, and high practicality and popularization are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a text processing method, device, medium and product. Background Art

[0002] In the context of artificial intelligence, natural language processing has gradually become more intelligent and large-scale, and with the popularization of the Internet, text data processing has also become a research focus.

[0003] Currently, when the prior art performs paragraph segmentation on text data, it fails to consider the information density distribution and semantic integrity, resulting in low effectiveness of the segmented paragraphs and inability to be used for subsequent data analysis and other operations.

[0004] Therefore, there is an urgent need for a text processing method that can quantify the richness of paragraph information to solve the above technical problems. Summary of the Invention

[0005] This application provides a text processing method, device, medium and product to solve at least the problems in the related art.

[0006] This application provides a text processing method, including: Traverse and select unit text data in the original text paragraph at a unit step length as candidate text data; Add the candidate text data to the current paragraph to obtain an extended paragraph, and the initial state of the current paragraph is empty; Calculate the first balance entropy corresponding to the extended paragraph and the second balance entropy corresponding to the current paragraph according to a first preset formula; Determine the entropy increase of the extended paragraph relative to the current paragraph according to the first balance entropy and the second balance entropy; Judge whether the first segmentation condition is satisfied according to the entropy increase and the text length of the current paragraph; In response to detecting that the first segmentation condition is satisfied, output one or more unit text data in the current paragraph as the first paragraph and set the current paragraph to a pending state; In response to not detecting that the first segmentation condition is satisfied, update the extended paragraph as the current paragraph until the original text paragraph is traversed.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above text processing methods when executing the computer program.

[0008] This application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above text processing methods.

[0009] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above text processing methods.

[0010] In the text processing method disclosed in the present application, by constructing a "balanced entropy" index system based on the fusion of word-level entropy, weighted entropy, and sentence-level entropy, the information entropy of text paragraphs is calculated from multiple dimensions. By comparing the entropy increase before and after paragraph expansion, the paragraph segmentation boundary is determined; avoiding the truncation of information-intensive content, improving paragraph integrity, and meeting the requirements of paragraph generation density and coverage in different application scenarios, with high practicality and promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 It is a flowchart of a text processing method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a text processing method provided by an embodiment of the present application; Figure 3 It is an architecture diagram of a text processing system provided by an embodiment of the present application; Figure 4 It is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] As described in the content disclosed in the background art, in the era of data explosion, how to automatically segment text has become the research focus. Especially in the process of personalized optimization of large language models, an important task in the data preprocessing stage is to segment the original corpus into multiple segments for subsequent generation and synthesis of question-answer pairs. Existing technologies mostly adopt fixed-length segmentation methods. For example, the text is segmented into multiple paragraphs according to the number of characters, the number of words in a dictionary, or the number of Tokens, and then input into the large language model at one time to generate corresponding questions. This solution is simple and effective and is commonly used to construct fine-tuning datasets in standard formats such as Alpaca (a dataset format based on instruction fine-tuning) and Self-Instruct (self-generated instruction framework).

[0016] On this basis, in order to enhance semantic coherence, a fixed overlapping window is introduced for text segmentation, that is, a certain number of Tokens at the end of the previous paragraph will appear at the beginning of the next paragraph. This technical solution alleviates the problem of semantic fragmentation to a certain extent, but essentially still belongs to a "position-driven" segmentation strategy and does not judge the actual information density of the text content. In addition, some studies have also tried to divide the text through structural information such as natural paragraphs, full stops, or line breaks, and even use topic boundary detection algorithms such as TextTiling to identify topic changes; however, these methods usually rely on heuristic rules such as word frequency distribution similarity and boundary value sliding windows between paragraphs and lack the ability to quantify the "information content" of paragraphs.

[0017] Embodiment 1 The embodiment of the present application provides a text processing method. By constructing a "balanced entropy" index system based on the fusion of word-level entropy, TF-IDF weighted entropy, and sentence-level entropy, the present application realizes controlling the segmentation granularity through information entropy, meets the requirements of paragraph generation density and coverage in different application scenarios, and has high practicality and promotion. Figure 1 As shown, the implementation of the method disclosed in the embodiment of the present application to achieve paragraph division of text specifically includes: S1. Traverse and select the unit text data in the original text paragraph at a unit step length as the candidate text data, and add the candidate text data to the current paragraph to obtain an extended paragraph.

[0018] Specifically, the above-mentioned original text paragraphs are usually text data from specific fields (such as the medical, financial, and industrial fields), OCR-extracted documents, and structured-to-natural language tables, etc. The content can be in multiple languages such as Chinese or English. The above unit step length is a custom unit used to set the sentence length of the text data added to the current paragraph each time, that is, the sentence length of a unit of text data. In this application, a unit of text data is sequentially selected as a candidate text data. Preferably, it can be set that one sentence is a unit step length; of course, it can also be set that two sentences are a unit step length, and this application does not make any limitations in this regard.

[0019] It should be noted that in the embodiments of this application, a paragraph (i.e., the current paragraph) is first constructed, and the initial state of this paragraph is empty; after adding the candidate text data to the current paragraph, an expected extended paragraph is formed. That is, if the current paragraph is S and the candidate text data is i, the extended paragraph can be expressed as S i , S i = S + i. Before adding each candidate text data to the current paragraph, it is necessary to calculate the balance entropy of the extended paragraph and the current paragraph to determine whether the current paragraph needs to be output as the first paragraph. If it does not need to be output, the extended paragraph is updated as the new current paragraph, and the candidate text data is continuously added to the updated current paragraph.

[0020] When obtaining the original text paragraph, in some implementation scenarios, such as Figure 2 shown, the embodiments of this application also propose to preprocess the original text paragraph before selecting the unit text data, specifically including: Performing text type verification on the original text paragraph; in response to detecting that the text type verification passes, performing denoising processing on the original text paragraph; in response to detecting that the text type verification fails, triggering a preset conversion operation according to the current text type of the original text paragraph to convert the text type of the original text paragraph. Specifically, the above denoising processing includes but is not limited to processing steps such as unified encoding, removing special characters, clause annotation, paragraph restoration, misspelling correction, and blank line filtering, etc. The above denoising processing can be performed using conventional technical means in this field, and this application will not expand on it here. By verifying the original text paragraph before dividing the text data in the original text paragraph to determine to convert the original text paragraph into a natural language text that is convenient to process, and further denoising the original text paragraph, the parsability and semantic coherence of the text data in the original text paragraph are ensured.

[0021] It should be noted that the text types of the above original text paragraphs usually include natural language text, structured text, and optical character text. The above text type verification is to verify whether the current text type of the original text paragraph is natural language text; if not, the text type verification fails (i.e., it does not pass), and if it is natural language text, the text type verification is successful (i.e., it passes the text type verification).

[0022] After the text verification fails, the embodiments of the present application also propose to trigger a preset conversion operation to ensure that the text type of the original text paragraph is natural language text. Specifically, in response to detecting that the original text paragraph is structured text, the template rules are called to convert the original text paragraph into natural language text and then it is determined that the text type verification passes. Structured text includes tables, lists, database export contents, etc., and needs to be converted into natural language form through specific rules before it can be segmented and entropy value calculated, etc. Among them, the above template rules are preset and stored in a certain medium for calling. For example, by setting table conversion rules, it is predefined how to convert the row and column information of the table into natural language descriptions, including header processing, data row conversion, and merged cell processing, etc.; by setting list conversion rules, it is predefined how to convert list items into natural language descriptions, including bullet point processing and hierarchical structure processing, etc.; and by setting database export content conversion rules, it is predefined how to convert database records into natural language descriptions, including field name processing, record merging, data formatting, etc. It should be noted that the number setting of the above template rules matches the number of types of structured text, covering the rules for converting all different types of structured text into natural language descriptions.

[0023] Furthermore, in response to detecting that the original text paragraph is optical character text, after cleaning and reconstructing the original text paragraph, it is determined that the text type verification passes. Optical character text, that is, OCR text, refers to the text data extracted from images or PDF files through optical character recognition technology. Such text often contains problems such as format disorders and semantic incompleteness, so it needs to be cleaned and reconstructed; the specific cleaning and reconstruction operations are common technical means and will not be elaborated in this application.

[0024] In addition, the present application also proposes to set dictionaries, stop word lists, and corpus statistical parameters in different languages to ensure that the calculation of the balanced entropy index is applicable to different corpus types such as Chinese and English.

[0025] S2. Calculate the first balanced entropy corresponding to the extended paragraph and the second balanced entropy corresponding to the current paragraph according to the first preset formula.

[0026] Specifically, the above first preset formula is H bal ( S )= αHword ( S ) + βH tfidf ( S ) + γH sent ( S ) ; where S represents a paragraph, including an extended paragraph and the current paragraph, H word ( S ) represents the word-level entropy matching the paragraph, H bal ( S ) represents the balanced entropy matching the paragraph, H tfidf ( S ) represents the weighted entropy matching the paragraph, H sent ( S ) represents the sentence-level entropy matching the paragraph, α 、 β and γ are all empirical parameters. It should be noted that the balanced entropy is used to reflect the information integrity and semantic diversity of the paragraph and is the core basis for paragraph boundary judgment. Generally speaking, in scenarios where a wide range of vocabulary needs to be covered, the weight of the word-level entropy is increased α , and the general value range is between 0.2 and 0.5; in scenarios where core information needs to be focused on, such as technical documents and medical reports, the weight of the TF-IDF entropy is increased β , and the general value range is between 0.3 and 0.6; in scenarios where content balance needs to be maintained, such as when multiple key points are described, the weight of the sentence-level entropy is increased γ , and the general value range is between 0.2 and 0.4. Of course, those skilled in the art can adjust the empirical parameters according to the actual scenario, and this application does not make any limitations in this regard.

[0027] It can be understood that the calculation formulas of the first balanced entropy and the second balanced entropy are the same, except that the paragraph data brought in is different.

[0028] It should be noted that the above word-level entropy is based on words and calculates the information entropy by statistically calculating the occurrence probability of each word in a text paragraph, and is used to measure the distribution richness of vocabulary in the paragraph, that is, to measure the diversity of vocabulary. The higher the entropy value of the word-level entropy, the stronger the discreteness and diversity of the word usage in the text paragraph.

[0029] Specifically, the method for obtaining the word-level entropy includes: assuming a paragraph is S, the set contained in it is marked as W, and the frequency of occurrence of a certain word is f(w), and the probability is calculated according to the third preset formula p (w ), the third preset formula is: ; Further calculate the word-level entropy according to the fourth preset formula H word ( S ), the fourth preset formula is: .

[0030] The above weighted entropy, that is, using the TF-IDF value as the importance weight of the word, calculates the weighted value to measure the distribution state of "high information value words" in a text paragraph, emphasizes semantic key information more, strengthens the influence of semantic core words in the paragraph, and helps to distinguish between texts with "rich surface content" and "rich semantic information". Among them, TF-IDF (Term Frequency-Inverse Document Frequency) is a statistical method used to measure the importance of a word to a document and is commonly used for keyword extraction. TF represents term frequency, and IDF represents inverse document frequency. The combined use can reflect the representativeness of a certain word in the current document.

[0031] Specifically, first calculate the weight of each word through the fifth preset formula , where the fifth preset formula is expressed as: ; Further calculate the weighted entropy of the paragraph according to the sixth preset formula, where the sixth preset formula is: .

[0032] The above sentence-level entropy takes the sentence as the information unit, calculates the information entropy after normalizing the importance score of each sentence to measure the balance of sentence content and the concentration of information distribution in a text paragraph, and is used to depict the structural complexity at the sentence level. The higher the entropy value, the more balanced the structure and content between sentences, which is beneficial for downstream generation to cover multiple key points.

[0033] Specifically, split a paragraph into a sentence set { S 1 , S 2 ,... ,S i}}, the score h( S i ) of each sentence comes from the sentence length or the number of keywords. Specifically, calculate the entropy value p( S i ) of each sentence through the seventh preset formula, where the seventh preset formula is expressed as: ; and calculate the sentence-level entropy of a paragraph text through an eighth preset formula H sent ( S ), where the eighth preset formula is: .

[0034] This application measures the information entropy of paragraph texts from different dimensions through the "balanced entropy" index algorithm based on the fusion of word-level entropy, TF-IDF weighted entropy, and sentence-level entropy, comprehensively measuring the information density of paragraphs; it provides a basis for implementing the paragraph segmentation strategy based on "information density driving", greatly improving the accuracy of subsequent paragraph segmentation.

[0035] S3. Determine the entropy increase of the extended paragraph relative to the current paragraph according to the first balanced entropy and the second balanced entropy. Specifically, the entropy increase can be expressed as the difference between the first balanced entropy and the second balanced entropy.

[0036] S4. Determine whether the first segmentation condition is satisfied according to the entropy increase and the text length of the current paragraph.

[0037] Specifically, if the entropy increase is less than the information gain saturation value and the text length of the current paragraph is greater than or equal to the second text threshold, it is determined that the first segmentation condition is satisfied; or, if the entropy increase is greater than the information overload preset value and the text length of the current paragraph is greater than or equal to the second text threshold, it is determined that the first segmentation condition is satisfied. The second text threshold is the minimum value for length protection, and preferably can be set to 512, which is set by those skilled in the art according to the actual scenario, and this application does not make a limitation on this.

[0038] Among them, the information gain saturation value is a preset value greater than zero, usually taking a value in the order of to . The information overload preset value H max is the maximum information entropy per paragraph preset. Specifically, the setting of the information overload preset value H max needs to balance two key objectives: preventing information overload and ensuring semantic integrity. In a specific field H max is determined through the following steps: In the representative texts of the target application field, randomly select at least 1000 text paragraphs, where the length of each paragraph is between 1 and N sentences (N usually takes 20); calculate the balanced entropy of each paragraph; sort all the obtained balanced entropy values in ascending order, and take the 95th percentile as H maxThe value of. The significance of taking the 95th percentile is that it can cover the vast majority of normal paragraphs and resist the interference of outliers. This percentile can also be adjusted according to the task requirements.

[0039] S5. In response to detecting that the first segmentation condition is met, output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to the pending state.

[0040] Specifically, the above-mentioned output of one or more unit text data within the current paragraph as the first paragraph and setting the current paragraph to the pending state includes: In response to detecting the generation of the first paragraph, calculate the number of overlapping units to be matched with the first paragraph; starting from the end of the first paragraph, select one or more unit text data that match the number of overlapping units to be matched in reverse order to generate overlapping text data; add the overlapping text data to the current paragraph to set the current paragraph to the pending state. That is, after each available paragraph (i.e., the first paragraph) is output, determine the text data that needs to be overlapped into the next paragraph as the context prefix of the next paragraph, thereby ensuring semantic continuity. Through the dynamic inter-segment overlap mechanism of this application, the context continuity is ensured, and the quality and semantic consistency among multiple first paragraphs matching the original text paragraphs are significantly improved, solving the problem of cross-segment information loss after segmentation.

[0041] Furthermore, the calculation process of the number of overlapping units to be matched with the first paragraph includes: Select the specified number of unit text data from the end of the first paragraph; according to the specified number of unit text data, the balance entropy matched with the first paragraph, and the second preset formula, determine the average entropy value at the end of the first paragraph; according to the average entropy value at the end of the paragraph, the information overload preset value, the maximum allowable number of overlapping units, and the third preset formula, determine the number of overlapping units to be matched with the first paragraph.

[0042] Among them, the value of the specified number of units m is generally set to 3 - 5 sentences at the end of the paragraph, with the setting standard of covering a complete small semantic unit, such as including the statement of the view + evidence + conclusion; in the implementation scenario where it is necessary to condense the technical key points, the specified number of units m is preferably set to 3 sentences; in the implementation scenario where more logical continuity is required, it is preferably set to 5 sentences. Of course, the specific setting standard can be set by those skilled in the art according to the actual scenario, and this application does not make any limitations in this regard.

[0043] The above-mentioned second preset formula is: ; Among them, H tail represents the average entropy value at the end of the paragraph, m represents the specified number of units, S 1 represents the first paragraph, Represents the balanced entropy for the first paragraph match.

[0044] The above third preset formula is: k = ceil ( H tail / H max × K max ); where, k represents the number of units to be overlapped, H max represents the preset value of information overload, K max represents the maximum allowable number of overlapping units, H tail represents the average entropy value at the end of the paragraph, ceil is the rounding function.

[0045] In addition, to avoid double counting of the balanced entropy, a penalty coefficient for the entropy weight increase is set for the overlapping part in the next paragraph, preferably set to 0.5, so as to improve the model inference efficiency. Of course, it can also be set by those skilled in the art according to the actual situation, and this application does not make any limitations in this regard. Specifically, an overlapping area identifier is set at the head of the newly created paragraph, which can be used as a separate paragraph for calculating the balanced entropy value and applied when calculating the balanced entropy of the overall paragraph: effective balanced entropy value = (penalty coefficient × balanced entropy of the overlapping part) + balanced entropy of the non-overlapping part.

[0046] S6. In response to not detecting that the first segmentation condition is satisfied, update the extended paragraph to the current paragraph until all the original text paragraphs are traversed.

[0047] After not detecting that the first segmentation condition is satisfied, in order to avoid the increased difficulty of subsequent analysis caused by too long paragraphs, this application also proposes to perform length protection on the current paragraph. Specifically: Compare the text length of one or more unit text data in the current paragraph with the first text threshold; in response to detecting that the text length is greater than or equal to the first text threshold, output one or more unit text data in the current paragraph as the first paragraph and set the current paragraph to a pending state. Wherein the first text threshold is the maximum value for protecting the length of a paragraph set in advance, preferably can be set to 4096, and is set by those skilled in the art according to the actual situation, and this application does not make any limitations in this regard.

[0048] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0049] Embodiment 2 The embodiments of the present application also propose to segment the original text paragraph in combination with distribution divergence. Specifically, after the step of not detecting the satisfaction of the first segmentation condition in step S6 disclosed in the foregoing embodiments is executed, the present application proposes: Calculate the distribution divergence between the current paragraph and the extended paragraph; specifically, it can be obtained by calculating the Jensen-Shannon divergence (JS divergence), and substitute the current paragraph and the extended paragraph as two probability distributions for calculation. The specific calculation method is a conventional means, and the present application will not expand it here. According to the distribution divergence and the preset divergence value, determine whether the second segmentation condition is satisfied; the preset divergence value is determined by factors such as text data characteristics (such as lexical richness and paragraph length) and the field involved in the text. In a loose scenario, it is preferably set to any value in the range of 0.2 to 0.3, and in a strict scenario, it is preferably set to any value in the range of 0.1 to 0.15. The specific value is determined by those skilled in the art according to the actual situation.

[0050] In response to detecting the satisfaction of the second segmentation condition, output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to a pending state; the setting of the pending state is consistent with the content disclosed in the foregoing embodiments, and the present application will not expand it here. In response to not detecting the satisfaction of the second segmentation condition, update the extended paragraph to the current paragraph until the original text paragraph is traversed; of course, in order to avoid increased difficulty in subsequent analysis due to overly long paragraphs, after not detecting the satisfaction of the second segmentation condition, the present application also proposes to perform length protection on the current paragraph, and the specific steps are consistent with the content disclosed in the foregoing embodiments, and the present application will not expand it here.

[0051] The present application combines the "balanced entropy" index after fusion and distribution divergence to achieve dynamic paragraph division, which can effectively perceive semantic changes and further improve the accuracy of paragraph segmentation.

[0052] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0053] Embodiment III The embodiments of the present application also propose to construct a training data set for a large language model by using the first paragraph generated after segmenting the original paragraph by the above method: It can be understood that in the current practice of fine-tuning large models, the automatic generation process of synthetic datasets is becoming the mainstream method. Especially in scenarios where data privacy is restricted and data in specific domains is scarce, this process usually takes private text data as the original text paragraphs, inputs them, generates question-and-answer pairs through large models, and is used for subsequent supervised fine-tuning training. Especially in the era of the rapid development of large language models, in order to ensure the quality of the model generation results, it is particularly important how to segment the text input into the model. It can be understood that different segmentation methods will affect the information organization form of the paragraphs, and further affect whether the model can extract key content for questions. Therefore, the design of the segmentation strategy is not only related to the generation efficiency, but also directly affects the quality and coverage of the generated answers. Most of the current mainstream paragraph segmentation methods adopt a fixed-length segmentation method, that is, the text is roughly truncated according to a fixed number of words (such as 512 words) or a fixed number of Tokens. There are also some that adopt a fixed sliding window with overlap strategy to avoid semantic fragmentation. These methods are simple to implement and have engineering feasibility. However, there are still the following disadvantages: Failure to consider the information density distribution: Fixed-length segmentation treats all texts equally, unable to recognize the reality that the information in some paragraphs is highly concentrated while that in other paragraphs is sparse, easily resulting in important information being truncated or wasting generation resources on invalid paragraphs. Insufficient guarantee of semantic integrity: Whether it is fixed-length truncation or a strategy based on natural paragraph division, it may cause key sentences to be split or topic sentences to be omitted, making the generation model lack the necessary context support and reducing the quality of the generated questions and answers. Weak structural adaptation ability: Fixed-rule methods have poor adaptability to structured texts such as formatted tables, nested long sentences, and list descriptions, and are prone to incorrect sentence cutting and block breaking, affecting the effects of downstream tasks. Coarse context control: The fixed overlap method fails to dynamically judge the necessity of overlapping paragraphs, often resulting in the repeated input of too much redundant context and increasing the generation burden. Unable to uniformly evaluate the effectiveness of paragraphs: Existing methods lack a unified evaluation index that can quantify the richness of paragraph information and cannot judge whether a segment is worth generating questions and answers before segmentation. In short, the existing solutions have different degrees of defects in terms of the intelligence, adaptability, efficiency, and information preservation of paragraph division. The embodiments of the present application provide a text processing method to segment the original paragraphs, systematically improve the coverage, accuracy, and cost-effectiveness of synthetic data generation, further improve the quality of the generated training dataset, and thus improve the effect of large language models.

[0054] The embodiments of the present application are extended based on the methods disclosed in Embodiment 1 and Embodiment 2 above. After the steps of traversing the original text paragraphs in Embodiment 1 and Embodiment 2 are executed, this embodiment further includes: Obtain and standardize one or more first paragraphs that match the original text paragraph to obtain one or more second paragraphs. Specifically, output the first paragraphs in JSON format (JavaScript Object Notation, a lightweight data exchange format) or TSV format (Tab-Separated Values, a plain text format using tab characters as separators).

[0055] Add a preset prompt word to the second paragraph to generate a third paragraph; for example, "Please generate 5 valuable questions and answers based on the following content". Input the third paragraph into the large model to obtain Q&A text data that matches the third paragraph. Collect the Q&A text data that matches the original text paragraph to construct a training dataset for fine-tuning the large model. It can be understood that the finally generated training dataset can be used for purposes such as supervised fine-tuning, instruction tuning, response quality evaluation, and data quality filtering, and has advantages such as traceability, controllability, and clear structure. The division of the first paragraph is consistent with the text processing methods disclosed in Embodiment 1 and Embodiment 2, and will not be elaborated again in this application.

[0056] The embodiment of the present application provides a method for segmenting paragraphs based on "balanced entropy" and then synthesizing a dataset. Through a multi-dimensional information entropy evaluation system, a dynamic segmentation boundary detection and overlap control mechanism, it avoids truncation of information-dense content, improves paragraph integrity, and ensures that the generated questions can cover all valid information points. In addition, it also reduces the risk of generating invalid Q&A for information-sparse paragraphs, reduces the number of large model calls, and improves the overall Q&A pair generation efficiency. Based on a more intelligent, accurate, and efficient data segmentation strategy, the effectiveness of the finally generated training dataset is greatly improved, thereby improving the training effect of the large model.

[0057] The embodiment of the present application provides a text processing system, as Figure 3 shown, including: A paragraph preparation module 310, configured to traverse and select unit text data in the original text paragraph at a unit step length as candidate text data; The above paragraph preparation module 310 is further configured to add the candidate text data to the current paragraph to obtain an extended paragraph, and the initial state of the current paragraph is empty; An information entropy measurement module 320, configured to calculate a first balanced entropy corresponding to the extended paragraph and a second balanced entropy corresponding to the current paragraph according to a first preset formula; The information entropy measurement module 320 is further configured to determine the entropy increase of the extended paragraph relative to the current paragraph according to the first balanced entropy and the second balanced entropy; A segmentation judgment module 330, configured to judge whether the first segmentation condition is satisfied according to the entropy increase and the text length of the current paragraph; The paragraph processing module 340 is configured to output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to a pending state in response to detecting that the first segmentation condition is met; The paragraph processing module 340 is further configured to update the extended paragraph to the current paragraph until the original text paragraph is traversed in response to not detecting that the first segmentation condition is met.

[0058] An embodiment of the present application further provides an electronic device, including: one or more processors; and a memory associated with the one or more processors, where the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the following operations are performed: Traverse and select unit text data in the original text paragraph as candidate text data according to a unit step size; Add the candidate text data to the current paragraph to obtain an extended paragraph, and the initial state of the current paragraph is empty; Calculate the first balance entropy corresponding to the extended paragraph and the second balance entropy corresponding to the current paragraph according to a first preset formula; Determine the entropy increase of the extended paragraph relative to the current paragraph according to the first balance entropy and the second balance entropy; Judge whether the first segmentation condition is met according to the entropy increase and the text length of the current paragraph; In response to detecting that the first segmentation condition is met, output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to a pending state; In response to not detecting that the first segmentation condition is met, update the extended paragraph to the current paragraph until the original text paragraph is traversed.

[0059] Among them, the first preset formula is: H bal ( S ) = αH word ( S ) + βH tfidf ( S ) + γH sent ( S ); Among them, S represents a paragraph, including the extended paragraph and the current paragraph, H bal ( S ) represents the balance entropy matching the paragraph, H word ( S ) represents the word-level entropy matching the paragraph, H tfidf ( Srepresents the weighted entropy matching the paragraph, H sent ( S ) represents the sentence-level entropy matching the paragraph, α 、 β and γ are all empirical parameters.

[0060] In some implementation scenarios, when the program instructions are read and executed by one or more processors, after detecting that the first segmentation condition is not satisfied, the following operations are further performed: Calculate the distribution divergence between the current paragraph and the extended paragraph; Judge whether the second segmentation condition is satisfied according to the distribution divergence and the preset divergence value; In response to detecting that the second segmentation condition is satisfied, output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to a pending state; In response to detecting that the second segmentation condition is not satisfied, update the extended paragraph to the current paragraph until all the original text paragraphs are traversed.

[0061] In some implementation scenarios, when the program instructions are read and executed by one or more processors, after detecting that the first segmentation condition is not satisfied and after detecting that the second segmentation condition is not satisfied, the following operations are further performed: Compare the text length of one or more unit text data within the current paragraph with the first text threshold; In response to detecting that the text length is greater than or equal to the first text threshold, output one or more unit text data within the current paragraph as the first paragraph and set the current paragraph to a pending state.

[0062] In some implementation scenarios, when the program instructions are read and executed by one or more processors, after traversing all the original text paragraphs, the following operations are further performed: Obtain and perform normalization processing on one or more first paragraphs matching the original text paragraphs to obtain one or more second paragraphs; Add a preset prompt word to the second paragraph to generate a third paragraph; Input the third paragraph into the large model to obtain the Q&A text data matching the third paragraph; Collect the Q&A text data matching the original text paragraphs to construct a training dataset for fine-tuning the large model.

[0063] In some implementation scenarios, when the program instructions are read and executed by one or more processors, the following operations are further performed: If the entropy increase is less than the information gain saturation value and the text length of the current paragraph is greater than or equal to the second text threshold, it is determined that the first segmentation condition is satisfied; Or, if the entropy increase is greater than the information overload preset value and the text length of the current paragraph is greater than or equal to the second text threshold, it is determined that the first segmentation condition is satisfied.

[0064] In some implementation scenarios, when the program instructions are read and executed by one or more processors, the following operations are also performed: In response to detecting the generation of the first paragraph, calculate the number of overlapping units to be matched with the first paragraph; Starting from the end of the first paragraph, select one or more unit text data that match the number of overlapping units to be generated in reverse order to generate overlapping text data; Add the overlapping text data to the current paragraph to set the current paragraph to a pending state.

[0065] In some implementation scenarios, when the program instructions are read and executed by one or more processors, the following operations are also performed: Select unit text data of a specified number from the end of the first paragraph; According to the unit text data of the specified number, the balanced entropy matched with the first paragraph, and the second preset formula, determine the average entropy value at the end of the first paragraph; According to the average entropy value at the end of the paragraph, the information overload preset value, the maximum allowable number of overlapping units, and the third preset formula, determine the number of overlapping units to be matched with the first paragraph.

[0066] The second preset formula is: ; Wherein, H tail represents the average entropy value at the end of the paragraph, m represents the specified number, S 1 represents the first paragraph, represents the balanced entropy matched with the first paragraph.

[0067] The third preset formula is: k = ceil ( H tail / H max × K max ) ; Wherein, k represents the number of overlapping units to be, H max represents the information overload preset value, K max represents the maximum allowable number of overlapping units, H tail represents the average entropy value at the end of the paragraph, ceil is the rounding function.

[0068] In some implementation scenarios, before the program instructions are read and executed by one or more processors, and traverse and select the unit text data in the original text paragraph as candidate text data at a unit step length, the following operations are also performed: Perform a text type check on the original text paragraph; In response to detecting that the text type check passes, perform denoising processing on the original text paragraph; In response to detecting that the text type check fails, trigger a preset conversion operation according to the current text type of the original text paragraph to convert the text type of the original text paragraph.

[0069] In some implementation scenarios, before the program instructions are read and executed by one or more processors, the following operations are also performed: In response to detecting that the original text paragraph is structured text, call a template rule to convert the original text paragraph into natural language text and then determine that the text type check passes; In response to detecting that the original text paragraph is optical character text, perform cleaning and reconstruction on the original text paragraph and then determine that the text type check passes.

[0070] Among them, Figure 4 An exemplary architecture of the electronic device is shown, which may specifically include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The above-mentioned processor 410, video display adapter 411, disk drive 412, input / output interface 413, network interface 414, and the memory 420 can be communicatively connected through a bus 430.

[0071] Among them, the processor 410 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application.

[0072] The memory 420 can be implemented in the form of ROM (ReadOnlyMemory), RAM (RandomAccessMemory), static storage devices, dynamic storage devices, etc. The memory 420 can store the operating system 421 for controlling the execution of the electronic device 400, and the basic input / output system (BIOS) 422 for controlling the low-level operations of the electronic device 400. Additionally, it can also store a web browser 423, a data storage management system 424, an icon font processing system 425, and so on. The above-mentioned icon font processing system 425 can be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 420 and called and executed by the processor 410.

[0073] The input / output interface 413 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices can include a display, a speaker, a vibrator, an indicator light, etc.

[0074] The network interface 414 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0075] The bus 430 includes a path for transmitting information between various components of the device (such as the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420).

[0076] In addition, the electronic device 400 can also obtain information on specific collection conditions from the virtual resource object collection condition information database for conditional judgment, and so on.

[0077] It should be noted that although the above device only shows the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, the memory 420, the bus 430, etc., in the specific implementation process, the device may also include other components necessary for normal execution. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the present application and does not necessarily include all the components shown in the figure.

[0078] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above-described embodiments of the text processing method when running.

[0079] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0080] Embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the text processing method are implemented.

[0081] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the text processing method are implemented.

[0082] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0083] The above provides a detailed introduction to a text processing method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A text processing method, characterized in that, The method includes: Traversing and selecting unit text data in the original text paragraph at a unit step length as candidate text data; Adding the candidate text data to the current paragraph to obtain an extended paragraph, where the initial state of the current paragraph is empty; Calculating a first balanced entropy corresponding to the extended paragraph and a second balanced entropy corresponding to the current paragraph according to a first preset formula; Determining the entropy increase of the extended paragraph relative to the current paragraph according to the first balanced entropy and the second balanced entropy; Judging whether a first segmentation condition is satisfied according to the entropy increase and the text length of the current paragraph; In response to detecting that the first segmentation condition is satisfied, outputting one or more unit text data in the current paragraph as the first paragraph and setting the current paragraph to a pending state; In response to not detecting that the first segmentation condition is satisfied, updating the extended paragraph as the current paragraph until the original text paragraph is traversed.

2. The method according to claim 1, characterized in that, After the step of in response to not detecting that the first segmentation condition is satisfied, the method further includes: Calculating the distribution divergence between the current paragraph and the extended paragraph; Judging whether a second segmentation condition is satisfied according to the distribution divergence and a preset divergence value; In response to detecting that the second segmentation condition is satisfied, outputting one or more unit text data in the current paragraph as the first paragraph and setting the current paragraph to a pending state; In response to not detecting that the second segmentation condition is satisfied, updating the extended paragraph as the current paragraph until the original text paragraph is traversed.

3. The method according to claim 2, wherein After the steps of in response to not detecting that the first segmentation condition is satisfied and in response to not detecting that the second segmentation condition is satisfied, the method further includes: Comparing the text length of one or more unit text data in the current paragraph with a first text threshold; In response to detecting that the text length is greater than or equal to the first text threshold, outputting one or more unit text data in the current paragraph as the first paragraph and setting the current paragraph to a pending state.

4. The method according to claim 1, wherein After traversing the original text paragraph, the method further includes: Obtaining and performing normalization processing on one or more of the first paragraphs matching the original text paragraph to obtain one or more second paragraphs; Adding a preset prompt word to the second paragraph to generate a third paragraph; Inputting the third paragraph into a large model to obtain question-and-answer text data matching the third paragraph; Collecting the question-and-answer text data matching the original text paragraph to construct a training data set for training and fine-tuning the large model.

5. The method according to claim 1, wherein The first preset formula is: H bal ( S )= αH word ( S )+ βH tfidf ( S )+ γH sent ( S ); in, S Represents a paragraph, including the extended paragraph and the current paragraph. H bal ( S ) represents the equilibrium entropy matching the paragraph, H word ( S ) represents the word-level entropy matching the paragraph, H tfidf ( S ) represents the weighted entropy of matching with the paragraph, H sent ( S ) represents the sentence-level entropy matching the paragraph, α , β and γ These are all empirical parameters.

6. The method according to claim 4, wherein The step of judging whether a first segmentation condition is satisfied according to the entropy increase and the text length of the current paragraph includes: If the entropy increase is less than the information gain saturation value and the text length of the current paragraph is greater than or equal to a second text threshold, it is determined that the first segmentation condition is satisfied; Or, if the entropy increase is greater than the information overload preset value and the text length of the current paragraph is greater than or equal to a second text threshold, it is determined that the first segmentation condition is satisfied.

7. The method according to claim 6, characterized in that, Outputting one or more unit text data within the current paragraph as the first paragraph and setting the current paragraph to a pending state includes: In response to detecting the generation of the first paragraph, calculating the number of overlapping units to be matched with the first paragraph; Starting from the end of the first paragraph, selecting one or more unit text data that match the number of overlapping units to be matched in reverse order to generate overlapping text data; Adding the overlapping text data to the current paragraph to set the current paragraph to a pending state.

8. The method according to claim 7, wherein The calculating the number of overlapping units to be matched with the first paragraph includes: Selecting unit text data of a specified number of units from the end of the first paragraph; Determining the average entropy value at the end of the first paragraph according to the unit text data of the specified number of units, the balance entropy matched with the first paragraph, and a second preset formula; Determining the number of overlapping units to be matched with the first paragraph according to the average entropy value at the end, the preset value of information overload, the maximum allowable number of overlapping units, and a third preset formula.

9. The method according to claim 8, wherein The second preset formula is: ; Among them, H tail represents the average entropy value at the end of the segment, m represents the specified number of units, S 1 represents the first paragraph, represents the balanced entropy matching the first paragraph.

10. The method according to claim 8, wherein The third preset formula is as follows: k = ceil ( H tail / H max × K max ); Among them, k represents the number of units to be overlapped, H max represents the preset value of information overload, K max represents the maximum allowable number of overlapping units, H tail represents the average entropy value at the end of the segment, ceil is the rounding function.

11. The method according to claim 1, wherein Before traversing and selecting unit text data in the original text paragraph as candidate text data according to a unit step size, the method further includes preprocessing the original text paragraph: Performing text type verification on the original text paragraph; In response to detecting that the text type verification passes, performing denoising processing on the original text paragraph; In response to detecting that the text type verification fails, triggering a preset conversion operation according to the current text type of the original text paragraph to convert the text type of the original text paragraph.

12. The method according to claim 11, wherein The triggering a preset conversion operation according to the current text type of the original text paragraph to convert the text type of the original text paragraph, the method includes: In response to detecting that the original text paragraph is a structured text, calling a template rule to convert the original text paragraph into a natural language text and then determining that the text type verification passes; In response to detecting that the original text paragraph is an optical character text, performing cleaning and reconstruction on the original text paragraph and then determining that the text type verification passes.

13. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the text processing method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the text processing method according to any one of claims 1 to 12 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the text processing method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • Micro-blog text classification method and system based on high-quality topic extension

    CN109344252A

  • Method and device for text paragraph division

    CN110674635A

  • Long text similarity comparison method based on paragraph division

    CN117688138A

  • Normalized log generation method based on entropy increase principle

    CN117709301A

  • Nested named entity identification method based on early leaving mechanism

    CN118261157A

Cited By

  • Cross-platform document management method and system

    CN120994623A