Text processing method and device, computer device, storage medium and computer program product
Patent Information
- Application Number
- CN202410713613.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-04
AI Technical Summary
使用关键词密度筛选的效果依赖于分词结果,如果某个词组被切分成了两个部分,那么在做关键词提取时无法将它们合并
[0059]上述文本处理方法、装置、计算机设备、存储介质和计算机程序产品,获取待处理文本集合后,确定质量检测指标,并计算每一第一文本对应的质量检测指标的指标值;获取文本过滤目标,并基于文本过滤目标确定各所述质量检测指标对应的初始过滤阈值;获取质量检测指标对应的初始过滤阈值处的第二文本的文本质量;基于质量检测指标与文本质量的相关关系以及所述第二文本的文本质量,调整初始过滤阈值;基于调整后的所述初始过滤阈值以及各所述第一文本对应的所述质量检测指标的指标值,对所述第一文本进行过滤得到第三文本,这样包括多个质量检测指标,且基于文本过滤目标来动态确定初始过滤阈值,从而提高过滤的质量,进而提高文本处理质量。
Smart Images

Figure CN118673130B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a text processing method, apparatus, computer device, storage medium, and computer program product. Background Technology
[0002] In fields such as natural language processing, data analysis, and machine learning, high-quality text data is one of the key elements that can significantly affect model training and system performance.
[0003] Traditional techniques for text processing typically employ methods such as keyword density filtering to sift out low-quality and advertising data from sampled data. The effectiveness of keyword density filtering depends on the word segmentation results; if a phrase is split into two parts, they cannot be merged during keyword extraction.
[0004] Therefore, there is an urgent need for a text processing method that can improve the quality of text processing. Summary of the Invention
[0005] Therefore, it is necessary to provide a text processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the quality of text processing in response to the above-mentioned technical problems.
[0006] Firstly, this application provides a text processing method, the method comprising:
[0007] Obtain a set of texts to be processed, the set of texts to be processed including a number of first texts;
[0008] Obtain quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each of the first texts;
[0009] Obtain the text filtering target, and determine the initial filtering threshold corresponding to each of the quality detection indicators based on the text filtering target;
[0010] Obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index;
[0011] Based on the correlation between the quality detection index and text quality, and the text quality of the second text, the initial filtering threshold is adjusted, including: when the correlation between the quality detection index and the text quality is positive and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is decreased; when the correlation between the quality detection index and the text quality is negative and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is increased.
[0012] Based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the first texts are filtered to obtain the third text.
[0013] In one embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes:
[0014] The third text is deduplicated to obtain the fourth text;
[0015] The deduplication process includes at least one of the following: performing deduplication based on the text summary of each of the third texts; or obtaining an information array corresponding to each of the third texts based on the index value of each of the quality detection indicators of each of the third texts, and performing deduplication on the third texts based on the information array.
[0016] In one embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes:
[0017] When the number of characters in the third text is greater than the target character processing capacity of the model, the text structure type corresponding to the third text is obtained;
[0018] Determine the text segmentation logic corresponding to the text structure type, and perform text segmentation on the third text based on the text segmentation logic.
[0019] In one embodiment, before obtaining the quality inspection indicators, the method further includes:
[0020] The first text is subjected to template recognition, and deduplication is performed based on the recognized template;
[0021] The first text after deduplication is subjected to text structure type identification, wherein the text structure type includes those with chapter structure and those without chapter structure.
[0022] In one embodiment, determining the text segmentation logic corresponding to the text structure type and segmenting the third text based on the text segmentation logic includes:
[0023] When the text structure type is a chapter structure, each layer of text is determined according to the order of the chapter structure;
[0024] If the number of characters in the current layer of text is greater than the target character processing capacity of the model, continue to segment the text in the current layer according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0025] In one embodiment, the method further includes:
[0026] When the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, the third text or the minimum level text is taken as the text to be segmented, the first order of each delimiter is determined, and the current delimiter is obtained based on the first order.
[0027] The text to be segmented is segmented based on the current delimiter to obtain the current segmented text;
[0028] When the number of characters in the current segmented text is greater than the target character processing capacity of the model, the next delimiter is obtained based on the first order of each delimiter as the current delimiter, and the current segmented text continues to be segmented until the number of characters in the current segmented text after each segmentation is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0029] In one embodiment, the method for segmenting the text to be segmented includes:
[0030] The number of slices is determined based on the number of characters in the text to be segmented and the target character processing capacity of the model.
[0031] The length of the segmented text is obtained based on the number of slices and the number of characters in the text to be segmented;
[0032] Based on the length of the segmented text, the text to be segmented is segmented to obtain the current segmented text.
[0033] In one embodiment, prior to obtaining the quality inspection indicators, the method further includes:
[0034] The first text is formatted, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked.
[0035] The text segmentation of the third text based on the text segmentation logic includes:
[0036] The table or formula is segmented based on the text segmentation logic, and the start and end points of each segmented table and formula are marked. The table header and continuation mark are marked on the next slice at the cut position of the table, and the continuation mark is marked on the next slice at the cut position of the formula.
[0037] In one embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes:
[0038] Determine the text type of the third text;
[0039] The third text is cleaned based on the text type.
[0040] Secondly, this application also provides a text processing apparatus, the apparatus comprising:
[0041] The module for obtaining a set of texts to be processed is used to obtain a set of texts to be processed, which includes several first texts.
[0042] The indicator value calculation module is used to obtain quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each of the first texts.
[0043] An initial filtering threshold determination module is used to obtain the text filtering target and determine the initial filtering threshold corresponding to each of the quality detection indicators based on the text filtering target.
[0044] The text quality determination module is used to obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index;
[0045] The threshold adjustment module is used to adjust the initial filtering threshold based on the correlation between the quality detection index and the text quality, as well as the text quality of the second text.
[0046] The filtering module is used to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index value of the quality detection index corresponding to each of the first texts.
[0047] In one embodiment, the apparatus further includes: a text deduplication module, used to perform deduplication processing on the third text to obtain a fourth text; wherein the deduplication processing includes at least one of the following: performing deduplication processing based on the text summary of each of the third texts; or obtaining an information array corresponding to each of the third texts based on the index value of each of the quality detection indicators of each of the third texts, and performing deduplication processing on the third texts based on the information array.
[0048] In one embodiment, the apparatus further includes: a text segmentation module, configured to, when the number of characters in the third text is greater than the target character processing capacity of the model, obtain the text structure type corresponding to the third text; determine the text segmentation logic corresponding to the text structure type; and perform text segmentation on the third text based on the text segmentation logic.
[0049] In one embodiment, the device further includes a structure classification module for performing template recognition on the first text and deduplication processing based on the recognized template; and for performing text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0050] In one embodiment, the text segmentation module is further configured to determine each layer of text according to the order of the chapter structure when the text structure type is a chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continue to segment the text in the current layer according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, and then the segmentation ends.
[0051] In one embodiment, the text segmentation module is further configured to: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, take the third text or the minimum level text as the text to be segmented, determine the first order of each delimiter, and obtain the current delimiter based on the first order; segment the text to be segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, obtain the next delimiter based on the first order of each delimiter as the current delimiter, and continue to segment the current segmented text until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0052] In one embodiment, the text segmentation module is further configured to determine the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtain the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segment the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0053] In one embodiment, the apparatus further includes a format conversion module for converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked.
[0054] The text segmentation module is also used to segment the table or the formula based on the text segmentation logic, and mark the start and end points of each segmented table and formula, and mark the table header and continuation mark at the cut position of the table, and mark the continuation mark at the cut position of the formula.
[0055] In one embodiment, the apparatus further includes a cleaning module for determining the text type of the third text and cleaning the third text based on the text type.
[0056] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method in any of the above embodiments.
[0057] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the methods in any of the above embodiments.
[0058] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods in any of the above embodiments.
[0059] The aforementioned text processing method, apparatus, computer equipment, storage medium, and computer program product, after acquiring a set of texts to be processed, determine quality detection indicators and calculate the indicator value of each first text corresponding to the quality detection indicator; acquire a text filtering target and determine an initial filtering threshold corresponding to each of the quality detection indicators based on the text filtering target; acquire the text quality of the second text at the initial filtering threshold corresponding to the quality detection indicator; adjust the initial filtering threshold based on the correlation between the quality detection indicator and the text quality and the text quality of the second text; and filter the first text to obtain a third text based on the adjusted initial filtering threshold and the indicator value of each of the first texts. This includes multiple quality detection indicators and dynamically determines the initial filtering threshold based on the text filtering target, thereby improving the filtering quality and ultimately improving the text processing quality. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is a flowchart illustrating a text processing method in one embodiment;
[0062] Figure 2 This is a schematic diagram of the text structure in one embodiment;
[0063] Figure 3 This is a schematic diagram of the text structure in another embodiment;
[0064] Figure 4 This is a flowchart of a text segmentation step with a chapter structure in one embodiment;
[0065] Figure 5 Here is a flowchart of the general segmentation steps in one embodiment;
[0066] Figure 6 A flowchart of a text processing method in another embodiment;
[0067] Figure 7 This is a structural block diagram of a text processing device in one embodiment;
[0068] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0070] In one embodiment, such as Figure 1 As shown, a text processing method is provided. This embodiment illustrates the method applied to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0071] S102: Obtain the set of texts to be processed, which includes several first texts.
[0072] The text set to be processed refers to the set of texts that need to be preprocessed. For ease of description, the texts in this text set to be processed are called the first text.
[0073] S104: Obtain the quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each first text.
[0074] Quality assessment metrics include at least one of the following: lexical richness, sentence structure complexity, paragraph structure complexity, text information entropy, text degradation, text redundancy, text readability, and stop word percentage. Lexical richness measures the diversity of words in the text and can be calculated as the proportion of deduplicated words to the total number of words, i.e., len(set(all_words)) / len(all_words). Generally, higher lexical richness indicates higher text quality. The total number of words (all_words) is the word segmentation result obtained after precise pattern segmentation of a given text using the Jieba Chinese word segmentation component.
[0075] Sentence structure complexity: This is an indicator of sentence length in a text, calculated as the proportion of complex sentences to the total number of sentences, i.e., len(complex_sentences) / len(all_sentences). Generally, higher sentence structure complexity indicates higher text quality.
[0076] The "all_sentences" are the sentence segments obtained after dividing a given text using sentence terminators (usually periods, exclamation marks, question marks, and ellipses in Chinese). The "complex_sentences" are sentences from the above all-sentence results whose character length exceeds a specified threshold.
[0077] Paragraph structural complexity: This is an indicator of paragraph length in a text, calculated as the proportion of complex paragraphs to the total number of paragraphs, i.e., len(complex_paragraphs) / len(all_paragraphs). Generally, higher paragraph structural complexity indicates higher text quality.
[0078] The `all_paragraphs` parameter is the segmented result of a given text after it has been split using paragraph terminators (newline characters).
[0079] Complex paragraphs: Paragraphs whose character length exceeds the specified threshold in all the above segmentation results.
[0080] Textual information entropy: Information entropy is an important concept in information theory, used to measure the amount of information contained in a random variable. In Natural Language Processing (NLP), information entropy can be used to measure the amount of information in a text. Generally, the higher the information entropy, the richer the information in the text, and the higher the quality of the text.
[0081] Text redundancy: represents the proportion of repeated sentences in the main clause and sub-clauses, i.e., len(duplicated_sentences) / len(all_sentences).
[0082] Among them, duplicated sentences are calculated as follows: for sentences that are part of a main clause, the similarity between sentences is calculated, and sentences with a similarity greater than a certain threshold are considered duplicated sentences.
[0083] Text degradation degree: the proportion of unique words in the total number of words, i.e.: len(unique_words) / len(all_words).
[0084] The unique words are those words that appear only once in the total word results.
[0085] Text readability refers to the ease or difficulty of reading text. It is calculated using the cntext library, and the result involves three indicators: readability1 (average number of words in each clause), readability2 (proportion of adverbs and conjunctions in each sentence), and readability3 (refer to Fog Index, readability3=(readability1+readability2)×0.5). In this example, readability3 is used.
[0086] Stop word percentage: The stop word percentage is the proportion of the number of stop words to the total number of words.
[0087] Stop words are characters that need to be filtered out when processing natural language text in information retrieval to save storage space and improve search efficiency. They mainly include punctuation marks and words that are used very frequently.
[0088] This application only describes the above eight quality inspection indicators. In other embodiments, other quality inspection indicators may be introduced, but no specific limitations are made here.
[0089] In this application, in order to improve computational efficiency, the maximum available resources can be determined, and then the number of threads can be determined based on the maximum available resources. Processing threads are generated based on the number of threads, and each first text is evenly distributed to each processing thread so that each processing thread can process the first text in parallel, thereby obtaining the index value of the quality detection index corresponding to each first text.
[0090] S106: Obtain the text filtering target and determine the initial filtering threshold corresponding to each quality detection indicator based on the text filtering target.
[0091] The text filtering target refers to the desired filtering ratio, that is, the desired filtering ratio of the first text. For example, if the goal is to filter out n% of the first text, then the text filtering target is n.
[0092] After determining the text filtering target, the number of quality inspection indicators can be identified. Then, based on the text filtering target and the number of quality inspection indicators, the initial filtering threshold for each quality inspection indicator is obtained. For example, assuming there are *a* quality inspection indicators, the initial filtering threshold can be *n% / a* of the extreme value. For instance, if the goal is to filter out 8% of low-quality text, then each indicator should filter out an average of 1% of the text, and the initial filtering threshold is 1% of the extreme value. Here, the extreme value can refer to the maximum and minimum values of a calculated quality inspection indicator.
[0093] In other embodiments, the text filtering target n% can be unevenly distributed, for example, by allocating the text filtering target n% based on the importance of a quality detection metric, and then determining the initial filtering threshold based on the allocated proportion and the extreme value of the quality detection metric.
[0094] S108: Obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index.
[0095] S110: Adjust the initial filtering threshold based on the correlation between quality detection indicators and text quality, as well as the text quality of the second text.
[0096] After the initial filtering threshold is determined, the text quality of the second text at the initial filtering threshold can be determined. The text quality of the second text can be characterized by the aforementioned quality detection indicators. Then, the initial filtering threshold is adjusted based on the text quality of the second text and its correlation.
[0097] In one optional embodiment, the initial filtering threshold is adjusted based on the correlation between the quality detection index and the text quality, as well as the text quality of the second text. This includes: lowering the initial filtering threshold when the correlation between the quality detection index and the text quality is positive and the text quality of the second text is higher than the quality threshold; and increasing the initial filtering threshold when the correlation between the quality detection index and the text quality is negative and the text quality of the second text is higher than the quality threshold.
[0098] Specifically, a positive correlation means that the higher the quality indicator, the higher the text quality. A negative correlation means that the higher the quality indicator, the lower the text quality. Generally, indicators with higher values for text quality include: word richness, sentence structure complexity, paragraph structure complexity, text information entropy, and text degradation; indicators with lower values for text quality include: text redundancy, text readability, and stop word percentage.
[0099] Because different datasets have different data distribution characteristics, this step cannot use a fixed threshold to filter text across all datasets. Instead, it requires manual sampling and observation to determine the threshold. The quality indicator values for each first text in the dataset are calculated sequentially. Following the rule of structure first, then content, and granularity from large to small, the observation order for each indicator is determined as follows: paragraph structure complexity, sentence structure complexity, word richness, stop word percentage, text redundancy, text degradation, text information entropy, and text readability. The observation starting point for each indicator value is determined based on the text filtering target. For example, if the text filtering target is to filter out 8% of low-quality text, then each indicator should filter out an average of 1% of the text, and the observation starting point is at the extreme value of 1%. For indicators with higher values indicating higher text quality, ascending order is used to observe the text quality at the initial filtering threshold. If the text quality is high, values are taken forward in a certain step size; conversely, values are taken backward in a certain step size until the text quality is low, thus determining the threshold. For indicators where smaller values indicate higher text quality, descending order is used to observe the text characteristics at the initial filtering threshold. If the text quality is high, values are taken sequentially with a certain step size; otherwise, values are taken sequentially with a certain step size until the text quality is low, thus determining the threshold.
[0100] Furthermore, it should be noted that the proportion of increase in positive correlation can be the same as the proportion of decrease in negative correlation, or the difference can be within a certain range. Similarly, the proportion of decrease in positive correlation can be the same as the proportion of increase in negative correlation, or the difference can be within a certain range, to ensure that the final filtering target remains constant.
[0101] S112: Based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each first text, filter the first text to obtain the third text.
[0102] Specifically, after determining the adjusted initial filtering threshold, the first text is filtered. Specifically, for positively correlated quality detection indicators, the first text that is less than the adjusted initial filtering threshold is deleted; for negatively correlated quality detection indicators, the first text that is greater than the adjusted initial filtering threshold is deleted.
[0103] The above text processing method, after obtaining the set of texts to be processed, determines quality detection indicators and calculates the indicator value of each first text corresponding to the quality detection indicator; obtains the text filtering target and determines the initial filtering threshold corresponding to each quality detection indicator based on the text filtering target; obtains the text quality of the second text at the initial filtering threshold corresponding to the quality detection indicator; adjusts the initial filtering threshold based on the correlation between the quality detection indicator and the text quality and the text quality of the second text; and filters the first text to obtain the third text based on the adjusted initial filtering threshold and the indicator value of each first text corresponding to the quality detection indicator. This method includes multiple quality detection indicators and dynamically determines the initial filtering threshold based on the text filtering target, thereby improving the filtering quality and ultimately improving the text processing quality.
[0104] In one optional embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index value of the quality detection index corresponding to each first text, the process includes: deduplicating the third text to obtain the fourth text; wherein the deduplication process includes at least one of the following: deduplicating based on the text summary of each third text; or obtaining the information array corresponding to each third text based on the index value of each quality detection index of each third text, and deduplicating the third text based on the information array.
[0105] In this embodiment, after processing the first text using quality detection metrics, most of the high-quality text is retained. To further reduce redundant information, the entire dataset needs to be deduplicated. The deduplication method can include at least one of precise deduplication or fuzzy deduplication. Precise deduplication can be based on the text summaries of each third text, for example, calculating the MD5 value of each third text, comparing them, removing texts with the same MD5 value, and retaining only one of them.
[0106] Fuzzy deduplication is based on quality detection metrics. The metric values of each quality detection metric of the third text are used to obtain the information array corresponding to each third text. For example, the eight metric values are formed into a list array [word richness, sentence structure complexity, paragraph structure complexity, text information entropy, text degradation degree, text redundancy, text readability, stop word ratio]. When the list arrays corresponding to any two texts are the same, they may be duplicated in content or have the same template. This can filter out texts that are not identified as having the same overall framework structure or texts with duplicate content.
[0107] In one optional embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: when the number of characters in the third text is greater than the target character processing capacity of the model, obtaining the text structure type corresponding to the third text; determining the text segmentation logic corresponding to the text structure type, and performing text segmentation on the third text based on the text segmentation logic.
[0108] Due to factors such as computing power, learning capacity, storage space, and inference efficiency, large models typically control the number of characters in the input text. Therefore, the final step in text processing is usually text segmentation. Assuming the maximum number of characters in a large model is `max_tokens`, if the total number of characters in the text is less than or equal to `max_tokens`, then the text does not need to be segmented and the entire text is retained. If the number of characters in the third text exceeds the model's target character processing capacity, then the third text needs to be segmented. The segmentation logic varies depending on the text structure type of the third text. For third texts with a chapter structure, segmentation is performed layer by layer based on the chapter structure hierarchy. For third texts without a chapter structure, segmentation is performed layer by layer based on the delimiter hierarchy. See below for details, until the number of characters in each segmented text is less than or equal to `max_tokens`.
[0109] In one optional embodiment, the text structure type identification step may be performed before text quality detection, such as before obtaining quality detection indicators, and may further include: performing template identification on the first text and performing deduplication processing based on the identified template; performing text structure type identification on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0110] The text structure has two aspects:
[0111] Firstly, this refers to the overall framework and structure of the text. Because financial texts often use the same templates, such as "xxx Credit Card Application Form," "xxx Loan Contract," and "xxx Corporate Business Agreement," most of the text within the same template is identical, differing only in the content filled in as personal / company information. In such cases, it's unnecessary to input large batches of these texts into the model; only one anonymized text needs to be selected for processing. Therefore, the corresponding template can be obtained by identifying the filename or subject name of each first text, and then the first text can be filtered based on the template.
[0112] Secondly, it refers to the document structure. This is mainly divided into two categories based on whether or not it has a chapter / section structure. Texts with a chapter / section structure generally include regulations, guidelines, manuals, and textbooks, and their possible formats are as follows: Figure 2 and Figure 3 As shown. The remaining data that is not organized according to the sequence number is classified into the category of not having a chapter structure. Among them, the chapter structure may be the required writing style, such as (1), (2), (3), ... First article, Second article, Third article, ... First part, Second part, Third part, ... a, b, c, ... (i), (ii), (iii), ... The above are just examples for illustration, and other chapter structures may also be used in other embodiments.
[0113] In one alternative embodiment, combined with Figure 4 As shown, Figure 4 This is a flowchart of a text segmentation step with a chapter structure in one embodiment. In this embodiment, the text segmentation logic corresponding to the text structure type is determined, and the third text is segmented based on the text segmentation logic, including: when the text structure type is a chapter structure, determining each layer of text according to the order of the chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continuing to segment the current layer of text according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0114] If the text has a chapter structure, it is first segmented according to the first-level flag. If the number of characters in the first-level text is greater than max_tokens, it is then segmented according to the second-level flag, and so on, until the number of characters in the current level is less than or equal to the model's target character processing capacity, at which point the segmentation ends. If the number of characters in the text is still greater than max_tokens when segmenting to the smallest level, the general segmentation method described below is used for segmentation.
[0115] In one alternative embodiment, combined with Figure 5 As shown, Figure 5 The flowchart of a general segmentation step in one embodiment further includes: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, taking the third text or the minimum level text as the text to be segmented, determining the first order of each delimiter, and obtaining the current delimiter based on the first order; segmenting the text to be segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, obtaining the next delimiter based on the first order of each delimiter as the current delimiter, and continuing to segment the current segmented text until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0116] In one optional embodiment, the method of segmenting the text to be segmented includes: determining the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtaining the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segmenting the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0117] In this embodiment, if the text does not have a chapter structure, it is directly segmented according to the delimiters. The order of the delimiters can be paragraph (line break) and sentence delimiters (period, question mark, ellipsis). First, it is segmented according to paragraphs. If the number of characters in a single paragraph exceeds max_tokens, the text is divided into sentences (delimiters: period, question mark, ellipsis) and concatenated as sentences.
[0118] In one optional embodiment, considering that when inputting the model, it is desirable for the slice information content of each text to be relatively uniform, the general paragraph segmentation method can be optimized: first, divide the number of text characters by max_tokens to obtain the number of slices num; then divide the number of text characters by num to obtain the average length of each slice average_length; use average_length as the max_tokens of the text to segment the text.
[0119] In one optional embodiment, before obtaining the quality inspection indicators, the method further includes: converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked; and performing text segmentation on the third text based on text segmentation logic, including: segmenting tables or formulas based on text segmentation logic, and marking the start and end points of each segmented table and formula, and marking the table header and continuation table identifier at the cut-off position of the table, and marking the continuation identifier at the cut-off position of the formula.
[0120] Format conversion is generally the first step in text processing. For unstructured documents, they first need to be converted into txt text format to facilitate subsequent cleaning work.
[0121] For example: docx documents can be converted using the python-docx library; html documents can be converted using the html2text library; pdf documents can be converted using OCR combined with a layout analysis model (such as PP-Structure); and epub ebooks can be converted using BeautifulSoup4 combined with the ebooklib library.
[0122] In particular, financial texts often involve a large number of tables and formulas. In order for large models to learn this kind of information, it is necessary to convert tables to Markdown format (which can be achieved through the texttable library), convert formulas to LaTeX format, and mark the beginning and end of tables and formulas. In other words, treat each table and formula as an independent paragraph.
[0123] Optionally, during the first step of format conversion, the table has been converted to Markdown format and the formula has been converted to LaTeX format, and marks have been made at the beginning and end of the table. The integrity of the table and formula needs to be maintained during the splitting process. If the number of characters in a single table exceeds max_tokens, the table will be truncated row by row, the original table header will be added to the truncated section, and a table title note will be added at the beginning of the slice to indicate that the table continues.
[0124] In one optional embodiment, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index value of the quality detection index corresponding to each first text, the process includes: determining the text type of the third text; and cleaning the third text based on the text type.
[0125] In this embodiment, the following issues need to be addressed during cleaning:
[0126] For HTML text, remove menu navigation information, page statistics, special characters, etc. from the webpage.
[0127] Remove headers, footers, table of contents, and special characters from PDF / docx / epub text files.
[0128] For ease of understanding, combined with Figure 6 As shown, Figure 6 The flowchart below shows a text processing method in another embodiment. In this embodiment, a set of texts to be processed is obtained, which includes several first texts. First, the format of each first text is converted. Then, the text structure type of each first text after format conversion is identified, including two types: chapter structure and chapter structure. Third, the first texts undergo quality detection, and the first texts are filtered based on the quality detection results to obtain third texts. The quality detection steps can be found in S102 to S112 above. Fourth, the third texts undergo deduplication, including precise deduplication and fuzzy deduplication. Fifth, the deduplicated third texts undergo text cleaning. Sixth, the cleaned third texts undergo text segmentation. The specific segmentation method can be combined with... Figure 4 and Figure 5 As shown.
[0129] In the above embodiments, a preprocessing method for unstructured document data was established, comprising six stages: document format conversion, text structure classification, text quality detection, text deduplication, text cleaning, and text slicing. In the document format conversion stage, various open-source tools were used to convert unstructured documents such as pdf / docx / epub / html into txt text, and tables were converted to Markdown format and formulas to LaTeX format, facilitating the learning of mathematical information by large models. In the text structure classification stage, templates with the same overall framework were identified to reduce the processing of repetitive text; furthermore, the text structure was classified, allowing for different processing schemes in the text slicing stage. In the text quality detection stage, eight indicators quantifying the quality of text structure and content were used. The calculation methods for these indicators are quick and simple, providing a direct basis for manually filtering low-quality text. In the text deduplication stage, in addition to conventional precise deduplication, a fuzzy deduplication method using the eight quality detection indicators was added, which helps to further identify texts with the same framework structure and remove redundant information from the dataset. In the text slicing stage, two slicing schemes were set according to the text structure, covering all text types and obtaining high-quality slicing results.
[0130] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0131] Based on the same inventive concept, this application also provides a text processing apparatus for implementing the text processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more text processing apparatus embodiments provided below can be found in the limitations of the text processing method described above, and will not be repeated here.
[0132] In one exemplary embodiment, such as Figure 7 As shown, a text processing apparatus is provided, including: a text set acquisition module 701, an index value calculation module 702, an initial filtering threshold determination module 703, a text quality determination module 704, a threshold adjustment module 705, and a filtering module 706, wherein:
[0133] The unprocessed text set acquisition module 701 is used to acquire the unprocessed text set, which includes several first texts;
[0134] The indicator value calculation module 702 is used to obtain quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each first text.
[0135] The initial filtering threshold determination module 703 is used to obtain the text filtering target and determine the initial filtering threshold corresponding to each quality detection index based on the text filtering target.
[0136] The text quality determination module 704 is used to obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index;
[0137] The threshold adjustment module 705 is used to adjust the initial filtering threshold based on the correlation between the quality detection index and the text quality and the text quality of the second text, including: when the correlation between the quality detection index and the text quality is positive and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is reduced; when the correlation between the quality detection index and the text quality is negative and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is increased.
[0138] The filtering module 706 is used to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text.
[0139] In one optional embodiment, the apparatus further includes: a text deduplication module for deduplicating the third text to obtain a fourth text; wherein the deduplication process includes at least one of the following: deduplication based on the text summary of each third text; or obtaining an information array corresponding to each third text based on the index value of each quality detection index of each third text, and deduplicating the third text based on the information array.
[0140] In one optional embodiment, the device further includes: a text segmentation module, used to obtain the text structure type corresponding to the third text when the number of characters in the third text is greater than the target character processing capacity of the model; determine the text segmentation logic corresponding to the text structure type; and perform text segmentation on the third text based on the text segmentation logic.
[0141] In one optional embodiment, the device further includes a structure classification module for performing template recognition on the first text and deduplication based on the recognized template; and for performing text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0142] In one optional embodiment, the text segmentation module is further configured to determine each layer of text according to the order of the chapter structure when the text structure type is a chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, the text in the current layer of text continues to be segmented according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0143] In one optional embodiment, the text segmentation module is further configured to, when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, take the third text or the minimum level text as the text to be segmented, determine the first order of each delimiter, and obtain the current delimiter based on the first order; segment the text to be segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, obtain the next delimiter based on the first order of each delimiter as the current delimiter, and continue to segment the current segmented text until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0144] In one optional embodiment, the text segmentation module is further configured to determine the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtain the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segment the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0145] In one optional embodiment, the apparatus further includes: a format conversion module for converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked;
[0146] The text segmentation module is also used to segment tables or formulas based on text segmentation logic, and to mark the start and end points of each segmented table and formula. It also marks the table header and continuation mark at the cut position of the table, and marks the continuation mark at the cut position of the formula.
[0147] In one alternative embodiment, the apparatus further includes a cleaning module for determining the text type of the third text and cleaning the third text based on the text type.
[0148] Each module in the aforementioned text processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0149] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a text processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0150] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: acquiring a set of texts to be processed, the set of texts to be processed including a plurality of first texts; acquiring quality detection indicators and calculating the indicator value of the quality detection indicator corresponding to each first text; acquiring a text filtering target and determining an initial filtering threshold corresponding to each quality detection indicator based on the text filtering target; acquiring the text quality of a second text at the initial filtering threshold corresponding to the quality detection indicator; adjusting the initial filtering threshold based on the correlation between the quality detection indicator and the text quality and the text quality of the second text, including: decreasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is positive and the text quality of the second text is higher than the quality threshold; increasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is negative and the text quality of the second text is higher than the quality threshold; and filtering the first texts to obtain a third text based on the adjusted initial filtering threshold and the indicator value of the quality detection indicator corresponding to each first text.
[0152] In one embodiment, after the processor executes a computer program to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each first text, the process includes: performing deduplication on the third text to obtain the fourth text; wherein the deduplication process includes at least one of the following: performing deduplication based on the text summary of each third text; or obtaining the information array corresponding to each third text based on the index values of each quality detection indicator of each third text, and performing deduplication on the third text based on the information array.
[0153] In one embodiment, after the processor executes the computer program to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: when the number of characters in the third text is greater than the target character processing capacity of the model, obtaining the text structure type corresponding to the third text; determining the text segmentation logic corresponding to the text structure type, and performing text segmentation on the third text based on the text segmentation logic.
[0154] In one embodiment, before the processor executes the computer program to acquire the quality detection index, the method further includes: performing template recognition on the first text and performing deduplication processing based on the recognized template; and performing text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0155] In one embodiment, when the processor executes a computer program, it implements text segmentation logic to determine the text structure type. Based on the text segmentation logic, it performs text segmentation on the third text, including: when the text structure type is a chapter structure, determining each layer of text according to the order of the chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continuing to segment the current layer of text according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0156] In one embodiment, when the processor executes the computer program, it further implements the following steps: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, the third text or the minimum level text is taken as the text to be segmented, the first order of each delimiter is determined, and the current delimiter is obtained based on the first order; the text to be segmented is segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, the next delimiter is obtained based on the first order of each delimiter as the current delimiter, and the current segmented text is continued to be segmented until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0157] In one embodiment, the method of segmenting the text to be segmented implemented by the processor when executing the computer program includes: determining the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtaining the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segmenting the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0158] In one embodiment, before the processor executes the computer program to acquire quality inspection indicators, the method further includes: converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked; the text segmentation of the third text based on text segmentation logic implemented by the processor when executing the computer program includes: segmenting tables or formulas based on text segmentation logic, and marking the start and end points of each segmented table and formula, and marking the table header and continuation table identifier at the cut-off position of the table, and marking the continuation identifier at the cut-off position of the formula.
[0159] In one embodiment, after the processor executes a computer program to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: determining the text type of the third text; and cleaning the third text based on the text type.
[0160] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program performs the following steps: acquiring a set of texts to be processed, the set of texts to be processed including a plurality of first texts; acquiring quality detection indicators and calculating the indicator value of the quality detection indicator corresponding to each first text; acquiring a text filtering target and determining an initial filtering threshold corresponding to each quality detection indicator based on the text filtering target; acquiring the text quality of a second text at the initial filtering threshold corresponding to the quality detection indicator; adjusting the initial filtering threshold based on the correlation between the quality detection indicator and the text quality and the text quality of the second text, including: decreasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is positive and the text quality of the second text is higher than the quality threshold; increasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is negative and the text quality of the second text is higher than the quality threshold; and filtering the first texts to obtain a third text based on the adjusted initial filtering threshold and the indicator value of the quality detection indicator corresponding to each first text.
[0161] In one embodiment, when the computer program is executed by the processor, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each first text, the process includes: performing deduplication on the third text to obtain the fourth text; wherein the deduplication process includes at least one of the following: performing deduplication based on the text summary of each third text; or obtaining the information array corresponding to each third text based on the index values of each quality detection indicator of each third text, and performing deduplication on the third text based on the information array.
[0162] In one embodiment, when the computer program is executed by the processor, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: when the number of characters in the third text is greater than the target character processing capacity of the model, obtaining the text structure type corresponding to the third text; determining the text segmentation logic corresponding to the text structure type, and performing text segmentation on the third text based on the text segmentation logic.
[0163] In one embodiment, before the computer program is executed by the processor to acquire quality detection indicators, the method further includes: performing template recognition on the first text and performing deduplication processing based on the recognized template; and performing text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0164] In one embodiment, when a computer program is executed by a processor, it implements text segmentation logic corresponding to the text structure type. Based on the text segmentation logic, it performs text segmentation on the third text, including: when the text structure type is a chapter structure, determining each layer of text according to the order of the chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continuing to segment the current layer of text according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0165] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, the third text or the minimum level text is taken as the text to be segmented, the first order of each delimiter is determined, and the current delimiter is obtained based on the first order; the text to be segmented is segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, the next delimiter is obtained based on the first order of each delimiter as the current delimiter, and the current segmented text is continued to be segmented until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0166] In one embodiment, the method of segmenting the text to be segmented implemented when the computer program is executed by the processor includes: determining the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtaining the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segmenting the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0167] In one embodiment, before the computer program is executed by the processor to acquire quality inspection indicators, the method further includes: converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked; and the computer program, when executed by the processor, performs text segmentation on the third text based on text segmentation logic, including: segmenting tables or formulas based on text segmentation logic, marking the start and end points of each segmented table and formula, and marking the table header and continuation table identifier at the cut-off position of the table, and marking the continuation identifier at the cut-off position of the formula.
[0168] In one embodiment, after the computer program is executed by the processor to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: determining the text type of the third text; and cleaning the third text based on the text type.
[0169] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps: acquiring a set of texts to be processed, the set of texts to be processed including a plurality of first texts; acquiring quality detection indicators and calculating the indicator value of the quality detection indicator corresponding to each first text; acquiring a text filtering target and determining an initial filtering threshold corresponding to each quality detection indicator based on the text filtering target; acquiring the text quality of a second text at the initial filtering threshold corresponding to the quality detection indicator; adjusting the initial filtering threshold based on the correlation between the quality detection indicator and the text quality and the text quality of the second text, including: decreasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is positive and the text quality of the second text is higher than the quality threshold; increasing the initial filtering threshold when the correlation between the quality detection indicator and the text quality is negative and the text quality of the second text is higher than the quality threshold; and filtering the first texts to obtain a third text based on the adjusted initial filtering threshold and the indicator value of the quality detection indicator corresponding to each first text.
[0170] In one embodiment, when the computer program is executed by the processor, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each first text, the process includes: performing deduplication on the third text to obtain the fourth text; wherein the deduplication process includes at least one of the following: performing deduplication based on the text summary of each third text; or obtaining the information array corresponding to each third text based on the index values of each quality detection indicator of each third text, and performing deduplication on the third text based on the information array.
[0171] In one embodiment, when the computer program is executed by the processor, after filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: when the number of characters in the third text is greater than the target character processing capacity of the model, obtaining the text structure type corresponding to the third text; determining the text segmentation logic corresponding to the text structure type, and performing text segmentation on the third text based on the text segmentation logic.
[0172] In one embodiment, before the computer program is executed by the processor to acquire quality detection indicators, the method further includes: performing template recognition on the first text and performing deduplication processing based on the recognized template; and performing text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
[0173] In one embodiment, when a computer program is executed by a processor, it implements text segmentation logic corresponding to the text structure type. Based on the text segmentation logic, it performs text segmentation on the third text, including: when the text structure type is a chapter structure, determining each layer of text according to the order of the chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continuing to segment the current layer of text according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
[0174] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, the third text or the minimum level text is taken as the text to be segmented, the first order of each delimiter is determined, and the current delimiter is obtained based on the first order; the text to be segmented is segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, the next delimiter is obtained based on the first order of each delimiter as the current delimiter, and the current segmented text is continued to be segmented until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
[0175] In one embodiment, the method of segmenting the text to be segmented implemented when the computer program is executed by the processor includes: determining the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtaining the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segmenting the text to be segmented based on the length of the segmented text to obtain the current segmented text.
[0176] In one embodiment, before the computer program is executed by the processor to acquire quality inspection indicators, the method further includes: converting the format of the first text, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked; and the computer program, when executed by the processor, performs text segmentation on the third text based on text segmentation logic, including: segmenting tables or formulas based on text segmentation logic, marking the start and end points of each segmented table and formula, and marking the table header and continuation table identifier at the cut-off position of the table, and marking the continuation identifier at the cut-off position of the formula.
[0177] In one embodiment, after the computer program is executed by the processor to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection index corresponding to each first text, the process includes: determining the text type of the third text; and cleaning the third text based on the text type.
[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0179] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0181] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A text processing method, characterized in that, The method includes: Obtain a set of texts to be processed, the set of texts to be processed including a number of first texts; Obtain quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each of the first texts; Obtain the text filtering target, and determine the initial filtering threshold corresponding to each of the quality detection indicators based on the text filtering target and the number of quality detection indicators; Obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index; Based on the correlation between the quality detection index and text quality, and the text quality of the second text, the initial filtering threshold is adjusted, including: when the correlation between the quality detection index and the text quality is positive and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is decreased; when the correlation between the quality detection index and the text quality is negative and the text quality of the second text is higher than the quality threshold, the initial filtering threshold is increased. Based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the first texts are filtered to obtain the third text.
2. The method according to claim 1, characterized in that, After filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes: The third text is deduplicated to obtain the fourth text; The deduplication process includes at least one of the following: performing deduplication based on the text summary of each of the third texts; or obtaining an information array corresponding to each of the third texts based on the index value of each of the quality detection indicators of each of the third texts, and performing deduplication on the third texts based on the information array.
3. The method according to claim 1, characterized in that, After filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes: When the number of characters in the third text is greater than the target character processing capacity of the model, the text structure type corresponding to the third text is obtained; Determine the text segmentation logic corresponding to the text structure type, and perform text segmentation on the third text based on the text segmentation logic.
4. The method according to claim 3, characterized in that, Before obtaining the quality testing indicators, the process also includes: The first text is subjected to template recognition, and deduplication is performed based on the recognized template; The first text after deduplication is subjected to text structure type identification, wherein the text structure type includes those with chapter structure and those without chapter structure.
5. The method according to claim 3, characterized in that, The step of determining the text segmentation logic corresponding to the text structure type, and performing text segmentation on the third text based on the text segmentation logic, includes: When the text structure type is a chapter structure, each layer of text is determined according to the order of the chapter structure; If the number of characters in the current layer of text is greater than the target character processing capacity of the model, continue to segment the text in the current layer according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
6. The method according to claim 5, characterized in that, The method further includes: When the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, the third text or the minimum level text is taken as the text to be segmented, the first order of each delimiter is determined, and the current delimiter is obtained based on the first order. The text to be segmented is segmented based on the current delimiter to obtain the current segmented text; When the number of characters in the current segmented text is greater than the target character processing capacity of the model, the next delimiter is obtained based on the first order of each delimiter as the current delimiter, and the current segmented text continues to be segmented until the number of characters in the current segmented text after each segmentation is less than or equal to the target character processing capacity of the model, and the segmentation ends.
7. The method according to claim 6, characterized in that, The methods for segmenting the text to be segmented include: The number of slices is determined based on the number of characters in the text to be segmented and the target character processing capacity of the model. The length of the segmented text is obtained based on the number of slices and the number of characters in the text to be segmented; Based on the length of the segmented text, the text to be segmented is segmented to obtain the current segmented text.
8. The method according to claim 3, characterized in that, Before obtaining the quality testing indicators, the process also includes: The first text is formatted, wherein tables in the first text are converted to Markdown format, formulas are converted to LaTeX format, and the start and end points of the tables and formulas are marked. The text segmentation of the third text based on the text segmentation logic includes: The table or formula is segmented based on the text segmentation logic, and the start and end points of each segmented table and formula are marked. The table header and continuation mark are marked on the next slice at the cut position of the table, and the continuation mark is marked on the next slice at the cut position of the formula.
9. The method according to claim 1, characterized in that, After filtering the first text to obtain the third text based on the adjusted initial filtering threshold and the index values of the quality detection indicators corresponding to each of the first texts, the process includes: Determine the text type of the third text; The third text is cleaned based on the text type.
10. A text processing device, characterized in that, The device includes: The module for obtaining a set of texts to be processed is used to obtain a set of texts to be processed, which includes several first texts. The indicator value calculation module is used to obtain quality inspection indicators and calculate the indicator value of the quality inspection indicator corresponding to each of the first texts. An initial filtering threshold determination module is used to obtain text filtering targets and determine the initial filtering threshold corresponding to each quality detection indicator based on the text filtering targets and the number of quality detection indicators. The text quality determination module is used to obtain the text quality of the second text at the initial filtering threshold corresponding to the quality detection index; A threshold adjustment module is used to adjust the initial filtering threshold based on the correlation between the quality detection index and the text quality and the text quality of the second text, including: lowering the initial filtering threshold when the correlation between the quality detection index and the text quality is positive and the text quality of the second text is higher than the quality threshold; and increasing the initial filtering threshold when the correlation between the quality detection index and the text quality is negative and the text quality of the second text is higher than the quality threshold. The filtering module is used to filter the first text to obtain the third text based on the adjusted initial filtering threshold and the index value of the quality detection index corresponding to each of the first texts.
11. The apparatus according to claim 10, characterized in that, The device further includes: a text deduplication module, used to perform deduplication processing on the third text to obtain a fourth text; wherein the deduplication processing includes at least one of the following: performing deduplication processing based on the text summary of each of the third texts; or obtaining an information array corresponding to each of the third texts based on the index value of each of the quality detection indicators of each of the third texts, and performing deduplication processing on the third texts based on the information array.
12. The apparatus according to claim 10, characterized in that, The device further includes: a text segmentation module, used to obtain the text structure type corresponding to the third text when the number of characters in the third text is greater than the target character processing capacity of the model; determine the text segmentation logic corresponding to the text structure type; and perform text segmentation on the third text based on the text segmentation logic.
13. The apparatus according to claim 12, characterized in that, The device further includes a structure classification module, used to perform template recognition on the first text and perform deduplication processing based on the recognized template; and to perform text structure type recognition on the deduplicated first text, wherein the text structure type includes having a chapter structure and not having a chapter structure.
14. The apparatus according to claim 12, characterized in that, The text segmentation module is further configured to determine each layer of text according to the order of the chapter structure when the text structure type is a chapter structure; when the number of characters in the current layer of text is greater than the target character processing capacity of the model, continue to segment the text in the current layer according to the order of the chapter structure until the number of characters in the current layer of text is less than or equal to the target character processing capacity of the model, at which point the segmentation ends.
15. The apparatus according to claim 14, characterized in that, The text segmentation module is further configured to: when the text structure type is without a chapter structure, or when the minimum level of the chapter structure is reached and the number of characters in the minimum level text is greater than the target character processing capacity of the model, take the third text or the minimum level text as the text to be segmented, determine the first order of each delimiter, and obtain the current delimiter based on the first order; segment the text to be segmented based on the current delimiter to obtain the current segmented text; when the number of characters in the current segmented text is greater than the target character processing capacity of the model, obtain the next delimiter based on the first order of each delimiter as the current delimiter, and continue to segment the current segmented text until the number of characters in each segmented current segmented text is less than or equal to the target character processing capacity of the model, and the segmentation ends.
16. The apparatus according to claim 15, characterized in that, The text segmentation module is further configured to determine the number of slices based on the number of characters in the text to be segmented and the target character processing capacity of the model; obtain the length of the segmented text based on the number of slices and the number of characters in the text to be segmented; and segment the text to be segmented based on the length of the segmented text to obtain the current segmented text.
17. The apparatus according to claim 12, characterized in that, The device further includes: a format conversion module, used to convert the format of the first text, wherein tables in the first text are converted to markdown format, formulas are converted to latex format, and the start and end points of the tables and formulas are marked; The text segmentation module is also used to segment the table or the formula based on the text segmentation logic, and mark the start and end points of each segmented table and formula, and mark the table header and continuation mark at the cut position of the table, and mark the continuation mark at the cut position of the formula.
18. The apparatus according to claim 10, characterized in that, The device further includes a cleaning module for determining the text type of the third text and cleaning the third text based on the text type.
19. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Text correction method and device, electronic equipment and storage medium
CN111832288A
Text processing model training method and device and computer equipment
CN115841109A