Text data processing method and device, electronic equipment and storage medium
By deduplication and purification of text data in the initial data set, quality evaluation and text classifier training, the problems of low efficiency and low quality of text data processing in the prior art are solved, and efficient and flexible text data processing and identification of high-quality data are achieved.
Patent Information
- Application Number
- CN202411978676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-02
AI Technical Summary
The processing methods for text data in the prior art have problems of low processing efficiency and low text quality.
By deduplication of the text data in the initial dataset, purifying the filtered dataset according to the preprocessing configuration rules, and quality evaluation is performed on each text data in the preprocessing dataset, high-quality text data is selected to train the text classifier, and finally using the text classifier to identify high-quality Chinese data.
It improves the processing efficiency of text data, reduces the workload of data cleaning, improves the flexibility of data cleaning, and avoids performance losses caused by large-scale data processing by identifying high-quality data, improving judgment accuracy and task execution efficiency.
Smart Images

Figure CN119917666A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology; specifically, to a text data processing method, device, electronic device and storage medium. Background Art
[0002] In the current field of artificial intelligence, the rapid development and application of large language models has led to an increasing demand for high-quality Chinese data. Large language models use deep learning technology to perform tasks such as language understanding, generation, and dialogue, but their performance and effectiveness directly benefit from the quality and quantity of training data; a typical application scenario is natural language processing tasks such as intelligent customer service, machine translation, and information retrieval. In these applications, the model needs to understand and generate complex Chinese language structures in order to effectively process user requests and provide accurate answers.
[0003] In order for large language models to perform well in the Chinese environment, large-scale text datasets are needed to train and fine-tune the models. The current methods for obtaining text datasets usually include: manually analyzing data, formulating cleaning rules, programming to implement cleaning rules, running programs for data cleaning, manual quality assessment and feedback, etc. The work is detailed and complicated, requiring a lot of manual participation and code development iteration, which is time-consuming and labor-intensive. In addition, there is a considerable amount of low-quality text data that cannot be filtered out by manually formulated rules.
[0004] It can be seen that the text data processing methods in the related art have the problems of low processing efficiency and low text quality. Summary of the invention
[0005] To solve the above technical problems, the embodiments of the present application provide a text data processing method, device, electronic device and storage medium, which solve the problems of low processing efficiency and low text quality in the text data processing methods in the related arts.
[0006] According to one aspect of an embodiment of the present application, a text data processing method is provided, the method comprising: performing deduplication processing on text data in an initial data set to obtain a filtered data set; wherein the initial data set comprises a data set constructed based on Internet data; performing purification processing on the filtered data set according to a preprocessing configuration rule to obtain a preprocessed data set; performing quality assessment on each text data in the preprocessed data set, and forming a text training sample set with text data having a high quality assessment result; training a text classifier based on the text training sample set to obtain a trained target text classifier; and evaluating and classifying the input text based on the target text classifier to obtain a quality classification result of the input text.
[0007] Optionally, deduplication processing is performed on the text data in the initial data set to obtain a filtered data set, including: performing word segmentation on each text data in the initial data set to obtain multiple feature words corresponding to each text data; obtaining a text signature corresponding to each text data based on the multiple feature words corresponding to each text data; calculating a similarity value between each text data based on the text signature corresponding to each text data; and deduplication of the text data in the initial data set based on the similarity value to obtain the filtered data set.
[0008] Optionally, based on the multiple feature words corresponding to each text data, a text signature corresponding to each text data is obtained, including: calculating the hash value of each feature word in the current text data according to the same hash function; calculating the target weight of each feature word according to the frequency of occurrence of each feature word in the current text data; obtaining the weighted feature value of each feature word according to the target weight and hash value of each feature word; accumulating each weighted feature value in the current text data to obtain a merged feature vector; and performing dimensionality reduction processing on the merged feature vector to obtain the text signature corresponding to the current text data.
[0009] Optionally, the target weight of each feature word is calculated according to the frequency of occurrence of each feature word in the current text data, including: obtaining the initial similarity of feature word x and feature word y according to the number of samples in which feature word x and feature word y appear in the current sample data; wherein feature word x and feature word y are respectively two different feature words from multiple feature words in the current sample data; according to the number of times feature word x and feature word y appear in the current sample data, the similarity of feature word x and feature word y is weighted to obtain the target similarity of feature word x and feature word y; according to the frequency of occurrence of each feature word in the current sample data, the initial weight of each feature word is obtained; according to the target similarity of feature word x and feature word y, the initial weight of each feature word is optimized to obtain the target weight of each feature word.
[0010] Optionally, a quality assessment is performed on each text data in the preprocessed data set, including: dividing the current text data into several paragraphs to obtain multiple short texts; obtaining an assessment result of each short text based on the syntactic structure level and tightness level of each short text; and weighting the assessment result of each short text to obtain a quality assessment result of the current text data.
[0011] Optionally, before obtaining the evaluation result of each short text according to the syntactic structure level and the closeness level in each short text, the method further includes: segmenting the current short text to obtain a segmentation list corresponding to the current short text; obtaining the number of predicates and the number of agent-patient relations of the predicates in the current short text according to the segmentation list; and obtaining the syntactic structure level of the current short text according to the ratio of the relationship number to the number of predicates.
[0012] Optionally, before obtaining the evaluation result of each short text according to the syntactic structure level and the closeness level in each short text, the method also includes: segmenting the current short text to obtain a segmentation list corresponding to the current short text; taking each segmentation as a node and calculating the closeness value of each node relationship; obtaining the sentence closeness value corresponding to the current short text according to the closeness value of each node relationship; and obtaining the closeness level corresponding to the current short text according to the sentence closeness value.
[0013] According to one aspect of an embodiment of the present application, a text data processing device is provided, the device comprising: a data filtering module, used to perform deduplication processing on text data in an initial data set to obtain a filtered data set; wherein the initial data set comprises a data set constructed based on Internet data; a data preprocessing module, used to perform purification processing on the filtered data set according to a preprocessing configuration rule to obtain a preprocessed data set; a data quality screening module, used to perform quality assessment on each text data in the preprocessed data set, and form a text training sample set for text data with high quality assessment results; and also used to train a text classifier based on the text training sample set to obtain a trained target text classifier; and also used to evaluate and classify input text based on the target text classifier to obtain a quality classification result of the input text.
[0014] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the text data processing method in the above technical solution is implemented.
[0015] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions so that the electronic device implements a text data processing method as in the above technical solution.
[0016] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the text data processing method in the above technical solution is implemented.
[0017] The technical solution provided by this application includes at least the following beneficial effects:
[0018] The present application improves the processing efficiency of text data by performing deduplication processing on a large amount of Internet data; the filtered data set is purified based on preprocessing configuration rules, and only the cleaning rules need to be defined in advance, and no coding is required. While greatly reducing the workload of data cleaning, it also improves the flexibility of data cleaning; further, the present application performs quality assessment on the purified text data, screens out text data with high quality assessment results to train the text classifier, so that the trained target text classifier has good classification efficiency, and finally uses the text classifier to identify high-quality Chinese data, which not only ensures the accuracy of data judgment, but also avoids performance loss caused by large models processing large amounts of data. While improving the judgment accuracy, it also improves the task execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0020] Figure 1 The figure is a flow chart of a text data processing method provided by an embodiment of the present application;
[0021] Figure 2 Shown Figure 1 An exemplary flow chart of step S10 in FIG.
[0022] Figure 3 Shown Figure 2 An exemplary flow chart of step S120 in FIG.
[0023] Figure 4 Shown Figure 1 An exemplary flow chart of step S30 in FIG.
[0024] Figure 5 The figure is a schematic diagram of the structure of a text data processing device provided in an embodiment of the present application;
[0025] Figure 6 The figure is a schematic diagram of a processing flow of a text data processing device provided in an embodiment of the present application;
[0026] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0027] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.
[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0030] It should also be noted that the "multiple" mentioned in this application refers to two or more than two. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship.
[0031] Key terms explained:
[0032] Large Language Model (LLM): refers to a deep learning model trained with a large amount of text data that can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc., and are an important path to artificial intelligence.
[0033] Figure 1 FIG. 1 is a flow chart of a text data processing method provided in an embodiment of the present application; Figure 1 As shown, the method specifically comprises the following steps:
[0034] Step S10: De-duplicate the text data in the initial data set to obtain a filtered data set.
[0035] In this embodiment, the initial data set includes a data set constructed based on Internet data, that is, a large amount of Internet data is obtained from the Internet in the form of a crawler to construct a basic data set as the initial data set; wherein the initial data set includes a number of text data.
[0036] It should be noted that since the text data in the initial data set is a large amount of Internet data obtained from the Internet, there is a lot of data with similar or identical content; in order to improve data processing efficiency, it is necessary to deduplicate redundant data. The deduplication processing in this embodiment includes deduplication of redundant data within the text, and also includes deduplication of redundant text between texts.
[0037] In one embodiment, before deduplication processing is performed on the text data, non-Chinese data is filtered on the initial data set, so that the text data in the initial data set are all Chinese monolingual data.
[0038] Step S20: purify the filtered data set according to the preprocessing configuration rule to obtain a preprocessed data set.
[0039] It should be noted that the preprocessing configuration rules include the number of text characters, sensitive words and domain keywords, that is, documents with fewer characters, erroneous characters or harmful information, and non-conforming domains are filtered out through pre-established configuration rules; in addition, the preprocessing configuration rules include a rule library for storing data processing rules with customized usage. The preprocessing controller calls the rule interpreter to purify the text data in the filtered data set based on the preprocessing configuration rules, thereby obtaining a preprocessed data set, thereby improving the quality of the text data in the preprocessed data set.
[0040] Step S30: Perform quality assessment on each text data in the preprocessed data set, and form a text training sample set with text data having high quality assessment results.
[0041] In this embodiment, during the text purification process, although the preprocessing configuration rules are used to filter out the explicit noise text in the data set, there is still a large amount of low-quality text data in the remaining data that cannot be filtered by the configuration rules; based on this, the quality of each text data in the preprocessing data set can be evaluated through a large language model, and high-quality text data can be screened out to form a text training sample set.
[0042] Step S40: training a text classifier according to the text training sample set to obtain a trained target text classifier.
[0043] Step S50: Evaluate and classify the input text based on the target text classifier to obtain a quality classification result of the input text.
[0044] It should be noted that the text classifier is trained with the text data in the text training sample set so that the trained text classifier can classify high-quality text data; finally, the target text classifier can be used to identify high-quality Chinese data, which not only improves the efficiency of text data classification, but also ensures the accuracy of data judgment.
[0045] It can be seen that the present application improves the processing efficiency of text data by performing deduplication processing on a large amount of Internet data; the filtered data set is purified based on the preprocessing configuration rules, and only the cleaning rules need to be defined in advance, and no coding is required. While greatly reducing the workload of data cleaning, it also improves the flexibility of data cleaning; further, the present application performs quality assessment on the purified text data, screens out text data with high quality assessment results to train the text classifier, so that the trained target text classifier has good classification efficiency, and finally uses the text classifier to identify high-quality Chinese data, which not only ensures the accuracy of data judgment, but also avoids performance loss caused by large models processing large amounts of data. While improving the judgment accuracy, it also improves the task execution efficiency.
[0046] In one embodiment, if Figure 2 As shown, deduplication processing is performed on the initial data set to obtain a filtered data set, which specifically includes the following steps:
[0047] Step S110: Segment each text data in the initial data set to obtain multiple feature words corresponding to each text data.
[0048] It should be noted that the Jieba word segmentation tool is used to perform word segmentation processing on each text data, and after removing special symbols, stop words and other irrelevant words, multiple feature words corresponding to each text data are obtained.
[0049] Step S120, obtaining a text signature corresponding to each text data according to the multiple feature words corresponding to each text data;
[0050] In one embodiment, if Figure 3 As shown, according to the multiple feature words corresponding to each text data, obtaining the text signature corresponding to each text data specifically includes the following steps:
[0051] Step S1211, calculating the hash value of each feature word according to the same hash function;
[0052] It should be noted that this embodiment uses the same hash function to calculate the hash value of each feature word.
[0053] Step S1212: Calculate the target weight of each feature word according to the occurrence frequency of each feature word in the current text data.
[0054] In this embodiment, according to the frequency of occurrence of each feature word in the current text data, calculating the target weight of each feature word specifically includes: obtaining the initial similarity of feature word x and feature word y according to the number of samples in which feature word x and feature word y appear in the current sample data; wherein feature word x and feature word y are respectively taken as two different feature words from multiple feature words in the current sample data; according to the number of times feature word x and feature word y appear in the current sample data respectively, the similarity of feature word x and feature word y is weighted to obtain the target similarity of feature word x and feature word y; according to the frequency of occurrence of each feature word in the current sample data, the initial weight of each feature word is obtained; according to the target similarity of feature word x and feature word y, the initial weight of each feature word is optimized to obtain the target weight of each feature word.
[0055] It should be noted that the calculation formula for the initial similarity between feature word x and feature word y is:
[0056]
[0057] Among them, I(x,y) represents the initial similarity between feature words x and feature words y, and f 11 represents the number of samples where x is 1 and y is 1, f 01 represents the number of samples where x is 0 and y is 1, f 10 It represents the number of samples where x is 1 and y is 0. In formula (1), if x is 1, it means that the sample contains the feature word x, otherwise it does not contain it. The same is true for y.
[0058] In practical applications, even if feature words x and y appear in multiple samples at the same time, their number of occurrences will be random. In order to eliminate the influence of dimension, the following improvements are made:
[0059]
[0060] In formula (2), J(x,y) represents the target similarity between feature words x and y, n represents the total number of samples containing both feature words x and y, and x k represents the number of times the feature word x appears in the kth sample, y k Indicates the number of times the feature word y appears in the kth sample.
[0061] Get the weight of the feature word in each text. The specific formula is:
[0062] ω dt =tf dt ×idf(N nt )
[0063] Among them, ω dtrepresents the weight of feature word t in text d, tf dt represents the frequency of the feature word t in the text d, N represents the total number of text data in the initial data set, idf(N nt ) represents the inverse document frequency, which is the logarithm of the ratio of the total number of text data N in the initial data set to the number of texts n in which the feature word t appears, and is used to weigh the importance of the feature word. In practical applications, in order to reduce the impact of text length, the feature word weight needs to be normalized. The specific formula is:
[0064]
[0065] Among them, m dt represents the number of times the feature word t appears in the text d, M d Represents the total number of feature words in text d, n t Indicates the number of texts in which feature word t appears in the text set.
[0066] The weight is optimized according to the similarity between two feature words. The optimized weight expression is:
[0067]
[0068] Among them, ω dx ,ω dy Respectively represent the initial weights of feature words x and y in text d, ω dy represents the target weight of feature word y in text d. The significance of formula (4) is that when the correlation between feature words x and y is high, if feature words x and y appear in the same text at the same time, the weight of one of the feature words can be reduced. In particular, when the correlation between feature words x and y is 1, selecting any feature word can largely represent the text information characteristics, so that attention can be focused on those feature words that cause text differences.
[0069] Step S1212: Obtain a weighted feature value for each feature word according to the target weight and the hash value of each feature word.
[0070] It should be noted that the hash value of each feature word is weighted according to the target weight of each feature word; specifically, when calculating each bit, if 1 is encountered, its weight value is added, and if 0 is encountered, its weight value is subtracted, so as to obtain the weighted feature value of each feature word.
[0071] Step S1213: Accumulate each weighted feature value in the current text data to obtain a combined feature vector.
[0072] Step S1214: perform dimensionality reduction processing on the merged feature vector to obtain a text signature corresponding to the current text data.
[0073] It should be noted that the merged feature vector is reduced to, for each bit, if it is greater than 0, the bit position is set to 1, otherwise it is set to 0, and the result is used as the signature of the text.
[0074] Step S130: Calculate the similarity value between each text data according to the text signature corresponding to each text data.
[0075] It should be noted that when calculating the similarity value between texts, the XOR operation is performed on different text signatures, and their signature values are compared bit by bit. If the value on the bit is different, it is recorded as 1, otherwise it is 0, and the number of 1s is the size of the Hamming distance. The larger the Hamming distance, the lower the similarity between the two texts, and vice versa.
[0076] Step S140: De-duplicate the text data in the initial data set according to the similarity value to obtain the filtered data set.
[0077] In this embodiment, each similarity value is compared with a preset threshold. If the similarity value is greater than or equal to the preset threshold, it means that the two text data are very similar, and one of them is deleted to obtain a filtered data set with redundant data removed.
[0078] In another embodiment, if Figure 4 As shown, quality assessment is performed on each text data in the preprocessed data set, specifically including the following steps:
[0079] Step S310: divide the current text data into several paragraphs to obtain multiple short texts.
[0080] Step S320: Obtain an evaluation result of each short text according to the syntactic structure level and the compactness level of each short text.
[0081] In one embodiment, before obtaining the evaluation result of each short text according to the syntactic structure level and the closeness level in each short text, the method further includes: segmenting the current short text to obtain a segmentation list corresponding to the current short text; obtaining the number of predicates and the number of agent-patient relations of the predicates in the current short text according to the segmentation list; and obtaining the syntactic structure level of the current short text according to the ratio of the relationship number to the number of predicates.
[0082] It should be noted that text quality is the main factor affecting information acquisition. High-quality text not only has clear, accurate and unambiguous semantics, but also has a smoother language experience. Therefore, the goal of this embodiment is to establish a universal text quality assessment method: first, determine the quality assessment criteria for sentences, and on this basis, analyze the syntactic structure and sequence tightness of the sentences based on abstract semantic representation to implement sentence assessment rules and algorithms, and then realize text quality assessment. Among them, the quality assessment criteria include: Criterion 1, the more complete the grammatical structure of the sentence, the higher the quality; Criterion 2, the stronger the tightness of the sequence of components in the sentence, the higher the quality.
[0083] In a sentence, the importance of the main part of the sentence (subject, predicate, and object) and the modifying components (attributive, adverbial, complement, etc.) are different, and the degree of influence of different components on the uncertainty of the expression content is also different; it can be further concluded that the different positions of the modifying components have different degrees of influence on the uncertainty of the expression content, and based on this, the second criterion is proposed.
[0084] In this embodiment, an abstract semantic representation graph (AMR graph) is constructed according to the segmentation list corresponding to the current short text, and the AMR graph includes the segmentation type corresponding to each segmentation and the relationship between the segmentations; wherein the segmentation type includes subject, predicate (predicate), object and adverbial; the agent relationship of the predicate represents the relationship between the subject and the predicate, and the patient relationship of the predicate represents the relationship between the predicate and the object. In a sentence, when the ratio of the number of agent-patient relationships of the predicate to the number of predicates is greater than 1.5, it indicates that the syntactic structure is good and its syntactic structure level is good; if the ratio is greater than 0.5 and less than or equal to 1.5, it indicates that the syntactic structure is general and the syntactic structure level is medium; if the ratio is less than or equal to 0.5, it indicates that the syntactic structure is poor and its syntactic structure level is poor.
[0085] In one embodiment, before obtaining the evaluation result of each short text according to the syntactic structure level and the closeness level in each short text, the method also includes: segmenting the current short text to obtain a segmentation list corresponding to the current short text; taking each segmentation as a node and calculating the closeness value of each node relationship; obtaining the sentence closeness value corresponding to the current short text according to the closeness value of each node relationship; and obtaining the closeness level corresponding to the current short text according to the sentence closeness value.
[0086] It should be noted that the quality of sentences is not only related to the syntactic structure, but also to the tightness of the sequence within the sentence, that is, different components have different effects on sentences. The stronger the sentence tightness, the greater the average sentence sequence value; the PEN-MAN tree representation of the AMR graph can fully reflect the hierarchical and sequential relationship of sentence fragments. According to the different hierarchical structures and node relationships of sentences, the following three rules are proposed: Rule 1: In AMR, different node relationships have different effects on sentence tightness; the relationship components are divided into four categories; the first category is the relationship between the main relationship in AMR, which represents the relationship between the node and the predicate, with a level value of x; the second category is the time, place and purpose explanation relationship, with a level value of y; the third category is the non-important relationship, which represents sentence modification, with a level value of z; the fourth category is the parallel or example relationship, with a level value of m; Rule 2: If the parent node of the current node relationship is not the root, the higher the influence of its parent node relationship on the sentence tightness, the more important the relationship node is; Rule 3: In the single sentence PENMAN tree form, the smaller the number of layers of the concept node, the greater the influence on the sentence tightness and the greater the weight.
[0087] In this embodiment, the calculation method of the closeness value of the current node relationship is as shown in formula (5):
[0088]
[0089] Among them, qr represents the closeness value of the current node relationship; q represents the current node relationship level value; qf represents its parent node framework value; N is the maximum number of layers in the AMR graph where the node is located; and n is the number of layers where the current node is located.
[0090] In this embodiment, according to the sentence closeness value, the calculation formula for obtaining the closeness level corresponding to the current short text is shown in formula (6):
[0091]
[0092] Among them, p represents the average compactness value; M is the number of all nodes, and ∑qr represents the sum of the compactness values of all nodes.
[0093] In one embodiment, the evaluation results of each short text include low, medium and high. The syntactic structure level and the compactness level of the short text are used as evaluation indicators to jointly determine the quality evaluation result of the short text. The specific evaluation criteria are shown in Table 1:
[0094] Table 1. Short text evaluation criteria
[0095] Syntactic structure level Tightness level Evaluation results Difference Low, Medium Low Difference high middle middle Low, Medium middle middle high high good Low, Medium middle good Medium, High high
[0096] Step S330: Weight the evaluation results of each short text to obtain the quality evaluation result of the current text data.
[0097] In this embodiment, the evaluation result of each short text is weighted by averaging, assigning different weights to different paragraphs, or assigning different weights according to the evaluation results, so as to obtain the quality evaluation result of the text data.
[0098] In one embodiment, the present application provides a text data processing device, which specifically includes the following embodiments:
[0099] Figure 5 The figure is a schematic diagram of the structure of a text data processing device provided in an embodiment of the present application; the text data processing device comprises:
[0100] A data filtering module, used to perform deduplication processing on the text data in the initial data set to obtain a filtered data set; wherein the initial data set includes a data set constructed based on Internet data;
[0101] A data preprocessing module, used for purifying the filtered data set according to a preprocessing configuration rule to obtain a preprocessed data set;
[0102] The data quality screening module is used to perform quality assessment on each text data in the preprocessed data set, and form a text training sample set with text data with high quality assessment results; it is also used to train a text classifier based on the text training sample set to obtain a trained target text classifier; it is also used to evaluate and classify input text based on the target text classifier to obtain a quality classification result of the input text.
[0103] It should be noted that if Figure 5 As shown, the text data processing device provided in this embodiment includes three core modules, specifically:
[0104] (1) Data filtering module: This module performs preliminary screening and filtering on the data crawled from the Internet, mainly using a language screening model to select Chinese data and a deduplication algorithm to remove duplicate data.
[0105] (2) Data preprocessing module: This module is responsible for preprocessing Chinese monolingual network data to form high-quality text data, that is, filtering out sensitive words, advertising content, and erroneous characters through manually formulated rules. It includes a rule library for storing customized data processing rules. The preprocessing controller calls the rule interpreter to purify the network data based on the processing rules.
[0106] (3) Data quality screening module: This module is used to screen high-quality data. During the preprocessing process, although some manual rules were used to remove explicit noise text from the dataset, there is still a large amount of low-quality text data in the remaining data that cannot be filtered out by manual rules. This module combines a large language model with a small model to filter the data quality and output a high-quality Chinese dataset, balancing accuracy and performance, and achieved good results.
[0107] In one embodiment, the processing flow of the text data processing device is as follows: Figure 6 As shown, the specific steps include:
[0108] Step 1: Obtain Internet data: Obtain Internet data through crawlers and build a basic data set as the input of the system.
[0109] Step 2: Data filtering: Perform preliminary filtering of the data, mainly to filter out Chinese data and deduplicate the data at the semantic level.
[0110] Step 2.1, Chinese data recognition: In order to improve performance, randomly sample n sentences, determine whether they are in Chinese, and filter out non-Chinese text data.
[0111] Step 2.2, data deduplication: Use the MinHashLSH algorithm to deduplicate data. It combines the minimum hash and local sensitive hash methods and is suitable for deduplication scenarios of large-scale data.
[0112] Step 3: Data preprocessing: Preprocess the Chinese monolingual network data based on the configuration rules to form high-quality text data.
[0113] Step 3.1, pre-processing rule configuration: formulate data pre-processing rules through the rule configuration module. This solution presets the following processing rules, where the numbers can be configured and changed, and other rules can also be added:
[0114] 1) If the average line length of the document is less than 15 characters or the total text length is less than 200 characters, it will be filtered out;
[0115] 2) Eliminate traditional Chinese characters and remove texts with less than 20% Chinese characters;
[0116] 3) The text is analyzed for the occurrence of harmful words from a predefined list, and any text with more than 1 occurrence of such words per line on average is classified as toxic content and removed from the list.
[0117] Step 3.2: Perform data preprocessing: call the interpreter to parse the rules and then complete the data processing.
[0118] Step 4: Data quality assessment: In this step, the large language model is used to screen some high-quality data. The screening results are output to the small model classifier for training the model.
[0119] Step 4.1, model training: Train the small model classifier based on the high-quality data filtered by the large model.
[0120] Step 4.1.1: The large language model scores the data quality: Based on the large language model's super language ability, the text data is scored. The score threshold T is preset. When the score is higher than T, the data will be sent to the small model classifier. In this step, due to the limited context of the large model (such as 8k), for longer texts, they will be divided into multiple text segments. For continuous text segments with a threshold greater than T in the same document, they will be automatically connected and output.
[0121] Step 4.1.2: Small model classifier training: Based on the high-quality data output in step 4.1.1, perform small model classifier training and output the trained model.
[0122] Step 4.1.3, output small model: export the trained small model for data quality identification.
[0123] Step 4.2: Determine data quality: Evaluate and classify the quality of the input text based on the small model classifier, and filter out data with good quality classification as the final output.
[0124] It can be seen that the text data processing device provided by this embodiment has at least the following beneficial effects:
[0125] 1. This embodiment proposes an automated data cleaning method based on a rule base and a rule interpreter. Through the introduction of the rule base and the development of the rule interpreter module, the user only needs to predefine the cleaning rules without coding implementation. The interpreter program can complete data cleaning based on these rules, greatly reducing the workload of data cleaning.
[0126] 2. This embodiment proposes a data quality judgment method that combines a large language model with a small text classification model. The large language model is used to filter and distill Chinese text data to obtain a high-quality Chinese data set. The high-quality data set is used to train the small model classifier so that the small model has a good classification effect. Finally, the small model is used to identify high-quality Chinese data, which ensures the accuracy of data judgment and avoids performance loss caused by the large model processing a large amount of data. While improving the judgment accuracy, it also improves the task execution efficiency.
[0127] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.
[0128] It should be noted that Figure 7 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0129] like Figure 7 As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 to the random access memory (RAM) 1003, such as executing the method described in the above embodiment. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, the ROM 1002 and the RAM 1003 are connected to each other through the bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0130] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read therefrom is installed into the storage section 1008 as needed.
[0131] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present application are executed.
[0132] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as a part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0133] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0134] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0135] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. A person skilled in the art can easily make corresponding changes or modifications based on the main concept and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.
Claims
1. A text data processing method, characterized in that: The method comprises: De-duplication processing is performed on the text data in the initial data set to obtain a filtered data set; wherein the initial data set includes a data set constructed based on Internet data; Purify the filtered data set according to the preprocessing configuration rule to obtain a preprocessed data set; Performing a quality assessment on each text data in the preprocessed data set, and forming a text training sample set with the text data having a high quality assessment result; Training a text classifier according to the text training sample set to obtain a trained target text classifier; The input text is evaluated and classified based on the target text classifier to obtain a quality classification result of the input text.
2. The method according to claim 1, characterized in that: The text data in the initial data set is deduplicated to obtain a filtered data set, including: Perform word segmentation on each text data in the initial data set to obtain multiple feature words corresponding to each text data; Obtaining a text signature corresponding to each text data according to the plurality of feature words corresponding to each text data; Calculate the similarity between each text data according to the text signature corresponding to each text data; According to the similarity value, the text data in the initial data set is deduplicated to obtain the filtered data set.
3. The method according to claim 2, characterized in that According to the multiple feature words corresponding to each text data, a text signature corresponding to each text data is obtained, including: According to the same hash function, the hash value of each feature word in the current text data is calculated; Calculate the target weight of each feature word according to the frequency of occurrence of each feature word in the current text data; According to the target weight and hash value of each feature word, the weighted feature value of each feature word is obtained; Accumulate each weighted feature value in the current text data to obtain a merged feature vector; The combined feature vector is subjected to dimensionality reduction processing to obtain a text signature corresponding to the current text data.
4. The method according to claim 3, characterized in that According to the frequency of occurrence of each feature word in the current text data, the target weight of each feature word is calculated, including: According to the number of samples in which the feature words x and the feature words y appear in the current sample data, the initial similarity of the feature words x and the feature words y is obtained; wherein the feature words x and the feature words y are respectively two different feature words from the multiple feature words in the current sample data; According to the number of times feature word x and feature word y appear in the current sample data, the similarity of feature word x and feature word y is weighted to obtain the target similarity of feature word x and feature word y; According to the frequency of occurrence of each feature word in the current sample data, the initial weight of each feature word is obtained; The initial weight of each feature word is optimized according to the target similarity between the feature word x and the feature word y to obtain the target weight of each feature word.
5. The method according to claim 1, characterized in that Performing a quality assessment on each text data in the preprocessed dataset includes: Divide the current text data into several paragraphs to obtain multiple short texts; According to the syntactic structure level and the compactness level of each short text, the evaluation result of each short text is obtained; The evaluation results of each short text are weighted to obtain the quality evaluation result of the current text data.
6. The method according to claim 5, characterized in that Before obtaining the evaluation result of each short text according to the syntactic structure level and the compactness level in each short text, the method further includes: Segment the current short text to obtain the segmentation list corresponding to the current short text; According to the word segmentation list, the number of predicates in the current short text and the number of agent-patient relationships of the predicates are obtained; The syntactic structure level of the current short text is obtained according to the ratio of the number of relations to the number of predicates.
7. The method according to claim 5, characterized in that Before obtaining the evaluation result of each short text according to the syntactic structure level and the compactness level in each short text, the method further includes: Segment the current short text to obtain the segmentation list corresponding to the current short text; Take each word as a node and calculate the closeness value of each node relationship; According to the closeness value of each node relationship, the sentence closeness value corresponding to the current short text is obtained; According to the sentence compactness value, the compactness level corresponding to the current short text is obtained.
8. A text data processing device, characterized in that: The device comprises: A data filtering module, used to perform deduplication processing on the text data in the initial data set to obtain a filtered data set; wherein the initial data set includes a data set constructed based on Internet data; A data preprocessing module, used for purifying the filtered data set according to the preprocessing configuration rules to obtain a preprocessed data set; The data quality screening module is used to perform quality assessment on each text data in the preprocessed data set, and form a text training sample set with text data with high quality assessment results; it is also used to train a text classifier based on the text training sample set to obtain a trained target text classifier; it is also used to evaluate and classify input text based on the target text classifier to obtain a quality classification result of the input text.
9. An electronic device, characterized in that: include: processor; as well as A memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to enable the electronic device to implement the text data processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the text data processing method according to any one of claims 1 to 7.