Text classification method based on compression length

Through the text classification method based on compressed length, CPC compressors are used to classify text data in each category, which solves the problem of low efficiency in large-scale data text classification in the prior art, and achieves more efficient and accurate text classification.

CN119988625APending Publication Date: 2025-05-13SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510016976.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing text classification methods are less efficient when processing large-scale data, and compressor-based methods require a lot of labor and time costs. Data compression may lead to a decrease in classification accuracy, and KNN algorithms are less efficient when processing large-scale data.

Method used

Using a text classification method based on compressed length, by obtaining training text data, CPC compressors are trained on labeled data of each category, and CPC classification models corresponding to n categories are obtained. The samples to be predicted are input to these classification models, the average length of each category is calculated, and the category corresponding to the smallest average length is taken as the classification result.

Benefits of technology

It improves the efficiency and accuracy of text classification, reduces the time complexity to O(n), maintains a low growth rate of time overhead during large-scale data processing, and avoids information loss caused by data compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988625A_ABST
    Figure CN119988625A_ABST
Patent Text Reader

Abstract

The invention relates to a text classification method based on compression length, and the method comprises the following steps: S1, obtaining training text data which comprises a plurality of types of labeled data; s2, CPC compressors are trained for the data with the labels of each category, CPC classification models corresponding to n categories are obtained, n represents the total number of the categories, and a to-be-predicted sample is obtained; s3, respectively inputting a to-be-predicted sample into CPC classification models respectively corresponding to the n categories to obtain CPC lengths respectively corresponding to the n categories, and calculating an average length corresponding to each category to obtain n average lengths; and S4, taking the minimum value of the n average lengths, and taking the category corresponding to the minimum average length as a classification result. Compared with the prior art, the method has the advantages of improving the text classification efficiency of large-scale data, guaranteeing the classification accuracy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification, and in particular to a text classification method based on compression length. Background Art

[0002] With the development of artificial intelligence, text classification tasks have gradually become one of the key technologies in various fields. Through text classification, we can effectively organize and manage massive amounts of text information, thereby achieving more efficient information utilization and processing. In the business field, text classification is widely used in marketing, customer service, public opinion monitoring, etc. For example, using text classification technology, products can be classified according to user comments and needs, so as to better understand market demand and user preferences; for example, in the financial field, text classification can be used to analyze news and announcements, timely capture market changes and risk factors, and provide support for investment decisions. In the field of medical health, text classification can be used for medical record classification, disease diagnosis, medical literature analysis, etc., to provide a basis for medical decision-making. At the same time, in the field of scientific research, text classification also has a wide range of applications, such as paper classification, literature analysis, etc., which helps to accelerate the progress of scientific research. In short, with the continuous advancement of artificial intelligence technology, the application scope of text classification tasks will be wider, bringing more convenience and value to all walks of life.

[0003] Existing Technology 1: AI is developing rapidly in text classification tasks, and many mature and efficient models have emerged. These models use technologies such as deep learning to continuously improve the performance and efficiency of text classification.

[0004] The following are some of the most popular and influential models:

[0005] 1. Recurrent Neural Network (RNN): RNN is a classic sequence model that is suitable for processing data with sequential or temporal relationships, such as natural language text. In text classification, RNN can process each word step by step along the text sequence and pass the previous information to the next time step. However, RNN has the problem of gradient disappearance or gradient explosion, which limits its ability to model long texts.

[0006] 2. Long Short-Term Memory Network (LSTM): LSTM was proposed to solve the gradient vanishing problem of RNN. LSTM introduces a gating mechanism that can effectively capture long-term dependencies. In text classification, LSTM can better process long text sequences and retain more text information, so it performs well in some tasks that require considering contextual information.

[0007] 3. Convolutional Neural Network (CNN): Although CNN was originally designed for image processing, it also performs well in text classification. CNN extracts local features through convolution operations, merges feature information through pooling operations, and finally uses a fully connected layer for classification. In text classification, CNN usually uses one-dimensional convolution operations to process local features of text, such as words and phrases, rather than two-dimensional features in images.

[0008] 4. Transformer model: The Transformer model is a major breakthrough in recent years. Its core is the self-attention mechanism. Through the self-attention mechanism, the Transformer can simultaneously take into account all position information in the input sequence, thereby better capturing long-distance dependencies in the text. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained model based on the Transformer model. It can achieve very good performance through large-scale unsupervised learning pre-training and then fine-tuning on specific tasks.

[0009] 5. BERT (Bidirectional Encoder Representations from Transformers): BERT is a pre-training model proposed by Google, based on the Transformer structure. Through large-scale unsupervised pre-training, BERT can learn rich language representations, including word meaning, syntax, and context information. In text classification tasks, the BERT model can be adapted to specific fields or tasks through fine-tuning, that is, performing a small amount of supervised learning on the target task. BERT has achieved significant performance improvements on multiple text classification tasks.

[0010] 6. GPT (Generative Pre-trained Transformer) series: GPT is a series of Transformer-based pre-trained language models proposed by OpenAI. These models learn rich language representations through large-scale unsupervised pre-training. In particular, GPT-3, with 175 billion parameters, is the largest natural language processing model to date. Although the GPT series of models is more prominent in generation tasks, it can also be used for text classification, especially when faced with complex or multi-category classification tasks, the GPT series of models can provide good results.

[0011] Prior art 2: By defining appropriate regular expression patterns, we can quickly and accurately identify these sensitive data in text data. Regular expression-based text classification tasks are a common, simple but effective method for finding specific patterns in text data. These patterns can represent sensitive data, such as phone numbers, email addresses, credit card numbers, etc. Information, thereby enhancing data security and privacy protection. For example, by matching the regular expression of the phone number, we can easily find and mark the phone number in the text, and then perform further processing, such as encryption or desensitization, to ensure that the data is not illegally obtained or abused. This method is not only simple and easy to use, but also very effective for some specific sensitive data identification tasks, becoming one of the important tools for data security and privacy protection. The regular expression-based text classification method has the following advantages: simple and intuitive: regular expressions are easy to understand and implement, without the need for complex model training processes; flexibility: a variety of different matching rules can be formulated as needed to adapt to different text classification tasks; high efficiency: regular expressions match quickly and are suitable for processing large-scale text data.

[0012] Prior art three: Text classification method based on compressor: The text classification method based on compressor and KNN (K-nearest neighbor) algorithm is a technology that combines data compression and nearest neighbor classification, which can efficiently and quickly classify text data of different categories. The main idea of ​​this method is based on data compression and similarity. The basic principle is divided into the following steps: 1) Data preparation. In order to smoothly carry out the classification task, it is necessary to prepare corpora of different categories in advance, such as email, mobile phone number, ipv4, etc., and a small amount of data is required for each category. 2) Data compression. Data compression technology is used to process text data. This will help reduce the dimension of the data and remove some irrelevant information. Data compression is performed on the text we want to predict and the corpus text to obtain the compressed length. 3) The obtained compressed length and NCD (Normalized Compression Distance) formula are used to calculate the information distance between the predicted text and the text of each category of the corpus, and the information distance matrix can be obtained. 4) Finally, the KNN algorithm and the information distance matrix are used to obtain the K categories that are closest to the predicted text and the corpus text. Among the K categories, the category with the largest number of categories belongs to the predicted text, and the prediction is completed. This approach is simple and easy to understand, and it helps reduce computational costs and improve classification performance by combining compression to reduce dimensionality.

[0013] Existing text classification methods use methods based on deep learning models, regular expression-based text methods, and compressor-based methods. The compressor-based method requires the preparation of corpora of different categories in advance, which may require a lot of manpower and time costs; data compression may lose some text information, resulting in a decrease in classification accuracy; the KNN algorithm may be less efficient when processing large-scale data. Summary of the invention

[0014] The purpose of the present invention is to provide a text classification method based on compression length in order to improve the efficiency of text classification of large-scale data while ensuring the classification accuracy.

[0015] The purpose of the present invention can be achieved by the following technical solutions:

[0016] A text classification method based on compression length, the method comprising the following steps:

[0017] S1. Obtain training text data, where the training text data includes labeled data of multiple categories;

[0018] S2. For each category of labeled data, CPC compressors are trained to obtain CPC classification models corresponding to n categories, where n represents the total number of categories, and samples to be predicted are obtained;

[0019] S3, input the samples to be predicted into CPC classification models corresponding to n categories respectively, obtain CPC lengths corresponding to n categories respectively, calculate the average length corresponding to each category, and obtain n average lengths;

[0020] S4. Take the minimum value of n average lengths, and the category corresponding to the minimum average length is the classification result.

[0021] Furthermore, the specific steps of training CPC compressors for labeled data of each category to obtain CPC classification models corresponding to n categories are:

[0022] The data of the i-th category in the training text data is divided into CPC parts on average, where i = 1, 2, 3...n. Each part of the data is input into the ZSTD compressor to obtain CPC dictionaries, and CPC compressors corresponding to the i-th category are obtained. The total number of CPC classification models corresponding to n categories is CPC*n.

[0023] Furthermore, the samples to be predicted are input into the CPC classification models corresponding to the n categories respectively, and the specific steps of obtaining the CPC lengths corresponding to the n categories are:

[0024] The samples to be predicted are respectively input into the CPC classification models corresponding to the i-th category, where i=1, 2, 3...n. The CPC classification models corresponding to the i-th category output the CPC lengths corresponding to the i-th category. The total number of CPC lengths corresponding to n categories is CPC*n.

[0025] Furthermore, the specific steps for calculating the average length corresponding to each category are:

[0026] Calculate the average of the CPC lengths corresponding to the i-th category to obtain the average length corresponding to the i-th category. The total number of average lengths is n.

[0027] Furthermore, the calculation process of the average length corresponding to the i-th category is:

[0028] Add the CPC lengths corresponding to the i-th category, divide the sum by CPC, and the quotient is the average length corresponding to the i-th category.

[0029] Furthermore, the specific steps of inputting each data into the ZSTD compressor to obtain CPC dictionaries are:

[0030] Filter k data from the j-th data, divide the k data into CPC parts, and input them into the ZSTD compressor, where j = 1, 2, 3, ... CPC, to obtain the dictionary corresponding to the j-th data, and for the ith category, obtain the corresponding CPC dictionaries.

[0031] Furthermore, k data are obtained by random screening.

[0032] Further, a dictionary is data used to seed the compressor state.

[0033] Furthermore, CPC is an integer greater than or equal to 1.

[0034] Furthermore, CPC is an integer greater than or equal to 2 and less than or equal to 5.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The present invention constructs one or more compressors for each class, that is, the same type of input data is divided into one or more blocks, a compressor is trained for each block, the average value of the compression ratio is taken to classify the text, and the category corresponding to the compressor with the highest compression ratio is selected as the classification result of the text, thereby improving the accuracy of text classification. At the same time, the time complexity of the present invention is O(n). When the number of input data to be predicted is large, the growth rate of time overhead is slow, which can improve the text classification efficiency of large-scale data. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1It is a schematic diagram of the ZSTD compressor flow of the present invention;

[0038] Figure 2 It is a flow chart of the classification training compressor of the present invention;

[0039] Figure 3 Pseudo code and schematic diagram for the actual reasoning of the present invention. DETAILED DESCRIPTION

[0040] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0041] In view of the fact that there is a large amount of text data that needs to be sorted, classified and archived, and it is hoped that the problem can be processed relatively quickly without the need to prepare a large amount of training data sets and long-term model training like neural networks, the present invention proposes a text classification method based on compression length, which includes the following steps:

[0042] S1. Obtain training text data, where the training text data includes labeled data of multiple categories;

[0043] S2. For each category of labeled data, CPC compressors are trained to obtain CPC classification models corresponding to n categories, where n represents the total number of categories, and samples to be predicted are obtained;

[0044] S3, input the samples to be predicted into CPC classification models corresponding to n categories respectively, obtain CPC lengths corresponding to n categories respectively, calculate the average length corresponding to each category, and obtain n average lengths;

[0045] S4. Take the minimum value of n average lengths, and the category corresponding to the minimum average length is the classification result.

[0046] For the classification of text data, the basic framework adopted by the present invention is based on the ZSTD compressor. Figure 1 shown.

[0047] A compressor (such as zstd) can generate a compression dictionary based on the data. A compression dictionary is essentially data used to seed the state of the compressor so that it can achieve better compression. The principle is that if you want to compress a large amount of similar data (such as JSON documents or any data that shares a similar structure), you can find common patterns in multiple objects and then use these common patterns during compression and decompression operations to achieve better compression rates. Dictionary compression is generally only suitable for small inputs, that is, data that does not exceed a few thousand bytes. Intuitively, this means that if the compressor is "trained" for a specific category of data, it can better compress this specific category of data. A different compressor can be trained for each category of data. Then given an input text, I can compare all the compressors. The category of the compressor with the highest compression rate is likely to be the subject of the input text. This set of compressors is the text classification model used in the present invention.

[0048] The CPC of the present invention is an integer greater than or equal to 1. When the CPC is 1, the steps of the present invention are:

[0049] Step 1 (prepare data): First prepare the training data. The training data contains labeled data of multiple categories, denoted as [text, class]. Package the data of each category together for easy training. ct_i represents all data with label i, then all data of class_1 can be represented as ct_1, all data of class_2 can be represented as ct_2,..., all data of class_n can be represented as ct_n.

[0050] Step 2 (training): Train each category separately. The result of the training is that a compression dictionary will be generated for each category of data separately. In this way, there will be n compressors for n categories of data. The steps are as follows:

[0051] (1) Setting the parameter k means taking k data from each category for training.

[0052] (2) By Figure 2 As shown, k data of category 1, i.e. ct_1, are randomly selected and input into the ZSTD compressor to obtain the corresponding dictionary 1. Then the compressor cmp_1 of class1 data is ready. Similarly, n compressors corresponding to n categories can be obtained. The compressor set can be expressed as [cmp_1, cmp_2, ······cmp_n].

[0053] Step 3 (Inference): There is a text, and its category needs to be predicted. For each compression dictionary, compress the input content and get the length. If there are n trained compressors, then input the text into these n compressors corresponding to different categories. The compressor with the best compression rate that can compress the text to the shortest is the category of the text. Figure 3 As shown, the following are the reasoning steps:

[0054] (1) Obtain the sample to be predicted, sample1.

[0055] (2) Obtain n trained compressors [cmp_1, cmp_2, ······cmp_n].

[0056] (3) Input sample1 into compressor cmp_1 to obtain the compressed file, and calculate the compressed bit stream length len1. Repeat this step to obtain the compressed bit stream lengths [len_1, len_2, ..., len_n] of n compressors.

[0057] (4) Take the minimum value in [len_1, len_2, ······, len_n]. If it is len_a (a=1, 2, ······, n), then the predicted category of sample1 is class_a.

[0058] (5) Output class_a.

[0059] When CPC is an integer greater than 1, the present invention does not build a compressor for each class, but builds several compressors. For example, a given class text is divided into 3 blocks, and 1 compressor is trained for each block. During inference, the average value of the compression ratio is taken (a voting method can also be used), which can stabilize the inference. The number of compressors per class is called CPC, and a small value between 2 and 5 is sufficient. The method in steps 2 and 3 is equivalent to CPC = 1. If the value of CPC is set to a value other than 1, assuming k = b (b = 2, 3, 4, 5), then change (2) in step 2 (training) to the following:

[0060] (2) Divide the k data taken from ct_1 into CPC parts on average, and input each data into the ZSTD compressor to obtain CPC dictionaries cmp_1_1, cmp_1_2, ..., cmp_1_CPC. Repeat this operation to obtain all the trained compressors {[cmp_1_1, cmp_1_2, ..., cmp_1_CPC], [cmp_2_1, cmp_2_2, ..., cmp_2_CPC], ... [cmp_n_1, cmp_n_2, ..., cmp_n_CPC]}, a total of CPC*n trained compressors.

[0061] Then (2) in step 3 also includes:

[0062] Get CPC*n trained compressors {[cmp_1_1, cmp_1_2,..., cmp_1_CPC], [cmp_2_1, cmp_2_2,..., cmp_2_CPC],... [cmp_n_1, cmp_n_2,..., cmp_n_CPC]}.

[0063] Change (3) and (4) in step 3 to the following:

[0064] (2) Input sample1 into the compressor [cmp_1_1, cmp_1_2, ..., cmp_1_CPC] to obtain CPC lengths and calculate the average length aveLen_1. Repeat this step to obtain n lengths [aveLen_1, aveLen_2, ..., aveLen_n].

[0065] Take the minimum value among [aveLen_1, aveLen_2, ······, aveLen_n]. If it is aveLen_a (a=1, 2, ······, n), then the predicted category of sample1 is class_a.

[0066] The key innovations of the present invention are:

[0067] 1. Text classification algorithm based on ZSTD compressor: This algorithm uses ZSTD compressor and its generated compression dictionary to achieve text classification.

[0068] 2. Training data preparation: It is necessary to prepare labeled data containing multiple categories, and the data of each category is packaged together for easy training.

[0069] 3. Train each category separately: Train the data of each category separately and generate a compressed dictionary for each category for use in subsequent reasoning.

[0070] 4. Reasoning process: For the text to be classified, it is input into compressors of various categories, and the category corresponding to the compressor with the highest compression rate is selected as the classification result of the text.

[0071] 5. Improved algorithm: An improved algorithm is introduced, that is, multiple compressors are constructed instead of one compressor for each category, which further improves the stability and accuracy of classification. The present invention divides the text of each category into multiple blocks and trains a compressor for each block, which improves the accuracy of text classification.

[0072] The beneficial effects of the present invention are:

[0073] 1. Compared with the prior art, the present invention has the advantages of fast training, fast reasoning, and less training sample requirements. Traditional deep neural networks require a large number of training samples, have many hyperparameters that need to be adjusted, and training is very time-consuming. The present invention is an excellent lightweight alternative to traditional deep neural network methods.

[0074] 2. Compared with the regular expression of the prior art, the design concept of the present invention is easier to understand, has a higher accuracy rate and a lower false alarm rate.

[0075] 3. Compared with the prior art three, the time complexity of the present invention is lower. The time complexity of the prior art three is O(n^2), and the time complexity of the present invention is O(n). As the number of samples increases, the time overhead of the prior art three will increase at a quadratic level, while the time overhead of the present invention only increases linearly. This is mainly because the present invention only needs to train multiple compressors in advance when predicting, while the scheme three needs to calculate the information distance between each pair.

[0076] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A text classification method based on compression length, characterized in that: The method comprises the following steps: S1. Obtain training text data, where the training text data includes labeled data of multiple categories; S2. For each category of labeled data, CPC compressors are trained to obtain CPC classification models corresponding to n categories, where n represents the total number of categories, and samples to be predicted are obtained; S3, input the samples to be predicted into CPC classification models corresponding to n categories respectively, obtain CPC lengths corresponding to n categories respectively, calculate the average length corresponding to each category, and obtain n average lengths; S4. Take the minimum value of n average lengths, and the category corresponding to the minimum average length is the classification result.

2. A text classification method based on compression length according to claim 1, characterized in that: The specific steps of training CPC compressors for labeled data of each category to obtain CPC classification models corresponding to n categories are: The data of the i-th category in the training text data is divided into CPC parts on average, where i = 1, 2, 3...n. Each part of the data is input into the ZSTD compressor to obtain CPC dictionaries, and CPC compressors corresponding to the i-th category are obtained. The total number of CPC classification models corresponding to n categories is CPC*n.

3. A text classification method based on compression length according to claim 2, characterized in that: The specific steps of inputting the samples to be predicted into the CPC classification models corresponding to n categories respectively and obtaining the CPC lengths corresponding to n categories respectively are as follows: The samples to be predicted are respectively input into the CPC classification models corresponding to the i-th category, where i=1, 2, 3...n. The CPC classification models corresponding to the i-th category output the CPC lengths corresponding to the i-th category. The total number of CPC lengths corresponding to n categories is CPC*n.

4. A text classification method based on compression length according to claim 3, characterized in that: The specific steps to calculate the average length corresponding to each category are: Calculate the average of the CPC lengths corresponding to the i-th category to obtain the average length corresponding to the i-th category. The total number of average lengths is n.

5. A text classification method based on compression length according to claim 4, characterized in that: The calculation process of the average length corresponding to the i-th category is: Add the CPC lengths corresponding to the i-th category, divide the sum by CPC, and the quotient is the average length corresponding to the i-th category.

6. A text classification method based on compression length according to claim 5, characterized in that: The specific steps for inputting each piece of data into the ZSTD compressor to obtain CPC dictionaries are: Filter k data from the j-th data, divide the k data into CPC parts, and input them into the ZSTD compressor, where j = 1, 2, 3, ... CPC, to obtain the dictionary corresponding to the j-th data, and for the ith category, obtain the corresponding CPC dictionaries.

7. A text classification method based on compression length according to claim 6, characterized in that: k data are obtained by random screening.

8. The text classification method based on compression length according to claim 1, characterized in that: The dictionary is the data used to seed the compressor state.

9. The text classification method based on compression length according to claim 1, characterized in that: CPC is an integer greater than or equal to 1.

10. A text classification method based on compression length according to claim 9, characterized in that: CPC is an integer greater than or equal to 2 and less than or equal to 5.