A Massive Text Clustering Method Based on Birch Clustering

By using birch clustering and locally sensitive hashing algorithms in massive text clustering methods, the problems of low and complex text clustering efficiency in the existing technology are solved, and efficient and accurate text clustering effect is achieved.

CN114266249BActive Publication Date: 2025-06-27NORTHEASTERN UNIV CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111586056.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-06-27
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

Existing text clustering methods are inefficient and complex when processing massive texts, especially when processing tens of millions of text data, the time consumes too much, and it is easy to lead to information loss and inaccurate clustering results.

Method used

A massive text clustering method based on birch clustering is adopted to calculate weights by word segmentation, removal of stop words, and TF-IDF, and a locally sensitive hash algorithm is used to reduce the dimensionality of text features, establish a hash cluster feature tree CF Tree, and realize the generation and update of the cluster center table.

Benefits of technology

It effectively reduces the dimension of text features, improves the speed and efficiency of clustering, reduces disk I/O operation and memory overhead, improves the accuracy of clustering results, and reduces the situation of poor clustering effects caused by artificial setting of hyperparameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266249B_ABST
    Figure CN114266249B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of machine learning, and proposes a method for clustering a large amount of text based on Birch clustering. For the processing of a large amount of text, an improved Birch clustering algorithm is used to build a hash clustering feature tree CF Tree. After the initial CF Tree is constructed, the threshold T is increased to reconstruct the CF tree to absorb more outlier sample values until the outlier disk no longer overflows; on the basis of word segmentation and stop word removal of the text, the weighted technology TF-IDF is used to calculate the keyword weights, and a part of the larger weight values is selected as the text features, and the dimensionality reduction of the features is carried out through the locality-sensitive hashing algorithm, removing the complex redundant information in the text, extracting the key information that can represent the text, reducing the dimensionality of the text features, and improving the clustering speed; a heuristic threshold increase mode is adopted, making the algorithm better adapt to the needs of different texts and reducing the probability of poor clustering results caused by artificially setting hyperparameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a method for clustering a large amount of text based on birch clustering. Background Art

[0002] With the continuous development of Internet technology, the data in the network is showing an explosive growth trend. A large number of unorganized and unclassified texts pose challenges for us to quickly obtain information in specific aspects. In order to improve the ability to obtain specified information, deduplication and classification of texts are essential. On the one hand, an excellent text clustering method can be used for deduplication to avoid the computer storing and accessing the same resources, wasting precious bandwidth and hard disk resources; on the other hand, it can be used for text classification to help us quickly and accurately locate specific information.

[0003] The clustering of texts mainly involves two aspects: one is to obtain the feature representation of texts; the other is to select an efficient clustering algorithm. At present, common text representation methods include part-of-speech representation, one-hot encoding, count vector representation, bag-of-words model vector representation, etc.; clustering algorithms are mainly common machine learning clustering algorithms and related variants. The Chinese patent "CN103226546A A suffix tree clustering method based on word segmentation and part-of-speech analysis" mainly extracts key information in texts by means of word segmentation, part-of-speech statistics, weight calculation, etc., reduces the dimension of the information to be clustered, and improves the accuracy of the clustering result. The Chinese patent "CN 106557777A A K-means document clustering method based on SimHash improvement" combines the characteristics of SimHash data compression and fast search for similar documents, improves K-means, and no longer calculates the text content, thus accelerating the clustering calculation time.

[0004] However, the Chinese patent "CN 103226546A A Suffix Tree Clustering Method Based on Word Segmentation and Part-of-Speech Analysis" uses word segmentation and part-of-speech analysis methods in the preprocessing stage of documents. Part-of-speech analysis only examines words of two parts of speech, namely nouns and verbs, and then conveys the processed string set to the subsequent suffix tree module. First, for some words in Chinese, there are often multiple meanings. Simply using part of speech to select target words as the main components of a document easily causes the loss of the main information of the document. Second, although the suffix tree itself is a document clustering algorithm with linear complexity, the process of string processing itself is a complex process. Especially when dealing with text data in the thousands or tens of thousands, the required time is unacceptable. In the solution described in the Chinese patent "CN 106557777A A Document Clustering Method Based on Improved SimHash K-means", when the K-means algorithm clusters documents, it calculates the distance between the fingerprint of the currently selected clustering document and each other document to be clustered. When the text data volume is very large, it will cause extremely large time consumption. Second, this solution does not stipulate the set threshold T, but the threshold T greatly affects the accuracy of the clustering algorithm. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention proposes a method for clustering massive texts based on birch clustering, including:

[0006] Step 1: Obtain text information and perform preprocessing;

[0007] Step 2: Perform dimensionality reduction processing on the text features of the words and phrases obtained from the preprocessing;

[0008] Step 3: Use the text SimHash value to create an initial clustering feature tree CF Tree and generate a clustering center table;

[0009] Step 4: Use the global clustering method to cluster all clusters and update the clustering center table to obtain the final classification result.

[0010] The preprocessing described in Step 1 includes word segmentation processing, removing stop words, and calculating weights.

[0011] Step 1 includes:

[0012] Step 1.1: Read each text in sequence. For the obtained English words, split them according to spaces, and for the obtained Chinese, use the Jieba word segmentation tool for word segmentation;

[0013] Step 1.2: Remove the stop words in the text data;

[0014] Step 1.3: Calculate the weight of each word in the text using the TF-IDF algorithm, and select the top M words with larger weights in each text as the features of the text.

[0015] The said step 2 includes:

[0016] Step 2.1: Use the Hash algorithm for each text feature to obtain an N-byte Hash value;

[0017] Step 2.2: Multiply the obtained Hash values by the corresponding weight values respectively to obtain the weighted values of the text features;

[0018] Step 2.3: Sum the weighted values of the text features of the same text bit by bit to obtain an N-byte string;

[0019] Step 2.4: Convert the N-byte string into a 0 / 1 string to obtain the SimHash value representing the text.

[0020] The said step 3 includes:

[0021] Step 3.1: Set the initial value of the threshold T and initialize a clustering feature tree CF Tree;

[0022] Step 3.2: Arbitrarily select a text to be clustered, take the SimHash value of the text as a sample point and put it into the initialized cluster, and take this value as the initialized cluster center;

[0023] Step 3.3: Cluster all texts to be clustered by calculating the Hamming distance between the SimHash value of the text and each cluster center;

[0024] Step 3.4: Judge whether the cluster with the newly added sample point exceeds the maximum number of sample points B that the cluster can accommodate; if it exceeds the maximum number of sample points B, delete the current cluster center, and select the two with the largest Hamming distance from all the sample points in the cluster as the new cluster centers;

[0025] Step 3.5: Calculate the Hamming distance between all sample points in the current cluster and the two cluster centers. For each sample point, add it to the cluster with the smaller Hamming distance and update the cluster center value;

[0026] Step 3.6: When all texts have been traversed, end the construction of the initial CF Tree and create a clustering center table according to the clustering results.

[0027] The said step 4 includes:

[0028] Step 4.1: Scan all clusters in the CF Tree, delete the clusters where the number of sample points within the cluster is much smaller than the average value, and recycle the sample points therein to the disk space as outliers. If the disk space for outliers does not overflow, execute Step 4.3; if the disk space overflows, execute Step 4.2;

[0029] Step 4.2: Increase the value of the threshold T, compress the CF Tree to absorb more sample points until the disk for outliers no longer overflows;

[0030] Step 4.3: Use the global clustering method to cluster all clusters and update the clustering center table to obtain the final classification result.

[0031] Furthermore, the specific description of Step 3.3 is as follows: Select the text to be clustered in sequence, calculate the Hamming distance between the SimHash value of the text and the centers of each cluster. If the Hamming distance between the text and the clustering center is less than or equal to the threshold T, then add the SimHash value of the text as a sample point to the cluster corresponding to the clustering center and update the cluster center value; if the Hamming distance between the text and the clustering center is greater than the threshold T, then use the SimHash of the text as a new cluster center and create a new cluster.

[0032] The beneficial effects of the present invention are:

[0033] The present invention proposes a method for clustering a large amount of text based on birch clustering. On the basis of word segmentation and stop word removal of the text, the weighted technology TF-IDF is used to calculate the keyword weights, and a part of the keywords with larger weight values are selected as keywords to reduce the dimension of the text features through the locality-sensitive hashing algorithm, removing the complex redundant information in the text, extracting the key information that can represent the text, reducing the dimension of the text features, and improving the clustering speed. Since the goal of the present invention is to process a large amount of text, the efficiency requirement for text clustering is relatively high. The improved birch clustering algorithm is used to establish the hashing clustering feature tree CF Tree, which can obtain a better clustering effect under the condition of linear scanning and save I / O and memory overhead. Since the clustering algorithm usually needs to set specific hyperparameters to determine the termination condition of clustering, when the inappropriate hyperparameters are set, it will affect the clustering effect. To further improve the clustering effect, the present invention also adopts a heuristic threshold promotion mode, making the algorithm better adapt to the needs of different texts and reducing the probability of poor clustering effect caused by artificially setting hyperparameters. Description of the Drawings

[0034] Figure 1 It is the flowchart of the method for clustering a large amount of text based on birch clustering in the present invention;

[0035] Figure 2 It is the main flowchart of the clustering process in the present invention. Detailed implementation manners

[0036] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation examples.

[0037] As Figure 2 shown is the main flowchart of the technical solution of the present invention. A method for clustering a large amount of text based on birch clustering mainly includes three parts: preprocessing of text information, feature dimensionality reduction, and feature clustering processing. The ultimate goal is to obtain a clustering center table of text data.

[0038] The preprocessing of text information includes three processes: word segmentation of the article, removal of stop words, and calculation of weights. When analyzing text-type data, whether it is English or Chinese, the first thing to do is to perform word segmentation. English words can be directly segmented according to spaces. For Chinese text, the jieba word segmentation tool is used in the present invention to perform word segmentation on text data.

[0039] Stop words include words that are widely used but are not helpful for searching themselves, such as personal pronouns like "you" and "I" in Chinese; and words that appear in a high proportion in the text but have little meaning themselves, such as modal particles, adverbs, prepositions, conjunctions, etc. Using this type of words will not only affect the clustering efficiency but also the accuracy of the clustering effect, and they should all be removed in the preprocessing stage.

[0040] In the present invention, TF-IDF is used to calculate weights. TF-IDF is a statistical method that can be used to evaluate the importance of a word for a text. The importance of a word is directly proportional to the number of times it appears in the text, but at the same time is inversely proportional to the frequency of its appearance in the corpus. Using TF-IDF can well measure the importance of a word for a document.

[0041] Feature dimensionality reduction refers to converting the words obtained from preprocessing into N-bit 0 / 1 strings using the locality-sensitive hashing algorithm to obtain text feature values, thereby achieving dimensionality reduction of text features.

[0042] The birch algorithm forms a clustering feature tree through clustering features (CF). In the feature clustering processing part, the present invention improves on the basis of the birch algorithm and uses an improved birch clustering algorithm to cluster the text. By improving the CF structure of the birch algorithm, the improved birch algorithm can process the feature data in the hash space; by adjusting the threshold T, the CF Tree can adaptively adjust the size of the cluster, improving the effect of the clustering algorithm.

[0043] For this reason, a method for clustering a large amount of text based on birch clustering proposed by the present invention, as Figure 1 shown, includes:

[0044] Step 1: Obtain text information and perform preprocessing; the preprocessing in Step 1 includes word segmentation, stop word removal, and weight calculation; specifically described as:

[0045] Step 1.1: Read each text in sequence. For the obtained English words, split them by spaces, and for the obtained Chinese, use the Jieba word segmentation tool for word segmentation.

[0046] Step 1.2: Remove stop words in the text data to reduce the computational amount of calculating the word segmentation weights by the computer.

[0047] Step 1.3: Use the TF-IDF algorithm to calculate the weight of each word in the text, and take the top M words with larger weights in each text as the features of the text; in this embodiment, select the first 20 words of each text as the features of the text, and obtain a weight matrix R of M * 20 (M represents the number of texts).

[0048] Step 2: Perform dimensionality reduction processing on the text features obtained by preprocessing; including:

[0049] Step 2.1: Use the Hash algorithm for each text feature to obtain a Hash value of N bytes, ensuring that each feature is unique here.

[0050] Step 2.2: Multiply the obtained Hash values by the corresponding weight values respectively to obtain the weighted values of the text features.

[0051] Step 2.3: Sum the weighted values of the text features of the same text bit by bit to obtain a string of N bytes.

[0052] Step 2.4: Convert the string of N bytes into a 0 / 1 string to obtain the SimHash value representing the text.

[0053] Step 3: Create an initial clustering feature tree CF Tree using the text SimHash value and generate a clustering center table; including:

[0054] Step 3.1: Set the initial value of the threshold T, set T to 3, and initialize a clustering feature tree CF Tree.

[0055] Step 3.2: Arbitrarily select a text to be clustered, put the SimHash value of the text as a sample point into the initialized cluster, and use this value as the initialized cluster center.

[0056] Step 3.3: Cluster all the texts to be clustered by calculating the Hamming distance between the SimHash value of the text and the cluster centers; specifically described as: sequentially select the texts to be clustered, calculate the Hamming distance between the SimHash value of the text and the cluster centers. If the Hamming distance between the text and the cluster center is less than or equal to the threshold T, then take the SimHash value of the text as a sample point and add it to the cluster corresponding to the cluster center, and update the cluster center value; if the Hamming distance between the text and the cluster center is greater than the threshold T, then take the SimHash of the text as a new cluster center and create a new cluster.

[0057] Step 3.4: Determine whether the number of sample points in the cluster where the newly added sample point is located exceeds the maximum number of sample points B that the cluster can accommodate; if it exceeds the maximum number of sample points B, delete the current cluster center, and select the two with the largest Hamming distance from all the sample points in the cluster as the new cluster centers.

[0058] Step 3.5: Calculate the Hamming distance between all the sample points in the current cluster and the two cluster centers. For each sample point, add it to the cluster with the smaller Hamming distance and update the cluster center value.

[0059] Step 3.6: When all the texts have been traversed once, end the construction of the initial CF Tree, and create a cluster center table according to the clustering results.

[0060] After the creation of the CF Tree, in order to absorb more leaf node entries, the CF Tree can be compressed by increasing the threshold T. For the initial threshold we selected, it may not conform to the actual situation of the text, resulting in many texts not being divided into clusters. After the initial CF Tree is constructed, the CF Tree is reconstructed by increasing the threshold T to absorb more leaf node entries until all the outliers are inserted into the tree.

[0061] Step 4: To ensure the clustering effect, use the global clustering method to cluster all the clusters and update the cluster center table to obtain the final classification result; including:

[0062] Step 4.1: Scan all the clusters in the CF Tree, delete the clusters with the number of sample points in the cluster much less than the average value, and recycle the sample points in them as outliers to the disk space. If the disk space for outliers does not overflow, execute Step 4.3; if the disk space overflows, execute Step 4.2.

[0063] Step 4.2: Increase the value of the threshold T to compress the CF Tree to absorb more sample points until the disk for outliers no longer overflows.

[0064] Step 4.3: Use the global clustering method to cluster all the clusters and update the cluster center table to obtain the final classification result.

[0065] To verify the effectiveness of the method of the present invention, the present invention is verified from two aspects: the text feature extraction effect and the clustering effect. The text feature evaluation indicators are mainly the number of feature words extracted and the feature dimension. The clustering effect includes the execution time of the algorithm and the silhouette coefficient.

[0066] Using the clustering method proposed by the present invention, in the text preprocessing stage, on the basis of covering more key information, it can effectively reduce the dimension of text features and is beneficial to improving the clustering speed. The comparison between the features of the traditional text clustering method and the features of the clustering method of the present invention is shown in Table 1:

[0067] Table 1 Comparison table of the part-of-speech suffix tree text clustering method and the birch text clustering method

[0068]

[0069] It can be seen from Table 1 that compared with the traditional part-of-speech suffix tree algorithm, birch can stably extract 20 feature words from the text as text features, and the feature dimension is always 64bit. The features of string type reduce the complexity of subsequent processing.

[0070] The present invention can cluster massive texts quickly and efficiently. On the dataset composed of 87,054 texts, the comparison of the clustering quality and clustering speed of the clustering method proposed by the present invention is shown in Table 2:

[0071] Table 2 Comparison result table of the k-means text clustering method and the bitch text clustering algorithm

[0072] K-means text clustering birch text clustering Execution time (s) 200.9 85.2 Time complexity O(l*n*k*m) O(n) Silhouette coefficient 0.62 0.71

[0073] In Table 2, O represents the time complexity of the algorithm, l is the number of iterations, n is the number of samples, k represents the number of clusters, and m represents the sample point dimension. It can be seen from Table 2 that the silhouette coefficient of the birch clustering algorithm is better than that of the K-means clustering algorithm, and the execution time is more than twice as fast. Generally speaking, the bitch algorithm has greatly improved compared with the K-means algorithm in processing massive texts.

[0074] In summary, the present invention proposes a method for clustering a large amount of text based on the birch clustering algorithm. Aiming at the problems of diverse text data features and high processing difficulty, the method uses the locality-sensitive hashing algorithm to extract the key information dimensions in the document, reducing the complexity of clustering; the improved birch algorithm is used to apply the pattern tree algorithm in the Euclidean space to the hash space, and a CF representation method for the hash space is proposed, solving the problem that the birch algorithm can only process numerical data; the use of the clustering feature CF and the CF tree reduces disk I / O operations and memory overhead, and the linear scanning mechanism enables a high-quality clustering effect to be generated in one traversal, overcoming the problem of time-consuming and laborious repeated calculation of the clustering center by traditional algorithms through multiple iterations.

Claims

1. A method for clustering a large amount of text based on birch clustering, characterized in that, Including: Step 1: Obtain text information and perform preprocessing; Step 2: Perform dimensionality reduction processing on the text features obtained from the preprocessing; Step 2.1: Use the Hash algorithm for each text feature to obtain an N-byte Hash value; Step 2.2: Multiply the obtained Hash values by the corresponding weight values respectively to obtain the weighted values of the text features; Step 2.3: Sum the weighted values of the text features of the same text bit by bit to obtain an N-byte string; Step 2.4: Convert the N-byte string into a 0 / 1 string to obtain the SimHash value representing the text; Step 3: Use the text SimHash value to create an initial clustering feature tree CF Tree and generate a clustering center table; Step 3.1: Set the initial value of the threshold T and initialize a clustering feature tree CF Tree; Step 3.2: Arbitrarily select a text to be clustered, take the SimHash value of this text as a sample point and put it into the initialized cluster, and use this value as the initialized cluster center; Step 3.3: Cluster all texts to be clustered by calculating the Hamming distance between the SimHash value of this text and the cluster centers of each cluster; Step 3.4: Determine whether the cluster where the newly added sample point is located exceeds the maximum number of sample points B that the cluster can accommodate; If it exceeds the maximum number of sample points B, delete the current cluster center, and select the two with the largest Hamming distance from all the sample points in the cluster as the new cluster centers; Step 3.5: Calculate the Hamming distance between all sample points in the current cluster and the two cluster centers. For each sample point, add it to the cluster with the smaller Hamming distance and update the cluster center value; Step 3.6: When all texts have been traversed once, end the construction of the initial CF Tree and create a clustering center table according to the clustering results; Step 4: Use the global clustering method to cluster all clusters and update the clustering center table to obtain the final classification result.

2. The mass text clustering method based on birch clustering according to claim 1, wherein The preprocessing described in Step 1 includes word segmentation processing, stop word removal, and weight calculation.

3. A method for clustering a large amount of text based on birch clustering according to claim 2, characterized in that, Step 1 includes: Step 1.1: Read each text in turn. For the obtained English words, split them according to spaces, and for the obtained Chinese, use the Jieba word segmentation tool for word segmentation; Step 1.2: Remove stop words from the text data; Step 1.3: Use the TF-IDF algorithm to calculate the weight of each word in the text, and take the first M words with larger weights in each text as the features of this text.

4. A method for clustering a large amount of text based on birch clustering according to claim 1, characterized in that, Step 4 includes: Step 4.1: Scan all clusters in the CF Tree, delete the clusters with the number of sample points in the cluster much smaller than the average value, and recycle the sample points among them to the disk space. If the disk space for outliers does not overflow, execute Step 4.3; if the disk space overflows, execute Step 4.2; Step 4.2: Increase the value of the threshold T, compress the CF Tree to absorb more sample points until the disk for outliers no longer overflows; Step 4.3: Use the global clustering method to cluster all clusters and update the clustering center table to obtain the final classification result.

5. A method for clustering a large amount of text based on birch clustering according to claim 1, characterized in that Step 3.3 is specifically described as follows: successively select the text to be clustered, calculate the Hamming distance between the SimHash value of the text and the center of each cluster. If the Hamming distance between the text and the cluster center is less than or equal to the threshold T, then take the SimHash value of the text as a sample point and add it to the cluster corresponding to the cluster center, and update the cluster center value; if the Hamming distance between the text and the cluster center is greater than the threshold T, then take the SimHash of the text as a new cluster center and create a new cluster.

Citation Information

Patent Citations

  • Suffix tree clustering method on basis of word segmentation and part-of-speech analysis

    CN103226546A

  • Improved Kmeans clustering method based on SimHash

    CN106557777A

  • Keyword calculation method based on document clustering

    CN105159998A

  • A method and a system for improving a Simhash algorithm in text deduplication

    CN109948125A