Data Clustering Method, Apparatus, Electronic Device, and Readable Storage Medium

By performing initial clustering of text data and reducing the dimensionality of high-dimensional feature vectors into low-dimensional feature vectors, and then performing re-clustering, the problem of low clustering efficiency in the existing technology is solved, and efficient and high-precision clustering effect is achieved.

CN115146692BActive Publication Date: 2025-07-01QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110352960.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-07-01
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

In the prior art, clustering efficiency is low, especially in scenarios where large data volumes are processed, resulting in large calculation volume and low efficiency.

Method used

By performing initial clustering of text data, the semantic features in the high-dimensional feature vector are replaced with cluster-like labels according to the initial clustering results, forming low-dimensional feature vectors, and then clustering again, reducing the calculation amount and improving efficiency.

Benefits of technology

This method effectively reduces the amount of clustering calculations, improves clustering efficiency, and ensures the accuracy of clustering effects, taking into account both efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115146692B_ABST
    Figure CN115146692B_ABST
Patent Text Reader

Abstract

The present application provides a data clustering method, apparatus, electronic device, and readable storage medium, relating to the technical field of data mining. After initially clustering text data, the method replaces semantic features of multiple dimensions in a high-dimensional feature vector with a cluster label of one dimension according to the initial clustering result, so as to achieve data dimensionality reduction. Then, the data after dimensionality reduction is clustered again, thereby effectively reducing the amount of calculation and improving the clustering efficiency during the process of re-clustering. Moreover, the clustering effect can be ensured through re-clustering. This implementation method can balance both clustering efficiency and clustering accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data mining technology. Specifically, it relates to a data clustering method, apparatus, electronic device, and readable storage medium. Background Art

[0002] Clustering refers to the process of dividing a set of physical or abstract objects into multiple classes composed of similar objects, which has wide applications in fields such as image analysis and text retrieval.

[0003] Currently, clustering is performed based on the high-dimensional vectors corresponding to the original data. Since the high-dimensional vectors have many dimensions, the amount of calculation during clustering is large, resulting in low clustering efficiency. This method is not suitable for clustering scenarios with a large amount of data. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a data clustering method, apparatus, electronic device, and readable storage medium to improve the problem of low clustering efficiency in the prior art.

[0005] In a first aspect, the embodiments of this application provide a data clustering method, which includes: obtaining a high-dimensional feature vector corresponding to text data to be clustered, where the high-dimensional feature vector includes semantic features of multiple dimensions; performing primary clustering on the text data according to the semantic features, and determining a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result; replacing the semantic features of multiple dimensions in the high-dimensional feature vector with a cluster label of one dimension to form a low-dimensional feature vector; and performing clustering on the low-dimensional feature vector again to obtain a final clustering result.

[0006] In the above implementation process, after performing primary clustering on the text data, according to the primary clustering result, the semantic features of multiple dimensions in the high-dimensional feature vector are replaced with a cluster label of one dimension to achieve data dimensionality reduction. Then, the dimensionality-reduced data is clustered again, so that the amount of calculation can be effectively reduced during the re-clustering process, improving the clustering efficiency. And through re-clustering, the clustering effect can be ensured. This implementation method can balance clustering efficiency and clustering accuracy.

[0007] Optionally, the performing primary clustering on the text data according to the semantic features and determining a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result includes:

[0008] Calculating a first similarity between the high-dimensional feature vector and potential clusters in a cluster set formed by clustering historical text data, where the potential clusters are clusters in the cluster set with the number of texts greater than a preset number;

[0009] Determine a target potential cluster, where the target potential cluster has the largest first similarity with the high-dimensional feature vector and is greater than a first preset similarity;

[0010] Cluster the text data into the target potential cluster, and determine the cluster label of the target potential cluster as the cluster label corresponding to the high-dimensional feature vector.

[0011] In the above implementation process, by first calculating the similarity between the text data and the potential clusters instead of all clusters, the computational amount can be reduced during the initial clustering process, improving the clustering efficiency of the initial clustering.

[0012] Optionally, the initial clustering of the text data according to semantic features and determining the cluster label corresponding to the high-dimensional feature vector according to the initial clustering result includes:

[0013] Calculate the first similarity between the high-dimensional feature vector and the potential clusters in the cluster set formed by clustering historical text data, where the potential clusters are the clusters in the cluster set with the number of texts greater than a preset number;

[0014] If there is no potential cluster with the first similarity greater than the first preset similarity, calculate the second similarity between the high-dimensional feature vector and the isolated clusters in the cluster set, where the isolated clusters are the clusters in the cluster set with the number of texts less than or equal to the preset number;

[0015] Determine a target isolated cluster, where the target isolated cluster has the largest second similarity with the high-dimensional feature vector and is greater than a second preset similarity;

[0016] Cluster the text data into the target isolated cluster, and determine the cluster label of the target isolated cluster as the cluster label corresponding to the high-dimensional feature vector.

[0017] In the above implementation process, when the text data is not similar to the potential clusters, calculating its similarity with the isolated clusters can avoid clustering the text data into the wrong potential clusters and ensure the clustering effect.

[0018] Optionally, after calculating the second similarity between the high-dimensional feature vector and the isolated clusters in the cluster set, it further includes:

[0019] If there is no isolated cluster with the second similarity greater than the second preset similarity, regard the text data as a new cluster, and determine the cluster label of the new cluster as the cluster label corresponding to the high-dimensional feature vector, thereby realizing the initial clustering of the text data.

[0020] Optionally, after clustering the text data into the target isolated cluster, it further includes:

[0021] If the number of texts in the target isolated cluster is greater than the preset number, then determine the target isolated cluster as a potential cluster, so as to ensure that when clustering subsequent text data, the potential cluster can participate in the similarity calculation first, so as to avoid the situation where the calculation amount increases due to the need to calculate the similarity with the isolated cluster again.

[0022] Optionally, after initially clustering the text data according to semantic features and obtaining the initial clustering result, it further includes:

[0023] Performing an update process on multiple clusters in the initial clustering result;

[0024] Among them, the update process includes at least one of the following: updating the cluster center of a cluster with a data volume greater than a first quantity; merging clusters with a similarity greater than a preset similarity and recalculating the cluster center of the merged cluster; deleting clusters with a data volume less than a second quantity.

[0025] In the above implementation process, performing an update process on each cluster can optimize the initial clustering result, so that the initial clustering accuracy can be improved or the calculation amount in the subsequent re-clustering process can be reduced.

[0026] Optionally, the re-clustering of the low-dimensional feature vectors to obtain the final clustering result includes:

[0027] Obtaining the text data in each cluster in the initial clustering result;

[0028] Performing density clustering on the low-dimensional feature vectors corresponding to the text data in each cluster to obtain the final clustering result.

[0029] In the above implementation process, in the case of a small data volume, performing density clustering on all data of each cluster can effectively improve the clustering accuracy and achieve a good clustering effect.

[0030] Optionally, the re-clustering of the low-dimensional feature vectors to obtain the final clustering result includes:

[0031] Obtaining the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the initial clustering result;

[0032] Performing density clustering on the low-dimensional feature vectors corresponding to the cluster centers of each cluster to obtain the final clustering result.

[0033] In the above implementation process, in the case of a large data volume, only performing density clustering on the cluster centers can improve the clustering efficiency.

[0034] In a second aspect, an embodiment of the present application provides a data clustering device, and the device includes:

[0035] A high-dimensional vector acquisition module for acquiring high-dimensional feature vectors corresponding to text data to be clustered, where the high-dimensional feature vectors include semantic features in multiple dimensions;

[0036] A primary clustering module for primarily clustering the text data according to the semantic features and determining the cluster label corresponding to the high-dimensional feature vector according to the primary clustering result;

[0037] A data dimensionality reduction module for replacing the semantic features in multiple dimensions in the high-dimensional feature vector with a cluster label in one dimension to form a low-dimensional feature vector;

[0038] A secondary clustering module for clustering the low-dimensional feature vector again to obtain a final clustering result.

[0039] Optionally, the primary clustering module is used to calculate a first similarity between the high-dimensional feature vector and a potential cluster in a cluster set formed by clustering historical text data, where the potential cluster is a cluster in the cluster set with the number of texts greater than a preset number; determine a target potential cluster, where the target potential cluster has the largest first similarity with the high-dimensional feature vector and is greater than a first preset similarity; cluster the text data into the target potential cluster, and determine the cluster label of the target potential cluster as the cluster label corresponding to the high-dimensional feature vector.

[0040] Optionally, the primary clustering module is used to calculate a first similarity between the high-dimensional feature vector and a potential cluster in a cluster set formed by clustering historical text data, where the potential cluster is a cluster in the cluster set with the number of texts greater than a preset number; if there is no potential cluster with the first similarity greater than the first preset similarity, then calculate a second similarity between the high-dimensional feature vector and an isolated cluster in the cluster set, where the isolated cluster is a cluster in the cluster set with the number of texts less than or equal to the preset number; determine a target isolated cluster, where the target isolated cluster has the largest second similarity with the high-dimensional feature vector and is greater than a second preset similarity; cluster the text data into the target isolated cluster, and determine the cluster label of the target isolated cluster as the cluster label corresponding to the high-dimensional feature vector.

[0041] Optionally, the primary clustering module is used to, if there is no isolated cluster with the second similarity greater than the second preset similarity, use the text data as a new cluster and determine the cluster label of the new cluster as the cluster label corresponding to the high-dimensional feature vector.

[0042] Optionally, the primary clustering module is used to, if the number of texts in the target isolated cluster is greater than the preset number, determine the target isolated cluster as a potential cluster.

[0043] Optionally, the initial clustering module is further configured to update multiple clusters in the initial clustering result;

[0044] Wherein, the update process includes at least one of the following: updating the cluster center of a cluster with a data volume greater than a first quantity; merging clusters with a similarity greater than a preset similarity, and recalculating the cluster center of the merged cluster; deleting a cluster with a data volume less than a second quantity.

[0045] Optionally, the re-clustering module is configured to obtain the text data in each cluster in the initial clustering result; perform density clustering on the low-dimensional feature vectors corresponding to the text data in each cluster to obtain a final clustering result.

[0046] Optionally, the re-clustering module is configured to obtain the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the initial clustering result; perform density clustering on the low-dimensional feature vectors corresponding to the cluster centers of each cluster to obtain a final clustering result.

[0047] In a third aspect, an embodiment of the present application provides a data clustering method, which is applied to a data clustering platform. The data clustering platform includes an application layer, a data mining layer, a computing layer, a feature representation layer, a preprocessing layer, and a data layer; the method includes:

[0048] Receiving, by the application layer, a clustering task sent by an upper-layer application program, where the clustering task is used to indicate clustering of specified text data;

[0049] Obtaining, by the data layer, the text data to be clustered from a database according to the clustering task;

[0050] Preprocessing, by the preprocessing layer, the text data to obtain preprocessed text data;

[0051] Performing vectorization processing on the preprocessed text data by the feature representation layer to obtain high-dimensional feature vectors corresponding to the text data;

[0052] Clustering, by the computing layer, the text data of the high-dimensional feature vectors according to a data clustering method to obtain a final clustering result;

[0053] Extracting topic words and / or key sentences for each cluster in the final clustering result by the data mining layer, and outputting the topic words and / or key sentences corresponding to each cluster through the application layer.

[0054] In a fourth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method provided in the first aspect above are run.

[0055] In a fifth aspect, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it runs the steps in the method provided in the first aspect as described above.

[0056] Other features and advantages of the present application will be described in the subsequent description, and in part, will be obvious from the description, or will be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 Schematic diagram of an electronic device for executing a data clustering method provided by an embodiment of the present application;

[0059] Figure 2 Flowchart of a data clustering method provided by an embodiment of the present application;

[0060] Figure 3 Schematic diagram of an initial clustering process and optimization of clusters provided by an embodiment of the present application;

[0061] Figure 4 Schematic diagram of data clustering provided by an embodiment of the present application;

[0062] Figure 5 Schematic diagram of a data clustering platform provided by an embodiment of the present application;

[0063] Figure 6 Schematic diagram of data clustering by process provided by an embodiment of the present application;

[0064] Figure 7 Block diagram of the structure of a data clustering device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application.

[0066] The embodiment of the present application provides a data clustering method. After initially clustering the text data, the semantic features of multiple dimensions in the high-dimensional feature vector are replaced with a class cluster label of one dimension according to the initial clustering result, so as to realize the dimensionality reduction of the data. Then, the dimensionality-reduced data is clustered again, which can effectively reduce the amount of calculation during the re-clustering process, and the clustering effect can be ensured through the re-clustering. Therefore, the implementation method of the present application can balance the clustering efficiency and clustering accuracy.

[0067] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of an electronic device for executing the data clustering method provided by the embodiment of the present application. The electronic device may include: at least one processor 110, such as a CPU, at least one communication interface 120, at least one memory 130, and at least one communication bus 140. Among them, the communication bus 140 is used to realize the direct connection communication of these components. Among them, the communication interface 120 of the device in the embodiment of the present application is used to communicate with other node devices for signaling or data. The memory 130 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 130 may also be at least one storage device located far from the aforementioned processor. The memory 130 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 110, the electronic device executes the following Figure 2 method process shown. For example, the memory 130 can be used to store text data, and the processor 110 can be used to vectorize the text data, convert it into a high-dimensional feature vector, and then reduce the dimensionality of the high-dimensional feature vector using the initial clustering result, and then perform secondary clustering to realize the clustering of the text data.

[0068] The electronic device may be a terminal device or a server and other devices with certain data processing capabilities.

[0069] It can be understood that Figure 1 the structure shown is only schematic, and the electronic device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown. Figure 1 Each component shown in

[0070] Please refer to Figure 2 , Figure 2 which is a flowchart of a data clustering method provided by the embodiment of the present application. The method includes the following steps:

[0071] Step S110: Obtain the high-dimensional feature vectors corresponding to the text data to be clustered, where the high-dimensional feature vectors include semantic features in multiple dimensions.

[0072] In some embodiments, the text data to be clustered may be text stored in an electronic device or text data received by the electronic device from the outside. In this case, clustering can be performed on the text data received from the outside in real time. The text data may be, for example, news text, user data, image data, device data, etc., or may also be text data converted from voice data.

[0073] If the electronic device directly obtains text data, it can also first perform vectorization processing on the text data to convert it into high-dimensional feature vectors. The vectorization processing can use word vector conversion methods, such as the Word2Vector algorithm, the doc2vec algorithm, etc., or can also use the bert multilingual model for vector conversion. After using the bert multilingual model to perform vectorization processing on the text data, the obtained high-dimensional feature vectors include rich semantic information. The input of the bert multilingual model is the original word vectors of each word / character in the text data, and the original word vectors can be obtained by performing vector processing on each word / character in the text data using the Word2Vector algorithm. The output of the bert multilingual model is the vector representation of each word / character after integrating the full-text semantic information. Since the full-text semantic information is relatively rich and the text data can be described in multiple dimensions, generally, high-dimensional vectors are used for representation, which is called high-dimensional feature vectors in the embodiments of the present application.

[0074] Of course, the electronic device may directly obtain the high-dimensional feature vectors corresponding to the text data. In this case, the high-dimensional feature vectors can be obtained after the external device performs vector conversion on the text data. Since the text data can be described in multiple dimensions, in order to accurately describe the text data, the converted vectors are high-dimensional feature vectors, and the high-dimensional feature vectors include semantic features in multiple dimensions. For example, the high-dimensional feature vectors include 512 dimensions, where 100 dimensions are used to represent the semantic features of the text data, and the other dimensions can represent other features of the text data, such as the upper position of each word, the time when the text data is generated, the source, etc.

[0075] In addition, the above text data may refer to a single piece of text data or multiple pieces of text data, and each piece of text data corresponds to a high-dimensional feature vector.

[0076] Step S120: Perform primary clustering on the text data according to the semantic features, and determine the cluster labels corresponding to the high-dimensional feature vectors according to the primary clustering results.

[0077] During the initial clustering, clustering can be performed based on semantic features. For example, text data with similar semantic features can be clustered into one category. When the text data is news text and clustering is performed on the news text, since news text is generated in real time, in this case, clustering is performed for each piece of text data obtained. This clustering process refers to clustering the text data into the clusters formed by clustering historical text data. At this time, clustering can be performed in real time or at regular intervals.

[0078] For example, if clustering is performed in real time, initially, if only one piece of text data is obtained, this piece of text data forms a single cluster during the initial clustering. For subsequent newly incoming text data, the similarity between this text data and the already formed clusters is calculated based on semantic features, and then the initial clustering is achieved. When clustering some existing text data, the similarity between these text data can be calculated based on semantic features (such as using 100-dimensional semantic features), and then text data with a similarity greater than a preset similarity is clustered into one category, thereby obtaining multiple clusters, that is, the initial clustering result.

[0079] Each cluster can be set with a cluster label. For example, for three clusters, their cluster labels are 1, 2, and 3. The semantic features represented by different cluster labels can be stored in an electronic device. For example, the cluster with cluster label 1 clusters text data related to a certain current event news together, then the semantic feature (or it can be called a topic) corresponding to the cluster label 1 is this current event news; the cluster with cluster label 2 clusters text data related to a certain star topic together, then the semantic feature corresponding to the cluster label 2 is this star topic. Therefore, the cluster label corresponding to the high-dimensional feature vector can be determined according to the initial clustering result. For example, if text data 1 is clustered into the cluster with cluster label 1, then the cluster label of the high-dimensional feature vector corresponding to this text data 1 is 1; if text data 2 is clustered into the cluster with cluster label 2, then the cluster label of the high-dimensional feature vector corresponding to this text data 2 is 2.

[0080] Step S130: Replace the semantic features of multiple dimensions in the high-dimensional feature vector with a cluster label of one dimension to form a low-dimensional feature vector.

[0081] Since directly using high-dimensional feature vectors for clustering will consume a large amount of resources of the electronic device and involve large computational complexity, and real-time stream clustering (stream clustering refers to clustering data streams, such as the scenario of clustering news texts) cannot be achieved for a large amount of data, it is necessary to first reduce the dimension of the high-dimensional feature vectors to reduce the computational complexity in the subsequent clustering process. In the embodiments of the present application, the semantic features of multiple dimensions in the high-dimensional feature vectors are replaced with a class cluster label of one dimension, thereby forming low-dimensional feature vectors. For example, if 100 dimensions are used to describe the semantic features in the above-mentioned high-dimensional feature vectors, then the semantic features of these 100 dimensions are replaced with a one-dimensional class cluster label, such as the class cluster label 1 or 2 mentioned above. In this way, the semantic features in the original high-dimensional feature vectors can be compressed to form simple one-dimensional eigenvalue, that is, using the class cluster label to represent the semantic features.

[0082] In the embodiments of the present application, the low-dimensional feature vectors can be three-dimensional feature vectors, which can be expressed as <time, service attribute, class cluster label>. These three dimensions constitute the low-dimensional feature vectors and can be described as the spatial features corresponding to the text data. Of course, the other two dimensions can also be flexibly set according to actual requirements. In the present application, considering that the text data is news text and is highly correlated with time, the time dimension is retained in the low-dimensional feature vectors. In addition, the service attribute can be related to the clustering task. For example, if text data in different languages are clustered into one category, the service attribute can be the language used in the text data; or if sensitive data is classified into one category, the service attribute is whether the text data is sensitive data, etc.

[0083] That is to say, in the initial clustering result, text data with similar semantic features are initially clustered into one category. According to the initial clustering result, the high-dimensional feature vectors are converted into low-dimensional feature vectors to achieve data dimension reduction. Thus, when performing re-clustering subsequently, the low-dimensional feature vectors can be clustered, which can reduce the calculation of the data volume and improve the clustering efficiency.

[0084] Step S140: Re-cluster the low-dimensional feature vectors to obtain the final clustering result.

[0085] After obtaining the low-dimensional feature vectors, the low-dimensional feature vectors can be re-clustered. After the initial clustering, the initial clustering result is obtained. Each text data in the initial clustering result corresponds to a low-dimensional feature vector. Therefore, when performing re-clustering, the similarity between two text data is calculated using the low-dimensional feature vectors, so that text data with high similarity can be clustered into one category, realizing batch clustering of text data. Batch clustering means that text data with a large data volume can be clustered and has a good clustering result.

[0086] In the above implementation process, after initially clustering the text data, according to the initial clustering results, the semantic features of multiple dimensions in the high-dimensional feature vectors are replaced with the cluster labels of one dimension to achieve data dimensionality reduction. Then, the data after dimensionality reduction is clustered again, so that the computational amount can be effectively reduced during the second clustering process, the clustering efficiency can be improved, and the clustering effect can be ensured through the second clustering. This implementation method can balance the clustering efficiency and clustering accuracy.

[0087] In some embodiments, the above-mentioned stream clustering may refer to a single-pass clustering algorithm, while batch clustering may refer to a density clustering algorithm. The clustering effect of batch clustering is better than that of stream clustering. However, if only batch clustering is performed on text data, repeated calculations are required during the calculation process, and the calculation is complex and requires a large amount of memory. It is difficult to obtain results using the batch clustering algorithm in the face of a large amount of data. For example, when using batch clustering to calculate incremental text data for several days, a large amount of hardware resources and a large amount of time are required to obtain the clustering results, and general servers are difficult to support such a large amount of calculation. For stream clustering, clustering can be achieved with limited memory and limited processing time, but its clustering effect is not as good as that of batch clustering, and the text data faced by both clusterings are high-dimensional vectors, resulting in a large amount of calculation and very low clustering efficiency. Therefore, in this application, the two clustering algorithms are combined. The stream clustering algorithm is used to achieve data dimensionality reduction, and then the batch clustering algorithm is used to achieve clustering, so that the clustering efficiency and clustering effect can be balanced.

[0088] The following will elaborate on the two processes of initial clustering and second clustering in detail.

[0089] In some embodiments, in order to be able to perform real-time clustering on news text data and improve the clustering efficiency, in the initial clustering in the embodiments of this application, a single-pass clustering algorithm is used as the framework, and the cluster structures of potential clusters and isolated clusters are introduced to achieve initial clustering.

[0090] The single-pass clustering algorithm belongs to non-hierarchical clustering. Its clustering process is an iterative process, and the algorithm efficiency is relatively high, suitable for processing text data with a relatively large amount of data. Moreover, the single-pass clustering algorithm is very sensitive to the data time order. If the data time order is different, the final clustering results may also be different. Therefore, the single-pass clustering algorithm is very suitable for clustering news text data and can better meet the application requirements for clustering news text.

[0091] The specific implementation method is as follows: Calculate the first similarity between the high-dimensional feature vector and the potential clusters in the set of clusters formed by clustering historical text data, where the potential clusters are the clusters in the set of clusters with the number of text data greater than a preset number; Determine the target potential cluster, where the first similarity between the target potential cluster and the high-dimensional feature vector is the largest and greater than the first preset similarity; Cluster the text data into the target potential cluster, and determine the cluster label of the target potential cluster as the cluster label corresponding to the high-dimensional feature vector.

[0092] The idea of the existing single-pass clustering algorithm is mainly as follows: If there is no cluster currently, take the text data to be clustered as the first cluster. If there is a set of clusters that have been clustered, calculate the similarity between the text data to be clustered and all clusters, and take the cluster with the largest similarity and the similarity greater than the preset threshold, and merge the text data into this cluster. Otherwise, generate a new cluster with this text data, and repeat this process until all the text data to be clustered is processed.

[0093] In the embodiments of the present application, the existing single-pass clustering algorithm is improved to reduce the computational complexity of the single-pass clustering algorithm and improve the clustering efficiency. The implementation method is to divide the clusters into potential clusters and isolated clusters. The potential clusters are the clusters in the set of clusters with the number of text data greater than a preset number, and the isolated clusters are the clusters in the set of clusters with the number of text data less than or equal to the preset number.

[0094] Since the clusters with a small number of text data may be caused by data noise or clustering deviation, the similarity between the text data and the potential clusters can be calculated first, so that it is not necessary to calculate the similarity with all clusters, thereby reducing some computational complexity and improving the clustering efficiency.

[0095] Among them, the set of clusters is formed by clustering historical text data. The historical text data can be the text data clustered before the current text data to be clustered. For example, if 50 pieces of text data need to be clustered, according to the idea of the above single-pass clustering algorithm, first take the first piece of text data as a cluster. At this time, the formed set of clusters includes one cluster, and the historical text data includes the first piece of text data. When clustering the second piece of text data, calculate the similarity between the second piece of text data and this cluster. If the second piece of text data is merged into this cluster, the historical text data includes these two pieces of text data. In this way, the subsequent text data can be clustered in turn. It should be noted here that since the number of text data in each cluster is relatively small initially, the clusters in the set of clusters may all be isolated clusters. Therefore, when clustering in the initial time period, the text data can be clustered into the isolated clusters first. After a certain period of time, if the number of text data in a certain isolated cluster is greater than the preset number, then convert this cluster into a potential cluster. For the clustering of subsequent text data, its similarity with the potential clusters can be calculated first, and the similarity with the isolated clusters is no longer calculated, which can reduce a certain amount of computational complexity.

[0096] Among them, calculating the first similarity refers to calculating the similarity between the high-dimensional feature vector corresponding to the text data and the cluster center of the potential cluster. The vector corresponding to the cluster center can refer to the average vector of the high-dimensional feature vectors corresponding to all the text data in this type of cluster. Therefore, the first similarity between the two vectors can be calculated by calculating the cosine distance or other methods, or the first similarity can also be calculated by calculating the Kullback-Leibler, Jaccard, Hellinger distance, etc. If there are multiple potential clusters, the first similarity between the text data and each potential cluster can be calculated, and then the maximum first similarity can be determined. For example, if the first similarity between potential cluster 1 and the text data is the largest, and the first similarity between potential cluster 1 and the text data is greater than the first preset similarity, then potential cluster 1 is called the target potential cluster. Then, the text data can be clustered into this potential cluster 1, and then the cluster label of potential cluster 1, such as 1, can be determined as the cluster label corresponding to the high-dimensional feature vector of the text data, so as to facilitate subsequent conversion into a low-dimensional feature vector.

[0097] In the above implementation process, calculating the similarity between the text data and the potential cluster first, rather than calculating the similarity with all clusters, can reduce the amount of calculation and improve the clustering efficiency of the initial clustering.

[0098] In some embodiments, after calculating the similarity between the text data and the potential cluster first, if there is no potential cluster with a first similarity greater than the first preset similarity, then calculate the second similarity between the high-dimensional feature vector and the isolated cluster in the cluster set. The isolated cluster is a cluster in the cluster set with the number of texts less than or equal to the preset number. Determine the target isolated cluster. The second similarity between the target isolated cluster and the high-dimensional feature vector is the largest and greater than the second preset similarity. Cluster the text data into the target isolated cluster, and determine the cluster label of the target isolated cluster as the cluster label corresponding to the high-dimensional feature vector.

[0099] Among them, the method of calculating the second similarity is similar to the method of calculating the first similarity, that is, calculating the similarity between the high-dimensional feature vector corresponding to the text data and the cluster center of the isolated cluster. If there are multiple isolated clusters, the second similarity between the text data and each isolated cluster can be calculated, and then the largest second similarity can be determined. For example, if the second similarity between isolated cluster 3 and the text data is the largest, and the second similarity between isolated cluster 3 and the text data is greater than the second preset similarity, then isolated cluster 3 can be determined as the target isolated cluster. Then, the text data can be clustered into this isolated cluster 3, and then the cluster label of isolated cluster 3, such as 3, can be determined as the cluster label corresponding to the high-dimensional feature vector of the text data, so as to facilitate subsequent conversion into a low-dimensional feature vector.

[0100] Among them, the first preset similarity and the second preset similarity in the above two methods can be flexibly set according to actual needs, and the first preset similarity and the second preset similarity can be different or the same. This application embodiment does not make special limitations on this.

[0101] The preset quantity mentioned above can also be flexibly set according to actual needs, such as 2 or 3, etc. This application embodiment does not make special limitations on this.

[0102] In the above implementation process, when the text data is not similar to the potential cluster, calculate its similarity with the isolated cluster, which can avoid clustering the text data into the wrong potential cluster and ensure the clustering effect.

[0103] In some implementation manners, after calculating the second similarity between the text data and the isolated cluster, if there is no isolated cluster with a second similarity greater than the second preset similarity, then the text data is used as a new cluster, and the cluster label of the new cluster is determined as the cluster label corresponding to the high-dimensional feature vector.

[0104] That is to say, when there is no cluster with a relatively high similarity to the text data in the current cluster set, the text data is used as the text in a new cluster alone, that is, a new cluster is constructed, and the cluster label of the new cluster can be assigned according to certain rules. For example, in the process of forming clusters, a data structure similar to an index will be constructed to record the cluster labels corresponding to each cluster, so that the cluster label of the new cluster can be generated according to the cluster labels of the existing clusters. For example, if the existing cluster labels are numerical numbers and the cluster label of the last cluster is 5, the cluster label of the generated new cluster can be 6. In this way, the text data can be clustered into the new cluster, and the cluster label of the new cluster is the cluster label corresponding to the high-dimensional feature vector of the text data.

[0105] To facilitate the management of each cluster, the electronic device can record the relevant information of each cluster, such as including the cluster label of the cluster, the number of texts in the cluster, the cluster center, the numbers of the text data in the cluster, the mark of the potential cluster or the isolated cluster, the expiration time of the cluster (that is, the time to delete the cluster), etc.

[0106] Among them, the electronic device can scan the number of texts in each cluster in real time and mark the cluster as a potential cluster or an isolated cluster according to the number of texts. When calculating the similarity, it can be known which clusters in the cluster set are potential clusters and which clusters are isolated clusters by obtaining the information recorded by the electronic device, so that the potential cluster or the isolated cluster can be quickly found.

[0107] In some embodiments, as clustering progresses, the number of texts in isolated clusters may gradually increase. The electronic device can also scan the number of texts in isolated clusters in real time, so as to determine an isolated cluster as a potential cluster after the number of texts in the isolated cluster exceeds a preset number. For example, after clustering the above text data into a target isolated cluster, if the number of texts in the target isolated cluster exceeds the preset number, the target isolated cluster is determined as a potential cluster.

[0108] In the specific implementation process, when the electronic device determines to determine an isolated cluster as a potential cluster, it can change the relevant information of the recorded isolated cluster. For example, change the label of the isolated cluster of this type of cluster recorded to the label of the potential cluster. For example, the initial label of this type of cluster recorded is 0, indicating that this type of cluster is an isolated cluster. If this type of cluster is converted into a potential cluster, the label is changed to 1, indicating that this type of cluster is a potential cluster. When clustering subsequent text data, this type of cluster can be included in the potential clusters, and the similarity with the text data is calculated first.

[0109] In the above implementation process, determining an isolated cluster as a potential cluster ensures that when clustering subsequent text data, the potential clusters can participate in the similarity calculation first, so as to avoid the situation where the calculation amount increases due to the need to calculate the similarity with isolated clusters.

[0110] In some embodiments, in order to achieve better clustering results, after obtaining the initial clustering result through the initial clustering, the multiple clusters in the initial clustering result can also be updated. The update process includes at least one of the following: updating the cluster center of the cluster with the data volume greater than the first quantity; merging at least two clusters with the similarity greater than the preset similarity (the value of the preset similarity here can be the same as or different from the above first preset similarity or second preset similarity), and recalculating the cluster center of the merged cluster; deleting the cluster with the data volume less than the second quantity.

[0111] For example, the electronic device can be responsible for updating each cluster generated in the initial clustering through a cluster center optimization thread. For example, updating the cluster center. Since the cluster center is the average vector of the high-dimensional feature vectors of each text data in the cluster, when the data volume in the cluster increases, the cluster center will shift (because as time goes by, the focus of the topic that may be concerned will also change continuously. In order to reflect the change characteristics of the text focus in this cluster over time, the cluster center needs to be updated), and in order to obtain more accurate similarity and be able to discover new topics subsequently, the cluster center needs to be updated. The update method is to recalculate an average vector by reusing the high-dimensional feature vectors of the text data in the current cluster, that is, recalculate a cluster center, and then replace the original cluster center with the new cluster center.

[0112] The electronic device can also update the cluster center when the data volume of the cluster increases, so that the update of the cluster center can be carried out in a timely manner. Or the electronic device can update the cluster center at regular intervals, or update the cluster center every time a certain amount of data grows. It can be understood that the update trigger condition of the cluster center can be flexibly set according to actual needs.

[0113] Of course, since clustering is essentially an approximate calculation, it is inevitable to introduce some text data with less relevance to the theme in a cluster. If the cluster center is updated frequently, it may cause a large deviation of the cluster center. Therefore, in the embodiments of the present application, the cluster center of the cluster can be updated only when the data volume is greater than the first quantity, so as to avoid the problem of large deviation of the cluster center.

[0114] Although the initial clustering divides the data into multiple clusters, the clustering accuracy of the initial clustering is not very high. There may be some clusters with relatively high similarity after clustering. At this time, these clusters with high similarity can be merged to reduce the computational complexity when clustering subsequent text data. For example, calculate the similarity between every two clusters. If the similarity is greater than the preset similarity, at least two clusters are merged. For example, the similarity between cluster 1 and cluster 2 is greater than the preset similarity, and the similarity between cluster 2 and cluster 3 is greater than the preset similarity, then these three clusters can be merged into one cluster; or, two clusters with similarity greater than the preset similarity can be merged first, and then calculate the similarity with other clusters, and then determine whether to merge. For example, calculate the similarity between every two clusters for the first time, and then merge them in pairs (here, if the similarity between a cluster and multiple other clusters is greater than the preset similarity, any one of the clusters is randomly selected for merging). After the first round of merging is completed, calculate the similarity between the new clusters, and continue the second round of merging in the same way until the similarity between every two clusters is less than or equal to the preset similarity. Among them, calculating the similarity is to calculate the similarity between the cluster centers of two clusters. Since the cluster center of the merged cluster has changed, it is also necessary to recalculate the cluster center of the merged cluster and use the calculated cluster center as the cluster center of the merged cluster.

[0115] In addition, since some clusters may be formed by some noise data points or clustering errors and have a small data volume, the clusters with a data volume less than the second quantity can also be deleted, which can also reduce the computational complexity in the subsequent clustering process of text data.

[0116] The above-mentioned first quantity and second quantity can be flexibly set according to actual needs. There is no certain correlation between the first quantity and the second quantity. The first quantity can be greater than the second quantity, less than the second quantity, or, in some cases, equal to the second quantity.

[0117] In the above implementation process, by performing an initial clustering on the text data and updating each cluster, the initial clustering result can be optimized, so as to improve the initial clustering accuracy or reduce the computational amount in the subsequent re-clustering process. The schematic of the initial clustering and optimization process can be as Figure 3 shown.

[0118] In the process of performing an initial clustering on the text data using the above method, in some other embodiments, the clustering algorithm used for the initial clustering can also be other clustering algorithms, such as the principal component analysis algorithm, the hierarchical clustering algorithm, etc. Alternatively, the text data can be first dimensionally reduced using the principal component analysis algorithm or the hierarchical clustering algorithm, and then the above method is used for the initial clustering and then continue to reduce the dimension, so as to effectively reduce the computational amount in the subsequent re-clustering process.

[0119] The process of re-clustering will be introduced below.

[0120] In some embodiments, the amount of data obtained after the initial clustering may be relatively small. To improve the clustering accuracy, all the data can be subjected to density clustering. The specific implementation method is: obtain the text data of each cluster in the initial clustering result, and then perform density clustering on the low-dimensional feature vectors corresponding to the text data in each cluster to obtain the final clustering result.

[0121] Among them, a typical density clustering algorithm is the Density-Based Spatial Clustering of Applications with Noise (DBSCAN). This density clustering algorithm uses the "neighborhood" probability to describe the tightness of the sample distribution, divides the regions with sufficient density into clusters, and can find clusters of any shape under noisy conditions. Its basic implementation idea is: derive the sample set with the maximum density connection from the density reachable relationship. There is one or more core objects in such a set. If there is only one core object, then other non-core objects in the cluster are all in the neighborhood of this core object. If there are multiple core objects, then the neighborhood of any one core object must contain another core object (otherwise it cannot be density reachable). These core objects and all the samples contained in its neighborhood form a cluster.

[0122] In this implementation method, density clustering is performed on the text data in all clusters. For example, if cluster 1 contains 100 pieces of data and cluster 2 contains 200 pieces of data, then these 300 pieces of data can be merged together for density clustering, that is, the object of density clustering is each piece of data. Since the vectors corresponding to these 300 pieces of data are low-dimensional vectors, the amount of calculation can be effectively reduced during the density clustering process, and the clustering effect can be ensured. In this way, the text data can be reclustered based on the low-dimensional feature vectors again. The text data is dimensionally reduced through the first clustering, and the clustering of the text data is achieved through the second clustering, so that both the clustering efficiency and the clustering accuracy can be taken into account.

[0123] Alternatively, in some other implementation methods, when performing density clustering, it is also possible to perform density clustering on the text data in each cluster in the result of the first clustering. For example, if cluster 1 contains 100 pieces of data, then density clustering is performed on these 100 pieces of data. If cluster 2 contains 200 pieces of data, then density clustering is performed on these 200 pieces of data. Density clustering is performed separately for each cluster. In this way, the text data in each cluster does not need to calculate the similarity with the text data in other clusters, which can reduce a certain amount of calculation. In this way, the clusters formed by the first clustering can be divided into smaller clusters, realizing the subdivision of the text data under a large topic into each small topic, which can make the clustering more accurate and improve the clustering accuracy.

[0124] It can be understood that the specific clustering process of density clustering can refer to the existing related implementation methods and will not be described in detail here.

[0125] In the above implementation process, when the amount of data is small, performing density clustering on all data in each cluster can effectively improve the clustering accuracy and achieve a good clustering effect.

[0126] In some other implementation methods, since the amount of data obtained after the first clustering may be relatively large, although the amount of calculation can be reduced after reducing the high-dimensional feature vectors to low-dimensional feature vectors, if all the data is clustered again, the amount of calculation is still relatively large. Therefore, it is also possible to perform density clustering only on the clusters. The specific implementation method is as follows: Obtain the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the result of the first clustering, perform density clustering on the low-dimensional feature vectors corresponding to the cluster centers of each cluster, and obtain the final clustering result.

[0127] In this way, the object of density clustering is the cluster. If the number of clusters formed after the first clustering is relatively large, then density clustering can be performed on these clusters. It calculates the distances between the cluster centers of each cluster to determine whether each cluster can be merged in turn. The process can be as Figure 4as shown (where some clusters can also be deleted). Therefore, the objects involved in its merging are clusters, which can effectively improve the clustering efficiency and achieve fast data clustering.

[0128] An embodiment of the present application also provides a data clustering platform, which can be understood as a software platform running in the above-mentioned electronic device. It can be used to run the above-mentioned data clustering method. The data clustering platform is as Figure 5 shown, and it includes an application layer, a data mining layer, a computing layer, a feature representation layer, a preprocessing layer, and a data layer.

[0129] Among them, the platform can obtain streaming data and knowledge-based data through the data layer. The data enters the preprocessing layer for processing such as feature selection and filtering. The preprocessed data enters the feature representation layer to represent features of users and then enters the computing layer for clustering processing. The clustering results enter the data mining layer for data mining work such as extracting subject words and key sentences. The application layer provides stream clustering or batch clustering services.

[0130] The functions of each layer will be introduced separately below.

[0131] Application layer: It can receive clustering tasks sent by upper-layer application programs. For example, the clustering task is used to indicate clustering of specified text data, or stream clustering (i.e., the above-mentioned single-pass clustering algorithm) or batch clustering (i.e., the above-mentioned density clustering algorithm) tasks for data. The application layer can perform functions such as task management and process allocation.

[0132] The application layer includes a master-slave backup module, which is used for functions such as master-slave backup of clustering results and realizes data interaction with the data mining layer, the computing layer, and the data layer. It can also perform task queue management, that is, manage the execution status data of tasks in the queue, etc.

[0133] The application layer also has a timed task function, which is used for timed recovery of leaked memory, cleaning up zombie processes caused by exceptions, cleaning system logs in the hard disk, cached data, timed monitoring or restarting daemon processes, etc.

[0134] The application layer also has a multi-process management function, which can instantiate multiple processes according to the task assignment results of task scheduling for subsequent computing work. For example, different processes are used for clustering different text data, as Figure 6 shown, and it can also manage the status of these processes and recycle resources in real time.

[0135] Data mining layer: It has the function of keyword extraction and is responsible for extracting keywords from each cluster in the final clustering result; it also has the function of key sentence extraction and is responsible for extracting key sentences from each cluster in the final clustering result; it also has the function of text deduplication and is responsible for removing duplicate texts in each cluster of the final clustering result, so as to facilitate subsequent data mining, data retrieval, etc.

[0136] Computing layer: It is used to execute the above data clustering method to achieve stream clustering and / or batch clustering of data. When performing batch clustering, the clustering computing process will perform clustering calculations on data separately according to the requirements of data and business. For example, after preliminary clustering, batch clustering can be performed in real time or after a period of time. Stream clustering and batch clustering use different processes.

[0137] Feature representation layer: It is used to vectorize text data, such as converting text data into high-dimensional feature vectors.

[0138] Preprocessing layer: It is used to preprocess text data, including language classification, word segmentation, filtering out interfering texts, feature selection, format conversion, etc.

[0139] Data layer: It is used to receive data from the outside, store data, and support data query. It can use kafka and a pre-constructed knowledge base for data storage. The data received from the outside can include data streams and data sets, etc. Data sets are suitable for situations with time sequence and small data sets, and data streams are suitable for real-time monitoring systems. The data has the characteristics of time and rapid change. The data stream changes continuously and has a huge amount of data, which can be called a data stream.

[0140] It can be understood that in practical applications, the functions of each layer can be flexibly divided according to the situation, and the functions of each layer can be flexibly increased or decreased, etc. The division of each layer of the data clustering platform is not limited to the several layers shown above, and some layers can also be flexibly increased or decreased according to requirements.

[0141] Please refer to Figure 7 , Figure 7 FIG. 22 is a schematic structural diagram of a data clustering device 200 provided by an embodiment of the present application. The device 200 can be a module, a program segment, or code on an electronic device. It should be understood that the device 200 corresponds to the above Figure 2 method embodiment and can execute Figure 2 each step involved in the method embodiment. The specific functions of the device 200 can be seen in the above description. To avoid repetition, the detailed description is appropriately omitted here.

[0142] Optionally, the device 200 includes:

[0143] A high-dimensional vector acquisition module 210, configured to acquire high-dimensional feature vectors corresponding to text data to be clustered, where the high-dimensional feature vectors include semantic features in multiple dimensions;

[0144] A primary clustering module 220, configured to perform primary clustering on the text data according to the semantic features, and determine a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result;

[0145] A data dimensionality reduction module 230, configured to replace the semantic features in multiple dimensions in the high-dimensional feature vector with a cluster label in one dimension to form a low-dimensional feature vector;

[0146] A secondary clustering module 240, configured to perform secondary clustering on the low-dimensional feature vector to obtain a final clustering result.

[0147] Optionally, the primary clustering module 220 is configured to calculate a first similarity between the high-dimensional feature vector and a potential cluster in a cluster set formed by clustering historical text data, where the potential cluster is a cluster in the cluster set with the number of texts greater than a preset number; determine a target potential cluster, where the target potential cluster has the largest first similarity with the high-dimensional feature vector and is greater than a first preset similarity; cluster the text data into the target potential cluster, and determine the cluster label of the target potential cluster as the cluster label corresponding to the high-dimensional feature vector.

[0148] Optionally, the primary clustering module 220 is configured to calculate a first similarity between the high-dimensional feature vector and a potential cluster in a cluster set formed by clustering historical text data, where the potential cluster is a cluster in the cluster set with the number of texts greater than a preset number; if there is no potential cluster with the first similarity greater than the first preset similarity, calculate a second similarity between the high-dimensional feature vector and an isolated cluster in the cluster set, where the isolated cluster is a cluster in the cluster set with the number of texts less than or equal to the preset number; determine a target isolated cluster, where the target isolated cluster has the largest second similarity with the high-dimensional feature vector and is greater than a second preset similarity; cluster the text data into the target isolated cluster, and determine the cluster label of the target isolated cluster as the cluster label corresponding to the high-dimensional feature vector.

[0149] Optionally, the primary clustering module 220 is configured to, if there is no isolated cluster with the second similarity greater than the second preset similarity, use the text data as a new cluster, and determine the cluster label of the new cluster as the cluster label corresponding to the high-dimensional feature vector.

[0150] Optionally, the primary clustering module 220 is configured to, if the number of texts in the target isolated cluster is greater than the preset number, determine the target isolated cluster as a potential cluster.

[0151] Optionally, the initial clustering module 220 is further configured to perform an update process on multiple clusters in the initial clustering result;

[0152] The update process includes at least one of the following: updating the cluster center of a cluster with a data volume greater than a first quantity; merging clusters with a similarity greater than a preset similarity and recalculating the cluster center of the merged clusters; deleting clusters with a data volume less than a second quantity.

[0153] Optionally, the re-clustering module 240 is configured to obtain the text data in each cluster in the initial clustering result; perform density clustering on the low-dimensional feature vectors corresponding to the text data in each cluster to obtain a final clustering result.

[0154] Optionally, the re-clustering module 240 is configured to obtain the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the initial clustering result; perform density clustering on the low-dimensional feature vectors corresponding to the cluster centers of each cluster to obtain a final clustering result.

[0155] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0156] An embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the method process executed by the electronic device in the method embodiment as Figure 2 shown.

[0157] This embodiment discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the foregoing method embodiments. For example, it includes: obtaining a high-dimensional feature vector corresponding to text data to be clustered, where the high-dimensional feature vector includes semantic features of multiple dimensions; performing initial clustering on the text data according to the semantic features and determining a cluster label corresponding to the high-dimensional feature vector according to the initial clustering result; replacing the semantic features of multiple dimensions in the high-dimensional feature vector with a cluster label of one dimension to form a low-dimensional feature vector; and performing clustering on the low-dimensional feature vector again to obtain a final clustering result.

[0158] In summary, the embodiments of the present application provide a data clustering method, apparatus, electronic device, and readable storage medium. After initially clustering text data, according to the initial clustering result, semantic features of multiple dimensions in the high-dimensional feature vector are replaced with a class cluster label of one dimension to achieve data dimensionality reduction. Then, the data after dimensionality reduction is clustered again, thereby effectively reducing the amount of calculation and improving the clustering efficiency during the second clustering process. Moreover, the clustering effect can be ensured through the second clustering. This implementation method can balance clustering efficiency and clustering accuracy.

[0159] In the embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some communication interfaces. The indirect coupling or communication connection of the apparatus or unit can be in an electrical, mechanical, or other form.

[0160] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] Furthermore, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0162] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0163] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data clustering method, characterized in that, The method includes: Obtaining a high-dimensional feature vector corresponding to text data to be clustered, where the high-dimensional feature vector includes semantic features in multiple dimensions; Performing primary clustering on the text data according to the semantic features, and determining a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result; when performing primary clustering, using a single-pass clustering algorithm as a framework, introducing a cluster structure of potential clusters and isolated clusters to implement primary clustering, where the potential cluster is a cluster in the cluster set with the number of texts greater than a preset number, and the isolated cluster is a cluster in the cluster set with the number of texts less than or equal to the preset number. When the number of texts in the isolated cluster is greater than the preset number, converting the isolated cluster into a potential cluster. The cluster set is formed by clustering historical text data, and the historical text data is the text data clustered before the current text data to be clustered; Replacing the semantic features in multiple dimensions in the high-dimensional feature vector with a cluster label in one dimension to form a low-dimensional feature vector; Performing clustering again on the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the primary clustering result or the text data of each cluster in the primary clustering result, using a stream clustering algorithm to achieve data dimensionality reduction, and then using a batch clustering algorithm to achieve clustering to obtain a final clustering result.

2. The method according to claim 1, wherein The performing primary clustering on the text data according to the semantic features and determining a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result includes: Calculating a first similarity between the high-dimensional feature vector and potential clusters in the cluster set formed by clustering historical text data, where the potential clusters are clusters in the cluster set with the number of texts greater than a preset number; Determining a target potential cluster, where the target potential cluster has the largest first similarity with the high-dimensional feature vector and is greater than a first preset similarity; Clustering the text data into the target potential cluster, and determining the cluster label of the target potential cluster as the cluster label corresponding to the high-dimensional feature vector.

3. The method according to claim 1, characterized in that The performing primary clustering on the text data according to the semantic features and determining a cluster label corresponding to the high-dimensional feature vector according to the primary clustering result includes: Calculating a first similarity between the high-dimensional feature vector and potential clusters in the cluster set formed by clustering historical text data; If there is no potential cluster with the first similarity greater than the first preset similarity, then calculating a second similarity between the high-dimensional feature vector and isolated clusters in the cluster set; Determining a target isolated cluster, where the target isolated cluster has the largest second similarity with the high-dimensional feature vector and is greater than a second preset similarity; Clustering the text data into the target isolated cluster, and determining the cluster label of the target isolated cluster as the cluster label corresponding to the high-dimensional feature vector.

4. The method according to claim 3, characterized in that, After calculating the second similarity between the high-dimensional feature vector and isolated clusters in the cluster set, it further includes: If there is no isolated cluster with the second similarity greater than the second preset similarity, then taking the text data as a new cluster, and determining the cluster label of the new cluster as the cluster label corresponding to the high-dimensional feature vector.

5. The method according to claim 3, wherein After clustering the text data into the target isolated cluster, it further includes: If the number of texts in the target isolated cluster is greater than the preset number, then determine the target isolated cluster as a potential cluster.

6. The method according to claim 1, wherein After obtaining the initial clustering result by performing initial clustering on the text data according to semantic features, it further includes: Performing an update process on multiple clusters in the initial clustering result; Among them, the update process includes at least one of the following: updating the cluster center of a cluster with a data volume greater than the first quantity; merging at least two clusters with a similarity greater than the preset similarity, and recalculating the cluster center of the merged cluster; deleting a cluster with a data volume less than the second quantity.

7. The method according to claim 1, characterized in that The re-clustering of the low-dimensional feature vectors to obtain the final clustering result includes: Obtaining the text data in each cluster in the initial clustering result; Performing density clustering on the low-dimensional feature vectors corresponding to the text data in each cluster to obtain the final clustering result.

8. The method according to claim 1, characterized in that The re-clustering of the low-dimensional feature vectors to obtain the final clustering result includes: Obtaining the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the initial clustering result; Performing density clustering on the low-dimensional feature vectors corresponding to the cluster centers of each cluster to obtain the final clustering result.

9. A data clustering device, characterized in that, The device includes: A high-dimensional vector acquisition module, configured to acquire high-dimensional feature vectors corresponding to text data to be clustered, where the high-dimensional feature vectors include semantic features of multiple dimensions; An initial clustering module, configured to perform initial clustering on the text data according to semantic features, and determine the cluster labels corresponding to the high-dimensional feature vectors according to the initial clustering result; when performing initial clustering, use a single-pass clustering algorithm as a framework, introduce a cluster structure of potential clusters and isolated clusters to implement initial clustering, where the potential cluster is a cluster in the cluster set with a text quantity greater than the preset quantity, and the isolated cluster is a cluster in the cluster set with a text quantity less than or equal to the preset quantity. When the text quantity in the isolated cluster is greater than the preset quantity, convert the isolated cluster into a potential cluster. The cluster set is formed by clustering historical text data, and the historical text data is the text data clustered before the current text data to be clustered; A data dimensionality reduction module, configured to replace the semantic features of multiple dimensions in the high-dimensional feature vectors with a one-dimensional cluster label to form low-dimensional feature vectors; A re-clustering module, configured to re-cluster the low-dimensional feature vectors corresponding to the cluster centers of each cluster in the initial clustering result or the text data of each cluster in the initial clustering result, use a stream clustering algorithm to implement data dimensionality reduction, and then use a batch clustering algorithm to implement clustering to obtain the final clustering result.

10. An electronic device, characterized in that, Including a processor and a memory, where the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-8 is run.

11. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-8 is run.

Citation Information

Patent Citations

  • Low-cost classification and clustering processing method for massive texts

    CN110377737A

  • Text processing method and device, electronic equipment and computer readable storage medium

    CN111737461A