Clustering method, system, device and storage medium based on density radius

By calculating the adjacent distance and density radius, automatically setting parameters and combining similar hash neural network and distance calculation, the problem that manually setting parameters in the existing technology is difficult to adapt to data of different distribution shapes is solved, and efficient clustering effect and multi-mapping are achieved.

CN114048318BActive Publication Date: 2025-09-16CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111430655.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-09-16
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

Existing clustering methods require manual input of setting parameters, which makes it difficult to adapt to data with different distribution shapes, resulting in poor clustering results.

Method used

By calculating the adjacency distance and density radius between cluster data, automatically setting parameters, using density radius for clustering processing, combining similar hash neural network and distance calculation method, the multi-mapping and deduplication judgment of clusters are realized.

Benefits of technology

It achieves the goal of improving clustering effects without the need for manual parameter input, is applicable to data with different shapes and distributions, and improves the multi-mapping and accuracy of clustered data in clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048318B_ABST
    Figure CN114048318B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence and provides a density radius-based clustering method, system, device, and storage medium. The method comprises: obtaining a sample data set, first cluster quantity data, and a cluster set, wherein the sample data set includes multiple cluster data; calculating the distance between any two cluster data to obtain multiple adjacent distance data; calculating density radius data on the adjacent distance data based on first sorting information and the first cluster quantity data; performing clustering processing with each cluster data as the center based on the density radius data and the adjacent distance data to obtain multiple clusters; when a cluster cluster meets a preset deduplication joining condition, adding the cluster cluster to the cluster set; and when the cluster set meets a preset cluster termination condition, outputting the cluster set. The present invention can automatically calculate the density radius for cluster data with different shapes, realize multi-mapping of cluster data in the cluster cluster, and improve the clustering effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a clustering method, system, device and storage medium based on density radius. Background Art

[0002] Cluster analysis refers to the analytical process of grouping a collection of physical or abstract objects into multiple classes consisting of similar objects. The goal of cluster analysis is to classify mobile data based on similarity, and it has a wide range of applications in unsupervised tasks in the field of Natural Language Processing (NLP). Clustering is the process of classifying data into different classes or clusters, so objects in the same cluster have great similarities, while objects in different clusters have great differences. The clustering methods in related technologies require manual input of setting parameters and are sensitive to the setting parameters. Therefore, it is difficult to determine the setting parameters for data with different distribution shapes, resulting in poor clustering results. Summary of the Invention

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] The embodiments of the present invention provide a density radius-based clustering method, system, device, and storage medium, which can automatically set parameters and improve the clustering effect for data with different distribution shapes.

[0005] In a first aspect, an embodiment of the present invention provides a clustering method based on density radius, the method comprising:

[0006] Acquire a sample data set, first cluster quantity data, and a cluster set, wherein the sample data set includes a plurality of cluster data;

[0007] Calculating the distance between any two cluster data to obtain multiple adjacent distance data;

[0008] Calculating density radius data from the adjacency distance data according to first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacency distance data;

[0009] Performing clustering processing with each of the cluster data as a center according to the density radius data and the adjacent distance data to obtain a plurality of clusters;

[0010] When the cluster meets the preset deduplication joining condition, the cluster is added to the cluster set;

[0011] When the cluster set meets the preset clustering termination condition, the cluster set is output.

[0012] According to some embodiments of the present invention, calculating the distance between any two cluster data to obtain a plurality of adjacent distance data includes:

[0013] Perform weighted processing on each cluster data according to word frequency-inverse text frequency TFIDF to obtain a weight value;

[0014] Importing the clustering data and the weight value into a similar hash neural network model to obtain transformed data;

[0015] According to the Hamming distance, the distance between any one of the conversion data and each of the remaining conversion data is calculated and processed to obtain a plurality of adjacent distance data.

[0016] The importance of clustered data within the sample dataset is assessed using TFIDF to determine the corresponding weights. The clustered data and corresponding weights are then fed into a similarity hashing neural network model, which outputs the transformed data. The Hamming distance between the transformed data and other transformed data is calculated to obtain the adjacency distance data, which can improve the accuracy of the adjacency distance data.

[0017] According to some embodiments of the present invention, the first cluster quantity data is obtained by the following steps:

[0018] Obtaining second cluster quantity data according to the preset category data and the cluster data;

[0019] Calculating the adjacency distance data according to the first sorting information and the second cluster quantity data to obtain a plurality of first density radius data;

[0020] Calculating the first density radius data according to second sorting information and a preset sorting threshold to obtain a density radius threshold, wherein the second sorting information is obtained by sorting the cluster data based on the first density radius data;

[0021] First cluster quantity data is obtained according to the adjacency distance data and the density radius threshold.

[0022] The first cluster quantity data is obtained by calculating and processing the preset category data, cluster data and adjacent distance data. There is no need to manually input the setting parameters. Cluster data with different shapes of distribution can be clustered to improve the clustering effect.

[0023] According to some embodiments of the present invention, obtaining first cluster quantity data according to the adjacency distance data and the density radius threshold includes:

[0024] Obtaining third cluster quantity data according to the adjacency distance data and the density radius threshold, wherein the third cluster quantity data includes a plurality of third cluster quantity data, and the third cluster quantity data corresponds one-to-one to the cluster data;

[0025] The third cluster quantity data is processed according to the third sorting information and the preset quantity condition to obtain the first cluster quantity data, and the third sorting information is obtained by sorting the third cluster quantity data.

[0026] Comparing the adjacency distance data with the density radius threshold to obtain third cluster quantity data, sorting the third cluster quantity data and performing calculation processing using a preset quantity condition to obtain the first cluster quantity data can improve the accuracy of the first cluster quantity data and enhance the clustering effect.

[0027] According to some embodiments of the present invention, clustering is performed based on the density radius data and the adjacency distance data, with each cluster data as a center, to obtain a plurality of clusters, including:

[0028] According to the fourth sorting information, the data to be clustered are clustered with each of the clustering data as the center in turn to obtain multiple cluster clusters; wherein, the fourth sorting information is obtained by sorting the clustering data based on the density radius data; the data to be clustered are the remaining clustering data corresponding to the adjacent distance data that is smaller than the density radius data.

[0029] The cluster data is sorted based on the density radius data, and clustered with each cluster data as the center in turn to obtain multiple clusters. In this way, one cluster data can exist in multiple clusters, realizing the multi-mapping of cluster data in cluster clusters.

[0030] According to some embodiments of the present invention, when the cluster satisfies a preset deduplication joining condition, adding the cluster to the cluster set includes:

[0031] Acquire a cluster center candidate set, wherein the cluster center candidate set includes all the cluster data, and the cluster data in the cluster center candidate set are arranged based on the density radius data;

[0032] According to the cluster center candidate set, the cluster set and the cluster clusters are processed in sequence using a distance-based similarity calculation method to obtain similarity data;

[0033] When the similarity data is less than a preset deduplication threshold, the cluster is added to the cluster set.

[0034] By sorting the cluster data in the cluster center candidate set, similarity calculations are performed on the corresponding cluster clusters and cluster sets in turn, and duplicate addition judgments are performed so that the cluster clusters corresponding to the cluster data with priority ranking can be entered into the cluster cluster first, thereby improving the clustering effect.

[0035] According to some embodiments of the present invention, the distance-based similarity calculation method includes at least one of the following types:

[0036] Euclidean distance calculation method;

[0037] Cosine distance calculation method;

[0038] Hamming distance calculation method;

[0039] Jaccard distance calculation method.

[0040] By using or combining Euclidean distance, cosine distance, Hamming distance, and Jaccard distance to calculate the similarity between two clusters, the accuracy of the similarity data can be improved, thereby improving the clustering effect.

[0041] According to some embodiments of the present invention, the preset clustering termination condition includes:

[0042] The number of clusters in the cluster set is equal to the preset number of clusters;

[0043] or,

[0044] The cluster radius data in the cluster set is greater than a preset radius threshold, and the cluster radius data is the density radius data corresponding to the cluster data in the cluster set.

[0045] When the number of clusters or the cluster radius data in the cluster set reaches the preset threshold, it is considered that the clustering termination condition is met, and the cluster cluster is output as the clustering result to avoid exceeding the set requirements and affecting the clustering effect.

[0046] According to some embodiments of the present invention, the method further comprises:

[0047] Acquire data to be labeled, where the data to be labeled comes from the cluster in the cluster set;

[0048] Performing category labeling processing on the data to be labeled to obtain label data;

[0049] The clusters are aggregated according to the label data to obtain a cluster set.

[0050] By extracting cluster data from the cluster set and labeling the categories, we can obtain label data. According to the label data in each cluster, we can aggregate the clusters to obtain a cluster set, thus improving the clustering effect.

[0051] In a second aspect, an embodiment of the present invention provides a density radius-based clustering system, including:

[0052] A sample acquisition module, configured to acquire a sample data set, first cluster quantity data, and a cluster set, wherein the sample data set includes a plurality of cluster data;

[0053] An adjacency distance calculation module is used to calculate the distance between any two cluster data to obtain a plurality of adjacency distance data;

[0054] a density radius calculation module, configured to calculate density radius data from the adjacency distance data according to first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacency distance data;

[0055] A cluster analysis module, configured to perform clustering processing based on the density radius data and the adjacent distance data, with each cluster data as the center, to obtain a plurality of clusters;

[0056] A deduplication judgment module, configured to add the cluster to the cluster set when the cluster meets a preset deduplication joining condition;

[0057] The cluster termination module is configured to output the cluster set when the cluster set meets a preset cluster termination condition.

[0058] In a third aspect, an embodiment of the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the density radius-based clustering method of the first aspect described above is implemented.

[0059] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the density radius-based clustering method of the first aspect.

[0060] According to the embodiment of the present invention, the clustering method based on density radius has at least the following beneficial effects: the distance between any two cluster data is calculated for all cluster data to obtain multiple adjacent distance data. Since the adjacent distance data corresponds to the cluster data, the cluster data can be sorted based on the adjacent distance data to obtain first sorting information, and the adjacent distance data are calculated in sequence according to the sorting information until the first cluster quantity data is met to obtain density radius data, thereby being able to automatically calculate the density radius of the data sample without manually inputting parameters. With each cluster data as the center, by comparing the density radius data and the adjacent distance data, the cluster data that meet the comparison result are clustered to obtain multiple cluster clusters, so that one cluster data can exist in multiple cluster clusters, realizing the multi-mapping property of cluster data in the cluster cluster. When a cluster cluster meets the preset deduplication joining condition, it is considered that there is no similar cluster cluster in the cluster set, and the cluster cluster can be added to the cluster set until the cluster set meets the preset cluster termination condition, and the cluster set is output as the clustering result. Therefore, the clustering method based on density radius can automatically calculate the density radius according to cluster data distributed in different shapes without manual input, realizing the multi-mapping of cluster data in cluster clusters and improving the clustering effect.

[0061] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.

[0063] Figure 1 is a flow chart of a density radius-based clustering method provided by an embodiment of the present invention;

[0064] Figure 2 yes Figure 1 Schematic diagram of the specific implementation process of step S200;

[0065] Figure 3 This is a schematic diagram of a specific process for forming first cluster quantity data provided by an embodiment of the present invention;

[0066] Figure 4 yes Figure 3 Schematic diagram of the specific implementation process of step S140;

[0067] Figure 5is a flow chart of a density radius-based clustering method provided by another embodiment of the present invention;

[0068] Figure 6 yes Figure 1 Schematic diagram of the specific implementation process of step S500;

[0069] Figure 7 yes Figure 1 Schematic diagram of the specific implementation process after step S600;

[0070] Figure 8 1 is a schematic structural diagram of a density radius-based clustering system provided in an embodiment of the present invention;

[0071] Figure 9 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0073] It should be noted that although the functional modules are divided in the module diagrams and the logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than the module division in the module or the order in the flowcharts. The terms "first," "second," etc. in the specification, claims, and drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0074] The present invention relates to artificial intelligence and provides a clustering method based on density radius, which obtains a sample data set, first cluster quantity data and a cluster set, wherein the sample data set includes multiple cluster data; calculates the distance between any two cluster data to obtain multiple adjacent distance data; calculates density radius data for the adjacent distance data based on first sorting information and first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacent distance data; performs clustering processing with each cluster data as the center based on the density radius data and the adjacent distance data to obtain multiple cluster clusters; when the cluster cluster meets a preset deduplication joining condition, the cluster cluster is added to the cluster set; when the cluster set meets a preset cluster termination condition, the cluster set is output. Therefore, the clustering method based on density radius can realize automatic calculation of the density radius according to cluster data distributed in different shapes without manual input, thereby realizing the multi-mapping property of the cluster data in the cluster cluster and improving the clustering effect.

[0075] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0076] It's important to note that AI also involves segmenting a dataset into classes or clusters based on specific criteria, such as distance criteria, to maximize the similarity of data objects within a cluster and the diversity of data objects in different clusters. This means that after clustering, data of the same class is grouped together as closely as possible, while different data is separated as much as possible.

[0077] Cluster analysis is a statistical analysis method for studying classification problems and is also a key algorithm for data mining. Cluster analysis consists of several patterns, typically a vector of metrics or a point in multidimensional space. Cluster analysis is based on similarity; patterns within a cluster are more similar than patterns in different clusters.

[0078] Clustering has a wide range of applications. For example, in business, clustering can help market analysts distinguish different consumer groups within a database and summarize the consumption patterns or habits of each consumer category. As a module in data mining, it can be used as a standalone tool to discover deeper information distributed within a database, summarize the characteristics of each category, or focus on a specific category for further analysis. Clustering analysis can also serve as a preprocessing step for other analysis algorithms within data mining.

[0079] Reference Figure 1 , Figure 1 The flowchart of the density radius-based clustering method provided by an embodiment of the present invention is shown. The density radius-based clustering method includes but is not limited to the following steps:

[0080] Step S100, obtaining a sample data set, first cluster quantity data and a cluster set, where the sample data set includes a plurality of cluster data;

[0081] Step S200, calculating the distance between any two cluster data to obtain a plurality of adjacent distance data;

[0082] Step S300: Calculate density radius data from the adjacency distance data based on the first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacency distance data;

[0083] Step S400, performing clustering processing with each cluster data as the center according to the density radius data and the adjacent distance data to obtain multiple clusters;

[0084] Step S500: When a cluster satisfies a preset deduplication joining condition, the cluster is added to the cluster set;

[0085] Step S600: When the cluster set meets the preset clustering termination condition, the cluster set is output.

[0086] It is understandable that, obtain sample data set, wherein sample data set includes multiple cluster data.Sample data set can be multiple articles, and cluster data then is corresponding to the text content of each article, and multiple articles are clustered by the text content of article.Calculate the distance between any two cluster data, obtain multiple adjacency distance data, namely select a cluster data as target data, calculate the distance between target data and remaining cluster data, obtain multiple adjacency distance data about target data, until obtain the adjacency distance data about all cluster data.For example, cluster data has 10, then each cluster data and its adjacent data are 9, therefore, each cluster data has 9 adjacency distance data corresponding thereto.In addition, can also construct and draw corresponding matrix according to adjacency distance data and cluster data, for example, cluster data has 10, then construct and draw 10th order matrix, thereby be conducive to carry out subsequent clustering process, improve processing efficiency.

[0087] Since the adjacency distance data corresponds to the cluster data, the adjacency distance data can represent the distance between the two cluster data, that is, the degree of similarity between the two cluster data. According to the adjacency distance data, the cluster data is sorted in ascending order to obtain the first sorting information. Therefore, a cluster data can be selected as the target data, and the target data is sorted in ascending order according to the distance between the target data and the other cluster data to obtain the first sorting information, that is, the cluster data is sorted in ascending order according to the numerical value of the adjacency distance data. If the cluster data has a higher similarity with the target data, the higher the corresponding ranking, the higher the probability of being in the same cluster cluster as the target data. The cluster data with higher similarity is selected for density radius calculation, and the accuracy of the obtained density radius data is higher. Therefore, the first cluster quantity data is obtained, which is used to determine the number of selected cluster data. According to the first sorting information, that is, according to the numerical value of the adjacency distance data, it is arranged from small to large, and the adjacency distance data is selected in sequence until the number of selected adjacency distance data reaches the number of the first cluster quantity data. Thus, the density radius data is obtained by calculation based on the selected adjacency distance data. Therefore, the density radius data can be automatically calculated based on the clustering data, without the need for customers to manually select the density radius in advance. The appropriate density radius can be selected for clustering data with different shapes, thereby improving the clustering effect.

[0088] Taking each cluster data as the center and its corresponding density radius data as the cluster boundary, the remaining cluster data is judged according to the adjacent distance data, and the cluster data within the cluster boundary is aggregated to obtain multiple cluster clusters, that is, each cluster data has a cluster cluster centered on itself, and the number of cluster clusters is the same as the number of cluster data, so that one cluster data can exist in multiple cluster clusters, so that the cluster data has multi-mapping in the cluster cluster and can be applied to a variety of situations.

[0089] Since each cluster data can exist in multiple clusters, there will be repeated clustering, resulting in multiple clusters being similar or identical, which affects the clustering effect. Therefore, it is necessary to obtain a cluster set and perform deduplication judgment on the cluster clusters, that is, to compare the similarity between the cluster cluster and the cluster clusters in the cluster set. When the cluster cluster meets the preset deduplication joining conditions, it is considered that there is no cluster cluster in the cluster set that is similar or identical to the current cluster cluster, and the cluster cluster can be added to the cluster set. When the cluster set meets the preset termination conditions, it is considered that the clustering processing of the cluster data is completed, and the cluster set is output as the clustering result. Therefore, the density radius can be automatically calculated without manual input, avoiding the pre-set density radius from being unable to match the cluster data, so that it can be applied to the situation where the cluster data is irregularly distributed, and at the same time realize the multi-mapping of the cluster data in the cluster cluster, thereby improving the clustering effect.

[0090] Reference Figure 2 , Figure 1 Step S200 in the illustrated embodiment includes but is not limited to the following steps:

[0091] Step S210, performing weighted processing on each cluster data according to word frequency-inverse text frequency TFIDF to obtain a weight value;

[0092] Step S220, importing the clustering data and weight values ​​into a similar hash neural network model to obtain transformed data;

[0093] Step S230 : calculating the distance between any one of the converted data and each of the remaining converted data based on the Hamming distance to obtain a plurality of adjacent distance data.

[0094] It is understood that Term Frequency-Inverse Document Frequency (TF-IDF) is a commonly used weighting technique for information retrieval and data mining. For example, clustered data can be text data. To cluster text data, topic mining is required. Therefore, the text data needs to be segmented. If a word or phrase has a high frequency (TF) in one article (i.e., clustered data) and rarely appears in other articles, it is considered to have good class distinction ability and is suitable for cluster data classification. Term Frequency (TF) indicates the frequency of a term in the first document. The fewer documents containing the first term, the greater the IDF, indicating that the first term has good class distinction ability. Therefore, TFIDF is used to weight each cluster data to obtain a corresponding weight value. The cluster data and the corresponding weight values ​​are used as input to a similarity hashing neural network model, which performs calculations to obtain transformed data, namely, a hash code for each cluster data, which can be used to calculate the distance between two clusters using the Hamming distance. By calculating the Hamming distance between any two hash codes, the adjacency distance data of the corresponding clustering data is obtained.

[0095] Reference Figure 3 , Figure 3 The first cluster quantity data can be obtained by the following steps:

[0096] Step S110, obtaining second cluster quantity data according to the preset category data and cluster data;

[0097] Step S120, calculating the adjacent distance data according to the first sorting information and the second cluster quantity data to obtain a plurality of first density radius data;

[0098] Step S130: Calculate the first density radius data according to the second sorting information and a preset sorting threshold to obtain a density radius threshold, where the second sorting information is obtained by sorting the cluster data based on the first density radius data.

[0099] Step S140: Obtain first cluster quantity data according to the adjacency distance data and the density radius threshold.

[0100] It is understandable that, according to the preset category data, each cluster data is divided into multiple categories, thereby obtaining the second cluster quantity data. The preset category data can be the number of categories of the required cluster data, and the second cluster quantity data can be the average number of cluster data for each category. For example, according to the preset category data, all cluster data can be divided into 5 categories, and there are 10 cluster data. Therefore, each category has an average of 2 cluster data. The density radius data of each cluster cluster is calculated based on the second cluster quantity data, so that each cluster cluster can have two cluster data, so that the number of cluster clusters reaches 5, which matches the number of categories corresponding to the preset category data and meets the requirements. Therefore, according to the first sorting information, the corresponding adjacent distance data are selected in sequence until the number of adjacent distance data reaches the number corresponding to the second cluster quantity data. The average value of all the selected adjacent distance data is calculated to obtain the first density radius data. The first density radius data is calculated for each cluster data respectively to obtain multiple first density radius data, and the cluster data is sorted in descending order based on the first density radius data to obtain the second sorting information. Since the smaller the density radius, the fewer cluster data may be included in the cluster, it is difficult to reflect the similarity of multiple cluster data. The larger the density radius, the more cluster data may be included in the cluster, and it is difficult to reflect the difference between multiple cluster data. Therefore, the appropriate density radius range is determined by presetting the sorting threshold to improve the clustering effect. According to the second sorting information and the preset sorting threshold, the first density radius data corresponding to the preset sorting threshold is selected for calculation to obtain the density radius threshold. For example, there are 10 cluster data, and the cluster data is numbered based on the second sorting information. The preset sorting threshold is 0.8, and the first density radius threshold corresponding to the cluster data with sequence number 8 is selected as the density radius threshold. The density radius threshold and the corresponding adjacent distance data are compared, and the number of adjacent distance data less than the density radius threshold is recorded to obtain the first cluster quantity data. For example, the first density radius threshold corresponding to the cluster data with serial number 8 is selected as the density radius threshold, and the adjacent distance data corresponding to the cluster data with serial number 8 is selected for comparison. The data in the adjacent distance data that is smaller than the density radius threshold is selected to obtain the number of adjacent distance data that is smaller than the density radius, that is, the first cluster quantity data, so that the appropriate number of clusters can be selected according to the cluster data, and the appropriate density radius range can be determined. This is suitable for situations where cluster data is irregularly distributed, reflects the similarity of cluster data in the same cluster cluster, and reflects the difference of cluster data in different cluster clusters, thereby improving the clustering effect.

[0101] Reference Figure 4 , Figure 3 Step S140 in the illustrated embodiment includes but is not limited to the following steps:

[0102] Step S141: obtaining third cluster quantity data according to the adjacency distance data and the density radius threshold, wherein the third cluster quantity data includes a plurality of cluster quantity data, and the third cluster quantity data corresponds one-to-one to the cluster data;

[0103] Step S142: Process the third cluster quantity data according to the third sorting information and the preset quantity condition to obtain the first cluster quantity data. The third sorting information is obtained by sorting the third cluster quantity data.

[0104] It can be understood that all the adjacent distance data corresponding to a cluster data are compared with the density radius threshold respectively, and the number of adjacent distance data less than the density radius threshold is recorded to obtain the third cluster quantity data corresponding to the cluster data. The adjacent distance data of all cluster data are compared to obtain a plurality of third cluster quantity data, wherein the third cluster quantity data corresponds to the cluster data one by one. The third cluster quantity data is sorted according to the numerical value to obtain the third sorting information. According to the third sorting information, the third cluster quantity data that meets the preset quantity condition is selected as the first cluster quantity data. For example, the third cluster quantity data is arranged in descending order according to the numerical value, and the preset quantity condition is 50%, that is, the third cluster quantity data ranked 50% is selected from the third sorting information as the first cluster quantity data. Therefore, according to the cluster data and its adjacent distance data, the appropriate first cluster quantity data is selected, so that it can be applicable to the situation where the cluster data is irregularly distributed, and the set parameters are automatically calculated without manual input, thereby improving the clustering effect.

[0105] Reference Figure 5 , Figure 5 FIG2 shows a flow chart of a clustering method based on density radius provided by another embodiment of the present invention. Figure 1 Step S400 in the illustrated embodiment includes but is not limited to the following steps:

[0106] Step S410, according to the fourth sorting information, cluster the data to be clustered with each cluster data as the center in turn to obtain multiple cluster clusters; wherein the fourth sorting information is obtained by sorting the cluster data based on the density radius data; the data to be clustered is the remaining cluster data corresponding to the adjacent distance data that is less than the density radius data.

[0107] It is understood that the cluster data is sorted based on the density radius data to obtain fourth sorting information, wherein the sorting can be performed in ascending order based on the numerical value of the density radius data. That is, the smaller the density radius of the cluster data, the higher the probability that the cluster cluster formed by it will join the cluster set. Based on the fourth sorting information, the data to be clustered are clustered with each cluster data as the center, where the data to be clustered is the cluster data within the density radius corresponding to the cluster data serving as the cluster center. For example, the first sample data is selected as the cluster center, and the cluster data within the density radius of the first sample data include the second sample data and the third sample data. Thus, the second sample data and the third sample data are clustered with the first sample data as the cluster center to form a first cluster. Clustering is performed sequentially with each cluster data as the center to form clusters. Therefore, a cluster data can exist in multiple clusters, achieving multi-mapped cluster data within the cluster clusters. Therefore, clustering can be performed based on the density radius corresponding to different cluster data, thereby selecting an appropriate density radius for irregularly distributed cluster data and improving the clustering effect.

[0108] Reference Figure 6 , Figure 1 Step S500 in the illustrated embodiment includes but is not limited to the following steps:

[0109] Step S510, obtaining a cluster center candidate set, wherein the cluster center candidate set includes all cluster data, and the cluster data in the cluster center candidate set are arranged based on density radius data;

[0110] Step S520: Based on the candidate set of cluster centers, the cluster set and the cluster clusters are processed in sequence using a distance-based similarity calculation method to obtain similarity data;

[0111] Step S530: When the similarity data is less than a preset deduplication threshold, the cluster is added to the cluster set.

[0112] It is understandable that, since a cluster data can exist in multiple clusters, there will be situations where multiple clusters are similar or identical, which affects the clustering effect. Perform deduplication judgment on cluster clusters and cluster sets to determine whether the current cluster cluster and the cluster clusters that have been added to the cluster set are similar or identical. If the current cluster cluster is not similar to the cluster clusters in the cluster set, the cluster cluster will be added to the cluster set. The current cluster cluster and the cluster clusters in the cluster set can be judged and processed based on the distance-based similarity calculation method to obtain similarity data, thereby judging whether they are similar based on the distance between the current cluster cluster and the cluster clusters in the cluster set. When the similarity data is less than the preset deduplication threshold, it is considered that the distance between the current cluster cluster and the cluster clusters in the cluster set is large, and the similarity between the current cluster cluster and the cluster clusters in the cluster set is low. Therefore, the cluster cluster is added to the cluster set.

[0113] It is understandable that the cluster center candidate set includes all cluster data, that is, all cluster data are the centers of cluster clusters, and the arrangement order of cluster data in the cluster center candidate set is obtained according to the corresponding density radius data, that is, the density radius data is sorted in ascending order to obtain the arrangement order of cluster data in the cluster center candidate set. Since the smaller the density radius of the cluster, the smaller the number of cluster data contained in the cluster, the higher the clustering accuracy. Therefore, the sorted cluster data is traversed in turn, and the cluster cluster centered on it is compared with the cluster clusters in the cluster set for similarity. The similarity between cluster clusters can be processed by a distance-based similarity calculation method, and the distance between each cluster cluster is calculated as the similarity data.

[0114] It is understandable that the distance-based similarity calculation method may include the Euclidean distance calculation method, the cosine distance calculation method, the Hamming distance calculation method, and the Jaccard distance calculation method. Among them, one calculation method or a combination of multiple calculation methods can be used for calculation. For example, to calculate the similarity between the first cluster and the second cluster, the texts of the cluster centers of the two clusters can be selected respectively, and the keywords are extracted, which are correspondingly recorded as set A and set B. The Jaccard distance calculation method is used to calculate the similarity between set A and set B, and the similarity is obtained by calculating the ratio of the size of the intersection of set A and set B to the size of the union of set A and set B. The specific formula of the Jaccard distance calculation method is as follows:

[0115]

[0116] Among them, when set A and set B are both empty sets, J(A,B) is defined as 1.

[0117] For example, when the value of J(A, B) is less than 0.3, the two clusters can be considered dissimilar. Alternatively, the Hamming distance calculation method can be used to perform a similarity determination again. If the Hamming distance calculation method determines that the cluster is dissimilar to all clusters in the cluster set, the cluster can be added to the cluster set. If the Hamming distance calculation method determines that the cluster is similar to a cluster in the cluster set, it is considered that there is a duplicate cluster and the cluster is not added to the cluster set. Therefore, the similarity between two clusters can be calculated using multiple similarity calculation methods to improve the accuracy of the similarity calculation, avoid adding the current cluster to the cluster set because it is a duplicate cluster, and improve the clustering effect.

[0118] It is understood that when a cluster cluster is added to a cluster set, a termination judgment is performed on the cluster set to determine whether the cluster set meets the preset cluster termination condition. When the cluster set meets the preset cluster termination condition, the cluster set is output as the clustering result. The preset cluster termination condition includes that the number of cluster clusters in the cluster set is equal to the preset number of clusters, or that the cluster radius data in the cluster set is greater than a preset radius threshold, where the cluster radius data is the density radius data corresponding to the cluster data in the cluster set. That is, when the number of cluster clusters in the cluster set is equal to the preset number of clusters, the cluster set is considered to meet the preset cluster termination condition, and the cluster set is output as the clustering result. When the density radius data corresponding to the cluster data in the cluster set is greater than the preset radius threshold, the cluster set is considered to meet the preset cluster termination condition, and the cluster set is output as the clustering result. For example, if the preset number of clusters is set to 5 cluster clusters, then when the current cluster cluster is added to the cluster set and the cluster set has 5 cluster clusters, the cluster set is considered to meet the preset cluster termination condition, and the 5 cluster clusters in the cluster set are output as the clustering result. For another example, if the preset radius threshold is 30, then when the density radius of the current cluster is 35 and the cluster is added to the cluster set, the cluster set is considered to meet the preset cluster termination condition, and the clusters added to the cluster set are output, where the cluster set includes the cluster with a density radius of 35. Therefore, by setting the preset cluster termination condition, over-clustering can be avoided and usage requirements can be met.

[0119] Reference Figure 7 , Figure 7 The density radius based clustering method includes but is not limited to the following steps:

[0120] Step S700: obtaining data to be labeled, where the data to be labeled comes from a cluster in a cluster set;

[0121] Step S800: performing category labeling on the data to be labeled to obtain label data;

[0122] Step S900: Aggregate the clusters according to the label data to obtain a cluster set.

[0123] It is understood that when a cluster cluster satisfies the preset cluster termination conditions and the cluster set is output, cluster data is extracted from the cluster clusters in the cluster set as data to be labeled. This extraction ratio can be set based on a preset ratio, such as 5% of the total number of cluster data in the cluster set. The extracted data to be labeled is then subjected to category labeling, i.e., the data to be labeled is labeled with a category label to obtain labeled data with the category label. Category labels can be used to distinguish the category to which the labeled data belongs. The category labels of the labeled data can be used to identify clusters of the same category, and clusters are aggregated to form a cluster set of the same category. For example, if the output cluster set contains 10 clusters, data to be labeled is extracted based on 5% of the total number of cluster data in the cluster set and then labeled with a category label. The labeled category labels include a total of five categories. When the category label of more than half of the labeled data in a cluster is labeled as the first category, the cluster is added to the first category cluster set. Therefore, cluster clusters can be automatically aggregated, improving the clustering effect.

[0124] Reference Figure 8 , Figure 8 FIG. 8 is a schematic structural diagram of a density radius-based clustering system 800 provided in an embodiment of the present invention.

[0125] The sample acquisition module 810 is used to acquire a sample data set, first cluster quantity data and a cluster set, where the sample data set includes a plurality of cluster data.

[0126] The adjacent distance calculation module 820 is used to calculate the distance between any two cluster data to obtain multiple adjacent distance data.

[0127] The density radius calculation module 830 is used to calculate the density radius data from the adjacent distance data according to the first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacent distance data.

[0128] The cluster analysis module 840 is used to perform clustering processing based on the density radius data and the adjacent distance data, with each cluster data as the center, to obtain multiple clusters.

[0129] The deduplication judgment module 850 is configured to add a cluster to the cluster set when the cluster meets the preset deduplication joining condition.

[0130] The clustering termination module 860 is configured to output a cluster set when the cluster set meets a preset clustering termination condition.

[0131] In addition, the adjacency distance calculation module 820 includes:

[0132] The weight value calculation module 821 is used to perform weighted processing on each cluster data according to the word frequency-inverse text frequency TFIDF to obtain a weight value.

[0133] The conversion data calculation module 822 is used to import the clustering data and weight values ​​into the similarity hash neural network model to obtain conversion data.

[0134] The distance data calculation module 823 is used to calculate the distance between any converted data and each remaining converted data according to the Hamming distance to obtain a plurality of adjacent distance data.

[0135] In addition, the density radius-based clustering system 800 further includes a cluster number calculation module 870, which includes:

[0136] The second cluster quantity calculation module 871 is configured to obtain second cluster quantity data according to the preset category data and cluster data.

[0137] The first density radius calculation module 872 is used to calculate the adjacent distance data according to the first sorting information and the second cluster quantity data to obtain multiple first density radius data.

[0138] The density radius threshold calculation module 873 is used to calculate the first density radius data according to the second sorting information and the preset sorting threshold to obtain the density radius threshold. The second sorting information is obtained by sorting the cluster data based on the first density radius data.

[0139] The first cluster quantity calculation module 874 is configured to obtain first cluster quantity data according to the adjacency distance data and the density radius threshold.

[0140] In addition, the first cluster number calculation module 874 includes:

[0141] The third cluster quantity calculation module 875 is used to obtain third cluster quantity data according to the adjacent distance data and the density radius threshold. The third cluster quantity data includes multiple third cluster quantity data, and the third cluster quantity data corresponds one-to-one to the cluster data.

[0142] The cluster quantity comprehensive calculation module 876 is used to process the third cluster quantity data according to the third sorting information and the preset quantity condition to obtain the first cluster quantity data. The third sorting information is obtained by sorting the third cluster quantity data.

[0143] In addition, cluster analysis module 840 is further configured to cluster the data to be clustered, based on the fourth sorting information, with each cluster data as the center, to obtain a plurality of clusters. The fourth sorting information is obtained by sorting the cluster data based on the density radius data. The data to be clustered is the remaining cluster data corresponding to the adjacent distance data that is less than the density radius data.

[0144] In addition, the deduplication determination module 850 includes:

[0145] The center candidate set module 851 is used to obtain a cluster center candidate set, wherein the cluster center candidate set includes all cluster data, and the cluster data in the cluster center candidate set is arranged based on density radius data.

[0146] The distance similarity calculation module 852 is used to perform a distance-based similarity calculation method on the cluster set and the cluster clusters in turn according to the cluster center candidate set to obtain similarity data.

[0147] The cluster adding module 853 is configured to add the cluster to the cluster when the similarity data is less than a preset deduplication threshold.

[0148] Reference Figure 9 , Figure 9 An electronic device 900 provided by an embodiment of the present invention is shown. The electronic device 900 includes a memory 910, a processor 920, and a computer program stored in the memory 910 and executable on the processor 920. When the processor 920 executes the computer program, the density radius-based clustering method in the above embodiment is implemented.

[0149] Memory 910, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the density radius-based clustering method in the above-mentioned embodiment of the present invention. Processor 920 implements the density radius-based clustering method in the above-mentioned embodiment of the present invention by executing the non-transitory software program and instructions stored in memory 910.

[0150] The memory 910 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data required to execute the density radius-based clustering method in the above embodiment, etc. In addition, the memory 910 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. It should be noted that the memory 910 may optionally include a memory remotely arranged relative to the processor 920, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0151] The non-transient software program and instructions required to implement the density radius based clustering method in the above embodiment are stored in the memory. When executed by one or more processors, the density radius based clustering method in the above embodiment is executed, for example, the above described Figure 1 Steps S100 to S600 of the method, Figure 2 Steps S210 to S230 of the method, Figure 3 Steps S110 to S140 of the method, Figure 4 Steps S141 to S142 of the method, Figure 5 Step S410 of the method, Figure 6 Steps S510 to S530 of the method, Figure 7 Method steps S700 to S900.

[0152] The present invention also provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are used to enable a computer to execute the density radius-based clustering method in the above embodiment, for example, to execute the above-described Figure 1 Steps S100 to S600 of the method, Figure 2 Steps S210 to S230 of the method, Figure 3 Steps S110 to S140 of the method, Figure 4 Steps S141 to S142 of the method, Figure 5 Step S410 of the method, Figure 6 Steps S510 to S530 of the method, Figure 7 Method steps S700 to S900.

[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0154] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0155] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in the technical field without departing from the spirit of the present invention.

Claims

1. A clustering method based on density radius, the method comprising: Acquire a sample data set, first cluster quantity data, and a cluster set, wherein the sample data set includes a plurality of cluster data, and the cluster data is text data; Perform weighted processing on each cluster data according to word frequency-inverse text frequency TFIDF to obtain a weight value; Importing the clustering data and the weight value into a similar hash neural network model to obtain transformed data; Calculating the distance between any one of the transformed data and each of the remaining transformed data according to the Hamming distance to obtain a plurality of adjacent distance data; Calculating density radius data from the adjacency distance data according to first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacency distance data; Performing clustering processing with each cluster data as a center according to the density radius data and the adjacent distance data to obtain multiple clusters; When the cluster meets the preset deduplication joining condition, the cluster is added to the cluster set; When the cluster set meets the preset clustering termination condition, output the cluster set; The first cluster quantity data is obtained by the following steps: Obtaining second cluster quantity data according to the preset category data and the cluster data; Calculating the adjacency distance data according to the first sorting information and the second cluster quantity data to obtain a plurality of first density radius data; Calculating the first density radius data according to second sorting information and a preset sorting threshold to obtain a density radius threshold, wherein the second sorting information is obtained by sorting the cluster data based on the first density radius data; First cluster quantity data is obtained according to the adjacency distance data and the density radius threshold.

2. The density radius-based clustering method according to claim 1, characterized in that: The obtaining of first cluster quantity data according to the adjacency distance data and the density radius threshold comprises: Obtaining third cluster quantity data according to the adjacency distance data and the density radius threshold, wherein the third cluster quantity data includes a plurality of third cluster quantity data, and the third cluster quantity data corresponds one-to-one to the cluster data; The third cluster quantity data is processed according to the third sorting information and the preset quantity condition to obtain the first cluster quantity data, and the third sorting information is obtained by sorting the third cluster quantity data.

3. The density radius-based clustering method according to claim 1, characterized in that: The clustering process is performed based on the density radius data and the adjacent distance data, with each cluster data as the center, to obtain multiple clusters, including: According to the fourth sorting information, the data to be clustered are clustered with each of the clustering data as the center in turn to obtain multiple cluster clusters; wherein, the fourth sorting information is obtained by sorting the clustering data based on the density radius data; the data to be clustered are the remaining clustering data corresponding to the adjacent distance data that is smaller than the density radius data.

4. The clustering method based on density radius according to claim 1, characterized in that: When the cluster meets the preset deduplication joining condition, adding the cluster to the cluster set includes: Acquire a cluster center candidate set, wherein the cluster center candidate set includes all the cluster data, and the cluster data in the cluster center candidate set are arranged based on the density radius data; According to the cluster center candidate set, the cluster set and the cluster clusters are processed in sequence using a distance-based similarity calculation method to obtain similarity data; When the similarity data is less than a preset deduplication threshold, the cluster is added to the cluster set.

5. The density radius-based clustering method according to claim 1, characterized in that: The preset clustering termination conditions include: The number of clusters in the cluster set is equal to the preset number of clusters; or, The cluster radius data in the cluster set is greater than a preset radius threshold, and the cluster radius data is the density radius data corresponding to the cluster data in the cluster set.

6. A clustering system based on density radius, characterized in that include: A sample acquisition module is used to acquire a sample data set, first cluster quantity data and a cluster set, wherein the sample data set includes a plurality of cluster data, and the cluster data is text data; an adjacency distance calculation module for weighting each of the clustered data according to term frequency-inverse text frequency (TFIDF) to obtain a weight value; importing the clustered data and the weight value into a similar hash neural network model to obtain transformed data; and calculating the distance between any one of the transformed data and each of the remaining transformed data according to the Hamming distance to obtain a plurality of adjacency distance data; a density radius calculation module, configured to calculate density radius data from the adjacency distance data according to first sorting information and the first cluster quantity data, wherein the first sorting information is obtained by sorting the cluster data based on the adjacency distance data; A cluster analysis module, configured to perform clustering processing based on the density radius data and the adjacent distance data, with each cluster data as the center, to obtain a plurality of clusters; A deduplication judgment module, configured to add the cluster to the cluster set when the cluster meets a preset deduplication joining condition; A cluster termination module, configured to output the cluster set when the cluster set meets a preset cluster termination condition; The first cluster quantity data is obtained by the following steps: Obtaining second cluster quantity data according to the preset category data and the cluster data; Calculating the adjacency distance data according to the first sorting information and the second cluster quantity data to obtain a plurality of first density radius data; Calculating the first density radius data according to second sorting information and a preset sorting threshold to obtain a density radius threshold, wherein the second sorting information is obtained by sorting the cluster data based on the first density radius data; First cluster quantity data is obtained according to the adjacency distance data and the density radius threshold.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the density radius-based clustering method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the density radius-based clustering method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Stream data clustering method based on density value dynamic change

    CN106203474A

  • Data clustering method and device

    CN110298371A