Data clustering method, device, storage medium and electronic device
By acquiring and updating the clustering center in the clustering model in the data clustering method, the problems of low data clustering efficiency and poor clustering results in the existing technology are solved, and clustering processing that automatically adapts to new data is realized, improving clustering efficiency and effect.
Patent Information
- Application Number
- CN202111677732.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In the prior art, data clustering methods cannot automatically meet the needs of new data clustering, resulting in low clustering efficiency and poor clustering results.
By obtaining the first optimal cluster cluster value predetermined by the original data of the data to be clustered, and performing secondary clustering processing after detecting the incremental data, the second optimal cluster value is obtained. The cluster center in the clustering model is updated based on the comparison results of the two, and the updated model is used for K-means clustering.
It realizes the clustering processing results that meet the needs of new data based on the new data characteristics, which improves the data clustering efficiency and clustering effect.
Smart Images

Figure CN114330584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data clustering method, device, storage medium and electronic device. Background Art
[0002] Data clustering or classification tasks are widely used algorithms in life and production. After analyzing entities and extracting feature information of entities of interest, clustering and classifying this information can greatly facilitate people's production and life. However, the data clustering methods in the existing technology, such as the K-means algorithm, can extract cluster centers and keep cluster categories unchanged, but cannot cope with the situation where new categories are contained in incremental data, and have certain limitations, such as Figure 1 As shown, if the number of categories in the incremental data changes, it is necessary to manually adjust the K value for re-clustering, and then use the manual intervention method to adjust the original classification according to the previous category discovery, which has low clustering efficiency and poor clustering results.
[0003] Currently, no effective solution has been proposed to address the problems of low clustering efficiency and poor clustering results in the prior art. Summary of the invention
[0004] The embodiments of the present invention provide a data clustering method, device, storage medium and electronic device to at least solve the technical problem of low clustering efficiency and poor clustering results caused by the inability of the data clustering method in the prior art to automatically meet the new data clustering requirements.
[0005] According to one aspect of an embodiment of the present invention, a data clustering method is provided, comprising: obtaining a first optimal clustering cluster value predetermined based on original data of the data to be clustered; after detecting incremental data of the data to be clustered, performing secondary clustering processing on the data to be clustered to obtain multiple second clustering index values, and selecting a second target clustering index value from the multiple second clustering index values; obtaining a second optimal clustering cluster value corresponding to the second target clustering index value; updating the cluster centers in the clustering model according to a comparison result of the first optimal clustering cluster value and the second optimal clustering cluster value, and performing K-means clustering processing on the data to be clustered using the updated clustering model to obtain a target clustering processing result.
[0006] Optionally, obtaining a first optimal clustering cluster value predetermined based on original data of the data to be clustered includes: performing a first clustering process on the original data to obtain multiple first clustering index values, and selecting a first target clustering index value from the multiple first clustering index values; obtaining a first optimal clustering cluster value corresponding to the first target clustering index value.
[0007] Optionally, after obtaining a first optimal cluster value predetermined based on the original data of the data to be clustered, the method further includes: performing K-means clustering processing on the original data using the first optimal cluster value to obtain a first clustering processing result.
[0008] Optionally, the cluster center in the clustering model is determined according to the comparison result of the above-mentioned first optimal cluster cluster value and the above-mentioned second optimal cluster cluster value, including: when the above-mentioned comparison result is that the above-mentioned first optimal cluster cluster value is less than the above-mentioned second optimal cluster cluster value, then a new cluster center is determined based on the above-mentioned incremental data and the initial cluster center, and the above-mentioned new cluster center and the above-mentioned initial cluster center are used as the initial cluster center for the next clustering processing; when the above-mentioned comparison result is that the above-mentioned first optimal cluster cluster value is greater than the above-mentioned second optimal cluster cluster value, it is determined that the distribution of the above-mentioned incremental data and the above-mentioned original data is inconsistent, and the clustering processing flow is canceled; when the above-mentioned comparison result is that the above-mentioned first optimal cluster cluster value is equal to the above-mentioned second optimal cluster cluster value, the cluster center determined by the first clustering processing of the above-mentioned original data will still be used as the initial cluster center for the next clustering processing.
[0009] Optionally, after using the newly added cluster center and the initial cluster center as the initial cluster center for the next clustering, the method further includes: calculating a third optimal cluster value based on the first optimal cluster cluster value and the second optimal cluster cluster value; according to the third optimal cluster cluster value, cyclically executing the step of selecting a distance target based on the incremental data and using the target as the initial cluster center for the next clustering to obtain a corresponding plurality of initial cluster centers; iterating the cluster center of the cluster model using the plurality of initial cluster centers to obtain the updated cluster model.
[0010] According to another aspect of an embodiment of the present invention, a data clustering device is also provided, including: a first acquisition module, used to obtain a first optimal clustering cluster value predetermined based on the original data of the data to be clustered; a first processing module, used to perform secondary clustering processing on the above-mentioned data to be clustered after detecting the incremental data of the above-mentioned data to be clustered, to obtain multiple second clustering index values, and select a second target clustering index value from the multiple above-mentioned second clustering index values; a second acquisition module, used to obtain a second optimal clustering cluster value corresponding to the above-mentioned second target clustering index value; a second processing module, used to update the cluster center in the clustering model according to the comparison result of the above-mentioned first optimal clustering cluster value and the above-mentioned second optimal clustering cluster value, and use the updated clustering model to perform K-means clustering processing on the above-mentioned data to be clustered to obtain the target clustering processing result.
[0011] Optionally, the above-mentioned first acquisition module also includes: a selection module, which is used to perform an initial clustering process on the above-mentioned original data to obtain multiple first clustering index values, and select a first target clustering index value from the multiple first clustering index values; a first acquisition submodule, which is used to obtain the first optimal clustering cluster value corresponding to the above-mentioned first target clustering index value.
[0012] Optionally, the device further includes: a second acquisition submodule, configured to perform K-means clustering processing on the original data using the first optimal clustering cluster value to obtain a first clustering processing result.
[0013] Optionally, the second processing module also includes: a comparison module, which is used to determine a new cluster center based on the incremental data and the initial cluster center when the comparison result is that the first optimal cluster cluster value is less than the second optimal cluster cluster value, and use the new cluster center and the initial cluster center as the initial cluster center for the next clustering processing; a first determination submodule, which is used to determine that the distribution of the incremental data and the original data are inconsistent and cancel the clustering processing flow when the comparison result is that the first optimal cluster cluster value is greater than the second optimal cluster cluster value; and a second determination submodule, which is used to still use the cluster center determined by the first clustering processing of the original data as the initial cluster center for the next clustering processing when the comparison result is that the first optimal cluster cluster value is equal to the second optimal cluster cluster value.
[0014] Optionally, the above-mentioned device also includes: a calculation module, which is used to calculate the third best clustering value based on the above-mentioned first best clustering cluster value and the above-mentioned second best clustering cluster value; a third acquisition submodule, which is used to cyclically execute the step of selecting a distance target based on the above-mentioned incremental data and using the above-mentioned target as the initial clustering center of the next clustering according to the above-mentioned third best clustering cluster value, and obtain the corresponding multiple initial clustering centers; a fourth acquisition submodule, which is used to iterate the clustering center of the above-mentioned clustering model using the multiple initial clustering centers to obtain the above-mentioned updated clustering model.
[0015] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing any one of the above-mentioned data clustering methods.
[0016] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute one of the data clustering methods described above.
[0017] In an embodiment of the present invention, a data clustering method is adopted, by obtaining a first optimal clustering cluster value predetermined based on the original data of the data to be clustered; after detecting the incremental data of the above-mentioned data to be clustered, performing secondary clustering processing on the above-mentioned data to be clustered to obtain multiple second clustering index values, and selecting a second target clustering index value from the multiple second clustering index values; obtaining the second optimal clustering cluster value corresponding to the above-mentioned second target clustering index value; according to the comparison result of the above-mentioned first optimal clustering cluster value and the above-mentioned second optimal clustering cluster value, updating the cluster center in the clustering model, and using the updated clustering model to perform K-means clustering processing on the above-mentioned data to be clustered to obtain the target clustering processing result, the purpose of determining the clustering processing result that meets the needs of the new data according to the characteristics of the new data is achieved, thereby realizing the technical effect of improving the data clustering efficiency and clustering effect, and further solving the technical problem of low clustering efficiency and poor clustering results caused by the inability of the data clustering method in the prior art to automatically meet the clustering needs of the new data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0019] Figure 1 is a flow chart of a data clustering method according to the prior art;
[0020] Figure 2 is a flow chart of a data clustering method according to an embodiment of the present invention;
[0021] Figure 3 is a schematic diagram of an optional system structure for implementing the above data clustering method according to an embodiment of the present invention;
[0022] Figure 4 is a flow chart of another optional data clustering method according to an embodiment of the present invention;
[0023] Figure 5 is a flow chart of another optional data clustering method according to an embodiment of the present invention;
[0024] Figure 6 is a schematic diagram of an application scenario of an optional data clustering method according to an embodiment of the present invention;
[0025] Figure 7 is a schematic diagram of an application scenario of an optional data clustering method according to an embodiment of the present invention;
[0026] Figure 8 It is a structural schematic diagram of a data clustering device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] First, to facilitate understanding of the embodiments of the present invention, some terms or nouns involved in the present invention are explained below:
[0030] Supervised classification and unsupervised clustering: In machine learning classification tasks, it is generally customary to distinguish algorithms based on whether or not sample labels are input for training. If sample labels are input during model training, that is, the algorithm that labels the sample categories during training to generate the model is called supervised classification. On the contrary, when unlabeled samples are provided during training, the classification algorithm that relies on the algorithm to automatically find the relationship between samples is called unsupervised clustering. In recent years, a series of classification methods that only input partially labeled samples have been developed, called semi-supervised learning.
[0031] Supervised classification: supervised learning gives the algorithm information about the sample category when the sample is input. This type of algorithm is generally more accurate than unsupervised algorithms. However, the input samples need to be manually labeled, which consumes a lot of manpower and material costs, and the task universality is poor. Representative classification algorithms include svm, linear discriminator, naive Bayes, decision tree classification, K nearest neighbor classification, etc.
[0032] Unsupervised clustering: Unsupervised clustering does not require input sample labels, but uses certain rules in the algorithm to maximize the number of clusters between clusters and minimize the number of clusters within clusters, so as to achieve the purpose of learning classification models. Representative algorithms include K-means, manifold learning, hierarchical clustering, DBSCAN, density clustering, covariance clustering, etc.
[0033] K-means (K-means clustering): K-means algorithm is a classic clustering algorithm. It first requires the user to set a K value, and the algorithm will cluster samples into K categories. First, the algorithm initializes K centers, and then repeats two steps until the centroid no longer changes. Step 1, calculate the distance from each sample to the centroid, and assign the sample to the category with the closest centroid; step 2, move the centroid to the center point (average value of each dimension) of this category of samples. The two steps are repeated until it stops.
[0034] Example 1
[0035] Clustering or classification tasks are widely used algorithms in life and production. After analyzing entities and extracting feature information of entities of interest, clustering and classifying this information can greatly facilitate people's production and life. For example, by learning from photos of cats and dogs, a model can be trained to recognize photos of cats or dogs. This model can be used to deal with subsequent problems, such as automatically using the model to distinguish whether it is a cat or a dog in some scenarios. Or in text recognition, the trained model can be used to distinguish the emotion or category of an article. There are countless kinds of data in life, and it is impossible for each category to have a large amount of labeled data for supervised learning to train models. Efficient, sustainable, and stable unsupervised methods are increasingly becoming one of the directions of popular research.
[0036] K-means algorithm is a classic clustering algorithm. It first requires the user to set a K value, and the algorithm will cluster the samples into K categories. First, the algorithm initializes K centers, and then repeats two steps until the center of mass does not change. Step 1, calculate the distance from each sample to the center of mass, and assign the sample to the category with the closest center of mass; step 2, move the center of mass to the center point of this category of samples (the average value of each dimension). The two steps are repeated until it stops. The constraint formula is as follows:
[0037]
[0038] In the face of the problem of multiple clustering, the traditional K-means has no continuity due to randomness, and the ability to automatically obtain the best K. The existing technical solution implementation flow chart is still as follows Figure 1As shown in the figure, when faced with incremental data, although the traditional model can extract cluster centers and keep cluster categories unchanged, it cannot cope with the situation where the incremental data contains new categories. It has certain limitations. If the number of categories in the incremental data changes, it is necessary to manually adjust the K value for re-clustering, and then use manual intervention methods to adjust the original classification according to the previous categories.
[0039] The above method has at least the following defects: the user is required to set the clustering K value. Since the sample categories of unsupervised samples are unknown in reality or the user setting is not the optimal value, the clustering results will be poor; unsupervised clustering randomly initializes the cluster centers, resulting in different clustering results each time; even if the user initializes the cluster centers, there is no response plan for new categories that appear in incremental samples; there is no continuity between the models trained each time. When new samples are added or new categories are added, retraining is required and the previous training results cannot be inherited, and the newly trained samples need to be manually recalibrated.
[0040] Based on the above problems, an embodiment of the present invention provides an embodiment of a method for data clustering. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] Figure 2 is a flow chart of a data clustering method according to an embodiment of the present invention. Figure 2 As shown, the method comprises the following steps:
[0042] Step S102, obtaining a first optimal cluster value predetermined based on the original data of the data to be clustered;
[0043] Step S104, after detecting the incremental data of the data to be clustered, performing secondary clustering processing on the data to be clustered to obtain a plurality of second clustering index values, and selecting a second target clustering index value from the plurality of second clustering index values;
[0044] Step S106, obtaining a second optimal cluster value corresponding to the second target cluster index value;
[0045] Step S108, based on the comparison result of the first best cluster value and the second best cluster value, the cluster center in the cluster model is updated, and the updated cluster model is used to perform K-means clustering processing on the data to be clustered to obtain a target clustering processing result.
[0046] Optionally, the value of the first optimal clustering cluster K1 described above can be, but is not limited to, the optimal number of clusters corresponding to the minimum first clustering index value; the value of the second optimal clustering cluster K2 described above can be, but is not limited to, the optimal number of clusters corresponding to the minimum second clustering index value.
[0047] Optionally, by clustering the original data multiple times, calculate the DBI value of the clustering after each clustering. The number of clusters corresponding to the clustering index value (i.e., the above-mentioned first target clustering index value) when the DBI value is the smallest is the optimal number of clusters, that is, the above-mentioned first optimal clustering cluster value K1.
[0048] Optionally, by clustering the original data and the incremental data multiple times, calculate the DBI value of the clustering after each clustering. The number of clusters corresponding to the clustering index value (i.e., the above-mentioned second target clustering index value) when the DBI value is the smallest is the optimal number of clusters, that is, the above-mentioned second optimal clustering cluster value K2.
[0049] Optionally, the above comparison results include any one of the following: K1>K2, K1<K2, K1 = K2. Different comparison results correspond to different clustering processing results.
[0050] Optionally, the method for obtaining the above-mentioned first optimal clustering cluster value K1 and the above-mentioned second optimal clustering cluster value K2 can also be the elbow method. Multiple clusterings can be performed to calculate the elbow value each time, and based on the above elbow value, the point with the largest slope change is obtained. The K value of this point is the optimal K value (i.e., the above-mentioned first optimal clustering cluster value K1 and / or the above-mentioned second optimal clustering cluster value K2).
[0051] In the embodiments of the present invention, a data clustering method is adopted. By obtaining a first optimal clustering cluster value预先 determined based on the original data of the data to be clustered; after detecting the incremental data of the data to be clustered, performing secondary clustering processing on the data to be clustered to obtain multiple second clustering index values, and selecting a second target clustering index value from the multiple second clustering index values; obtaining a second optimal clustering cluster value corresponding to the second target clustering index value; according to the comparison result of the first optimal clustering cluster value and the second optimal clustering cluster value, updating the clustering center in the clustering model, and using the updated clustering model to perform K-means clustering processing on the data to be clustered to obtain a target clustering processing result, achieving the purpose of determining a clustering processing result that meets the requirements of the new data according to the characteristics of the new data, thereby realizing the technical effect of improving the data clustering efficiency and clustering effect, and further solving the technical problem of low clustering efficiency and poor clustering results caused by the inability of the existing data clustering method to automatically meet the clustering requirements of new data.
[0052] As an optional embodiment, Figure 3is a schematic diagram of an optional system structure for implementing the above data clustering method according to an embodiment of the present invention, such as Figure 3 As shown, the above system specifically includes two modules, namely the data module and the clustering module. Among them, in the above data module, the data processing module, the data storage module, the data discovery module and the data index classification module complement each other. The data processing module is used to process the data, such as data conversion, reading and normalization; the data discovery module is used to provide incremental data other than the first clustering, and by calculating the unique hash value of the file in the data pool, the newly added data in the data pool is discovered to prepare for the next incremental training, and the results of the data discovery are provided to the data index classification module, so that the clustering module can quickly obtain data from different parts; the model part (i.e., the model storage module and the model loading module) is used to provide management for the trained model, and the basic model of incremental training can be selected to facilitate the management and rollback of the model. In the above clustering module, in addition to the K-means clustering module, the optimal K value acquisition part is also added to automatically calculate the optimal degree of aggregation that the clustering should meet; the optimal cluster center initialization module is used to calculate and initialize the new cluster center.
[0053] In an optional embodiment, obtaining a first optimal cluster value predetermined based on original data of the data to be clustered includes:
[0054] Step S202, performing a first clustering process on the original data to obtain a plurality of first clustering index values, and selecting a first target clustering index value from the plurality of first clustering index values;
[0055] Step S204: Obtain a first optimal cluster value corresponding to the first target cluster index value.
[0056] Optionally, the first clustering index value is a Davies Bouldin index value (ie, a DBI value), and the first target clustering index value may be, but is not limited to, a clustering index value corresponding to a minimum DBI value.
[0057] Optionally, by clustering the original data multiple times, the DBI value of the cluster is calculated after each clustering, and the number of clusters corresponding to the clustering index value when the DBI value is the smallest (that is, the first target clustering index value) is the optimal number of clusters, that is, the first optimal clustering cluster value K1.
[0058] It should be noted that the method for obtaining the above-mentioned first optimal cluster value K1 (ie, automatically finding the optimal K value) in the embodiment of the present invention is applicable to the optimal parameter selection stage of all unsupervised algorithms of convex clustering.
[0059] In an optional embodiment, after obtaining the first optimal cluster value predetermined based on the original data of the data to be clustered, the method further includes:
[0060] Step S302: Perform K-means clustering processing on the original data using the first optimal cluster value to obtain a first clustering processing result.
[0061] Optionally, the above-mentioned first clustering processing result at least includes: a clustering model and cluster centers in the clustering model.
[0062] Optionally, the first best cluster value corresponding to the K1 value when the DBI value is the smallest is used to perform K-means clustering processing on the original data, and the clustering model and the clustering center in the clustering model are obtained and saved to prepare for the next clustering.
[0063] Optionally, the above clustering processing method may also be a Mini Batch K-means clustering method.
[0064] In an optional embodiment, determining the cluster center in the clustering model according to the comparison result of the first best cluster cluster value and the second best cluster cluster value includes:
[0065] Step S402: when the comparison result is that the first best cluster value is less than the second best cluster value, a target cluster center is determined based on the incremental data and the initial cluster center, and the newly added cluster center and the initial cluster center are used as the initial cluster centers for the next clustering process;
[0066] Step S404, when the comparison result is that the first best cluster value is greater than the second best cluster value, it is determined that the distribution of the incremental data is inconsistent with the distribution of the original data, and the clustering process is canceled;
[0067] Step S406: when the comparison result is that the first best cluster value is equal to the second best cluster value, the cluster center determined by the first clustering process on the original data is still used as the initial cluster center for the next clustering process.
[0068] Optionally, the newly added cluster center can be but is not limited to being obtained by calculating the distance between the above-mentioned incremental data and the initial cluster center, and the point in the above-mentioned incremental data farthest from the above-mentioned initial cluster center is used as the above-mentioned newly added cluster center, and the above-mentioned newly added cluster center and the above-mentioned initial cluster center are used as the initial cluster centers for the next clustering processing.
[0069] It should be noted that when the above first optimal clustering cluster value is less than the above second optimal clustering cluster value, that is, when K1 < K2, it indicates that a new classification has been added to the incremental data. At this time, new clustering centers need to be generated; when the above comparison result is that the above first optimal clustering cluster value is greater than the above second optimal clustering cluster value, that is, when K1 > K2, it indicates that the data distribution of the added incremental data does not match the original data. At this time, the clustering process is exited to avoid damaging the clustering model; when the above comparison result is that the above first optimal clustering cluster value is equal to the above second optimal clustering cluster value, that is, when K1 = K2, it indicates that the incremental data has not generated new categories, and the original clustering center can be used as the initial clustering center for the new clustering.
[0070] As an alternative embodiment, Figure 4 is a flowchart of another alternative data clustering method according to an embodiment of the present invention. As Figure 4 shown, the method specifically includes the following steps: By clustering the original data multiple times, calculate the DBI value of the clustering after each clustering. The number of clusters corresponding to the clustering index value (i.e., the above first target clustering index value) when the DBI value is the smallest is the optimal number of clusters, that is, the above first optimal clustering cluster value K1; By clustering the original data and incremental data multiple times, calculate the DBI value of the clustering after each clustering. The number of clusters corresponding to the clustering index value (i.e., the above second target clustering index value) when the DBI value is the smallest is the optimal number of clusters, that is, the above second optimal clustering cluster value K2; By comparing the magnitude relationship between K1 and K2, determine the clustering center and subsequent clustering operations: When K2 = K1, it indicates that the incremental data has not generated new categories, so the original clustering center is used as the initial clustering center for the new clustering, and then cluster according to the clustering rules of K-means until the clustering center no longer changes; When K2 < K1, generally when the initial data distribution is satisfied, this situation should not occur, but once the incremental data distribution does not match the original data, this situation will occur. At this time, a prompt that the incremental data distribution does not match the original data needs to be issued, and at the same time, the current clustering process is exited to avoid damaging the original clustering model; When K2 > K1, it indicates that a new classification has been added to the incremental data, so new clustering centers need to be generated. Then, select the point farthest from each existing clustering center in the new added data as the new clustering center, and add the newly selected clustering center to the clustering center.
[0071] The embodiments of the present invention can at least achieve the following technical effects: automatically calculate the optimal number of clusters K, which no longer needs to be specified manually, and can automatically discover the optimal value hidden in the data; cluster the same data while keeping the category of the first clustering unchanged, solving the problem that the category of the cluster center of K-means changes each time; support incremental clustering, automatically discover new data, automatically discover new categories without destroying the original clustering results; if the new data does not conform to the classification, the user will be reminded, maintaining the continuity and maintainability of the clustering model.
[0072] As an optional embodiment, Figure 5 is a flowchart of another optional data clustering method according to an embodiment of the present invention. Figure 5 As shown, after the newly added cluster center and the initial cluster center are used as the initial cluster center for the next clustering, the method further includes:
[0073] Step S502, based on the first best cluster value and the second best cluster value, a third best cluster value is calculated;
[0074] Step S504, according to the third best cluster value, cyclically executing the step of selecting a distance target based on the incremental data and taking the target as the initial cluster center of the next clustering, to obtain a plurality of corresponding initial cluster centers;
[0075] Step S506, iterating the clustering center of the clustering model using the multiple initial clustering centers to obtain the updated clustering model.
[0076] Optionally, the third best cluster value may be, but is not limited to, a difference between the second best cluster value and the first best cluster value, that is, the third best cluster value is K2-K1.
[0077] Optionally, the step of selecting a distance target based on the incremental data and using the target as the initial cluster center for the next clustering is executed in a loop K2-K1 times, K2-K1 new cluster centers are added, and the K2-K1 new cluster centers are used as the initial cluster centers of K-means (i.e., the initial cluster centers).
[0078] As an optional embodiment, Figure 4As shown in the figure, when K2>K1, it indicates that new categories have been added to the incremental data, so new cluster centers need to be generated. Then the points in the newly added data that are farthest from the existing cluster centers are selected as new cluster centers, and the newly selected cluster centers are added to the cluster centers. The above method is repeated K2-K1 times, K2-K1 new cluster centers are added, and these centers are used as the initial cluster centers of K-means, and the cluster centers are iterated to obtain the best clustering effect.
[0079] As an optional embodiment, Figure 6 is a schematic diagram of an application scenario of an optional data clustering method according to an embodiment of the present invention, such as Figure 6 As shown, the above method or system can be used in combination with supervised learning. It should be noted that supervised learning generally requires a large number of labeled samples, and unsupervised learning can be responsible for an initial clustering of disorganized data. The user only needs to label each cluster. When new incremental data is added, the original data label will not change, and will not affect the category and model of the original supervised learning algorithm, which greatly reduces the manual workload and can provide relatively accurate data for the supervised classification algorithm in a continuous and stable manner.
[0080] As an optional embodiment, Figure 7 is a schematic diagram of an application scenario of an optional data clustering method according to an embodiment of the present invention, such as Figure 7 As shown, the above method can be applied to the user classification system. In the actual application scenario, the number of people is huge, and it is impossible to label each user. The basic users can be clustered first, and then the content preferred by different users can be recommended. When new users come, the clustering function can be used to observe whether new user categories are added. The content data they like can be recommended according to their preferences, and incremental learning can be performed on the new users in a timely manner to maintain the accuracy and continuous updating of the model.
[0081] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0082] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0083] Example 2
[0084] In this embodiment, a data clustering device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the terms "unit" and "device" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0085] According to an embodiment of the present invention, a device embodiment for implementing the above data clustering method is also provided. Figure 8 is a structural schematic diagram of a data clustering device according to an embodiment of the present invention. Figure 8 As shown, the data clustering device includes: a first acquisition module 60, a first processing module 62, a second acquisition module 64, and a second processing module 66, wherein:
[0086] The first acquisition module 60 is used to acquire a first optimal cluster value predetermined based on the original data of the data to be clustered;
[0087] The first processing module 62 is used to perform secondary clustering processing on the data to be clustered after detecting the incremental data of the data to be clustered, obtain multiple second clustering index values, and select a second target clustering index value from the multiple second clustering index values;
[0088] The second acquisition module 64 is used to obtain the second best cluster value corresponding to the second target cluster index value;
[0089] The second processing module 66 is used to update the cluster center in the clustering model according to the comparison result of the first optimal cluster cluster value and the second optimal cluster cluster value, and use the updated clustering model to perform K-means clustering processing on the data to be clustered to obtain the target clustering processing result.
[0090] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0091] It should be noted that the first acquisition module 60, the first processing module 62, the second acquisition module 64, and the second processing module 66 correspond to steps S102 to S108 in Example 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the modules as part of the device can be run in a computer terminal.
[0092] It should be noted that the optional or preferred implementation of this embodiment can refer to the relevant description in Example 1, which will not be repeated here.
[0093] The above-mentioned data clustering device may also include a processor and a memory. The above-mentioned first acquisition module 60, first processing module 62, second acquisition module 64, second processing module 66, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0094] In an optional embodiment, the above-mentioned first acquisition module also includes: a selection module, which is used to perform an initial clustering process on the above-mentioned original data to obtain multiple first clustering index values, and select a first target clustering index value from the multiple first clustering index values; a first acquisition sub-module, which is used to obtain the first optimal clustering cluster value corresponding to the above-mentioned first target clustering index value.
[0095] In an optional embodiment, the apparatus further includes: a second acquisition submodule, configured to perform K-means clustering processing on the original data using the first optimal cluster value to obtain a first clustering processing result.
[0096] In an optional embodiment, the second processing module further includes: a comparison module, which is used to determine a newly added cluster center based on the incremental data and the initial cluster center when the comparison result is that the first optimal cluster cluster value is less than the second optimal cluster cluster value, and use the newly added cluster center and the initial cluster center as the initial cluster center for the next clustering processing; a first determination submodule, which is used to determine that the distribution of the incremental data and the original data do not match and cancel the clustering processing flow when the comparison result is that the first optimal cluster cluster value is greater than the second optimal cluster cluster value; and a second determination submodule, which is used to still use the cluster center determined by the first clustering processing of the original data as the initial cluster center for the next clustering processing when the comparison result is that the first optimal cluster cluster value is equal to the second optimal cluster cluster value.
[0097] In an optional embodiment, the above-mentioned device also includes: a calculation module, which is used to calculate a third optimal clustering value based on the above-mentioned first optimal clustering cluster value and the above-mentioned second optimal clustering cluster value; a third acquisition submodule, which is used to cyclically execute the step of selecting a distance target based on the above-mentioned incremental data and using the above-mentioned target as the initial clustering center of the next clustering according to the above-mentioned third optimal clustering cluster value, and obtain the corresponding multiple initial clustering centers; a fourth acquisition submodule, which is used to iterate the clustering center of the above-mentioned clustering model using the multiple initial clustering centers to obtain the above-mentioned updated clustering model.
[0098] The processor includes a kernel, which retrieves the corresponding program unit from the memory. The kernel may be one or more. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory includes at least one memory chip.
[0099] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute any of the data clustering methods described above.
[0100] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group, and the non-volatile storage medium includes a stored program.
[0101] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: obtaining a first optimal clustering cluster value predetermined based on the original data of the data to be clustered; after detecting the incremental data of the above-mentioned data to be clustered, performing secondary clustering processing on the above-mentioned data to be clustered to obtain multiple second clustering index values, and selecting a second target clustering index value from the multiple above-mentioned second clustering index values; obtaining a second optimal clustering cluster value corresponding to the above-mentioned second target clustering index value; according to the comparison result of the above-mentioned first optimal clustering cluster value and the above-mentioned second optimal clustering cluster value, updating the cluster center in the clustering model, and using the updated clustering model to perform K-means clustering processing on the above-mentioned data to be clustered to obtain the target clustering processing result.
[0102] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: perform an initial clustering process on the above-mentioned original data to obtain multiple first clustering index values, and select a first target clustering index value from the multiple first clustering index values; obtain the first optimal clustering cluster value corresponding to the above-mentioned first target clustering index value.
[0103] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following function: performing K-means clustering processing on the original data using the first optimal clustering cluster value to obtain a first clustering processing result.
[0104] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: when the comparison result is that the value of the first optimal cluster cluster is less than the value of the second optimal cluster cluster, a new cluster center is determined based on the incremental data and the initial cluster center, and the new cluster center and the initial cluster center are used as the initial cluster center for the next clustering process; when the comparison result is that the value of the first optimal cluster cluster is greater than the value of the second optimal cluster cluster, it is determined that the distribution of the incremental data and the original data are inconsistent, and the clustering process is canceled; when the comparison result is that the value of the first optimal cluster cluster is equal to the value of the second optimal cluster cluster, the cluster center determined by the first clustering process on the original data will still be used as the initial cluster center for the next clustering process.
[0105] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: based on the above-mentioned first optimal cluster cluster value and the above-mentioned second optimal cluster cluster value, a third optimal cluster value is calculated; according to the above-mentioned third optimal cluster cluster value, the step of selecting a distance target based on the above-mentioned incremental data and using the above-mentioned target as the initial cluster center of the next clustering is cyclically executed to obtain the corresponding multiple initial cluster centers; the above-mentioned multiple initial cluster centers are used to iterate the cluster center of the above-mentioned cluster model to obtain the above-mentioned updated cluster model.
[0106] According to an embodiment of the present application, an embodiment of a processor is also provided. Optionally, in this embodiment, the processor is used to run a program, wherein the program executes any one of the data clustering methods described above when it is run.
[0107] According to an embodiment of the present application, an embodiment of a computer program product is also provided. When executed on a data processing device, it is suitable for executing a program that initializes any one of the above-mentioned data clustering method steps.
[0108] Optionally, the above-mentioned computer program product, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: obtaining a first optimal clustering cluster value predetermined based on the original data of the data to be clustered; after detecting the incremental data of the above-mentioned data to be clustered, performing secondary clustering processing on the above-mentioned data to be clustered to obtain multiple second clustering index values, and selecting a second target clustering index value from the multiple second clustering index values; obtaining a second optimal clustering cluster value corresponding to the above-mentioned second target clustering index value; updating the cluster centers in the clustering model based on the comparison result of the above-mentioned first optimal clustering cluster value and the above-mentioned second optimal clustering cluster value, and using the updated clustering model to perform K-means clustering processing on the above-mentioned data to be clustered to obtain the target clustering processing result.
[0109] According to an embodiment of the present application, an embodiment of an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any one of the above-mentioned data clustering methods.
[0110] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0111] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0112] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0113] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0114] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a non-volatile storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0116] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A data clustering method, characterized in that: include: Acquire a first optimal cluster value predetermined based on original data of the data to be clustered, wherein the data to be clustered is user data; After detecting the incremental data of the data to be clustered, performing secondary clustering processing on the data to be clustered to obtain a plurality of second clustering index values, and selecting a second target clustering index value from the plurality of second clustering index values; Obtaining a second optimal cluster value corresponding to the second target cluster index value; According to the comparison result of the first best cluster cluster value and the second best cluster cluster value, the cluster center in the cluster model is updated, and the updated cluster model is used to perform K-means clustering processing on the data to be clustered to obtain a target clustering processing result; Wherein, updating the cluster center in the clustering model according to the comparison result of the first best cluster cluster value and the second best cluster cluster value includes: when the comparison result is that the first best cluster cluster value is less than the second best cluster cluster value, determining a new cluster center based on the incremental data and the initial cluster center, and using the new cluster center and the initial cluster center as the initial cluster center for the next clustering process; when the comparison result is that the first best cluster cluster value is greater than the second best cluster cluster value, determining that the distribution of the incremental data and the original data do not match, and canceling the clustering process; when the comparison result is that the first best cluster cluster value is equal to the second best cluster cluster value, still using the cluster center determined by the first clustering process on the original data as the initial cluster center for the next clustering process.
2. The method according to claim 1, characterized in that: Obtaining a first optimal cluster value predetermined based on original data of the data to be clustered, including: Performing a first clustering process on the original data to obtain a plurality of first clustering index values, and selecting a first target clustering index value from the plurality of first clustering index values; Obtain a first optimal clustering cluster value corresponding to the first target clustering index value.
3. The method according to claim 1, characterized in that: After obtaining a first optimal cluster value predetermined based on the original data of the data to be clustered, the method further includes: The original data is subjected to K-means clustering processing using the first optimal cluster value to obtain a first clustering processing result.
4. The method according to claim 1, characterized in that After the newly added cluster center and the initial cluster center are used as the initial cluster center for the next clustering, the method further includes: Based on the first best cluster value and the second best cluster value, a third best cluster value is calculated; According to the third best cluster value, cyclically executing the step of selecting a distance target based on the incremental data and taking the target as the initial cluster center of the next clustering, to obtain a corresponding plurality of the initial cluster centers; The clustering model is iterated using the multiple initial clustering centers to obtain the updated clustering model.
5. A data clustering device, characterized in that: include: A first acquisition module, used to acquire a first optimal cluster value predetermined based on original data of the data to be clustered, wherein the data to be clustered is user data; A first processing module is used to perform secondary clustering processing on the data to be clustered after detecting the incremental data of the data to be clustered, obtain a plurality of second clustering index values, and select a second target clustering index value from the plurality of second clustering index values; A second acquisition module, used to acquire a second optimal clustering cluster value corresponding to the second target clustering index value; A second processing module is used to update the cluster center in the clustering model according to the comparison result of the first best clustering cluster value and the second best clustering cluster value, and use the updated clustering model to perform K-means clustering processing on the data to be clustered to obtain a target clustering processing result; Among them, the second processing module also includes: a comparison module, which is used to determine a new cluster center based on the incremental data and the initial cluster center when the comparison result is that the first optimal cluster cluster value is less than the second optimal cluster cluster value, and use the new cluster center and the initial cluster center as the initial cluster center for the next clustering processing; a first determination submodule, which is used to determine that the distribution of the incremental data and the original data is inconsistent and cancel the clustering processing flow when the comparison result is that the first optimal cluster cluster value is greater than the second optimal cluster cluster value; and a second determination submodule, which is used to still use the cluster center determined by the first clustering processing of the original data as the initial cluster center for the next clustering processing when the comparison result is that the first optimal cluster cluster value is equal to the second optimal cluster cluster value.
6. The device according to claim 5, characterized in that The first acquisition module also includes: A selection module, configured to perform a first clustering process on the original data to obtain a plurality of first clustering index values, and select a first target clustering index value from the plurality of first clustering index values; The first acquisition submodule is used to acquire a first optimal cluster value corresponding to the first target cluster index value.
7. The device according to claim 5, characterized in that The device also includes: The second acquisition submodule is used to perform K-means clustering processing on the original data using the first optimal cluster cluster value to obtain a first clustering processing result.
8. The device according to claim 5, characterized in that The device also includes: A calculation module, configured to calculate a third optimal cluster value based on the first optimal cluster value and the second optimal cluster value; A third acquisition submodule is used to cyclically execute the step of selecting a distance target based on the incremental data and taking the target as the initial cluster center of the next clustering according to the third optimal cluster value, so as to obtain a corresponding plurality of the initial cluster centers; The fourth acquisition submodule is used to iterate the clustering center of the clustering model using the multiple initial clustering centers to obtain the updated clustering model.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the data clustering method according to any one of claims 1 to 4.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the data clustering method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Incremental data clustering method, apparatus and device and readable storage medium
CN110866555A
Target clustering number obtaining method and device and computer system
CN110895706A