Method, apparatus, electronic device, and storage medium for clustering data

Multiple preliminary clustering results are generated through multiple clustering algorithms and the membership matrix is fused, which solves the problem of clustering algorithm selection and parameter influence, and achieves more accurate and more adaptable clustering results.

CN114548276BActive Publication Date: 2025-07-29GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210163273.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-07-29
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

In the prior art, the initialization parameters after selection of the clustering algorithm have a significant impact on the clustering results, resulting in difficulty in selecting suitable clustering methods and parameters, and a single clustering algorithm is insufficiently adaptable to different data structures.

Method used

Multiple clustering algorithms are used to initially cluster the cluster data, multiple first clustering results are generated, and the membership degree of data in each cluster is represented by the membership degree matrix, and multiple membership degree matrices are fused for re-clustering to determine the final clustering result.

Benefits of technology

By fusing the membership matrix information of multiple clustering algorithms, more data structure characteristics are retained, the problem of insufficient adaptability of a single algorithm is avoided, and the accuracy and adaptability of clustering results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548276B_ABST
    Figure CN114548276B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, electronic device, and storage medium for clustering data, belonging to the technical field of data processing. The method includes: obtaining a plurality of data to be clustered for a target clustering event; respectively clustering the plurality of data to be clustered by a plurality of clustering algorithms to obtain a plurality of first clustering results; for each first clustering result, determining a membership degree matrix of the plurality of data to be clustered under the first clustering result, where the membership degree matrix represents the membership degree of each data to be clustered relative to each cluster of the first clustering result under the first clustering result; based on the plurality of membership degree matrices, clustering the plurality of data to be clustered to obtain a second clustering result of the target clustering event, so as to determine the categories of the plurality of data to be clustered. In this way, clustering the plurality of data to be clustered again based on the membership degree matrix that fuses various partitioning information of the data to be clustered retains more partitioning information and avoids the problem that a single clustering algorithm is not suitable for the data structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of data processing, and particularly to a method, apparatus, electronic device, and storage medium for clustering data. Background Art

[0002] With the development of data processing technology, data collection means have gradually matured, and the amount of collected data has increased significantly. As the amount of collected data increases significantly, extracting useful information from the collected data to interpret the data has become the most difficult problem. Clustering data can reveal the internal relationship between data and features and plays an important role in the process of extracting information.

[0003] In related technologies, many clustering algorithms have been developed to handle different problems. For example, partitioning clustering, density clustering, or hierarchical clustering, etc. These clustering algorithms use different distances or similarities as measurement parameters and use different objective functions for measurement. Different clustering algorithms will produce different clustering results for the same data set, and often show different performances for data sets with different data structures. Therefore, when clustering data, it is necessary to select the corresponding clustering method for clustering.

[0004] In the above related technologies, once the clustering algorithm is selected, the initialization parameters have a significant impact on the clustering result. Therefore, it is difficult to select a suitable clustering algorithm and various parameters during the clustering process. Therefore, there is an urgent need for a new clustering method. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, electronic device, and storage medium for clustering data, which avoid the problem that a single clustering algorithm is not suitable for the data structure. The technical solution is as follows:

[0006] On the one hand, a method for clustering data is provided, and the method includes:

[0007] Obtain multiple data to be clustered for a target clustering event;

[0008] Cluster the multiple data to be clustered respectively by multiple clustering algorithms to obtain multiple first clustering results;

[0009] For each first clustering result, determine the membership degree matrix of the multiple data to be clustered under the first clustering result, where the membership degree matrix represents the membership degree of each data to be clustered relative to each cluster of the first clustering result under the first clustering result;

[0010] Cluster the multiple data to be clustered based on multiple membership degree matrices to obtain a second clustering result of the target clustering event, so as to determine the categories of the multiple data to be clustered.

[0011] On the other hand, a device for clustering data is provided. The device includes:

[0012] An acquisition module, configured to acquire a plurality of data to be clustered for a target clustering event;

[0013] A first clustering module, configured to cluster the plurality of data to be clustered respectively by multiple clustering algorithms to obtain a plurality of first clustering results;

[0014] A determination module, configured to, for each first clustering result, determine a membership degree matrix of the plurality of data to be clustered under the first clustering result, where the membership degree matrix represents the membership degree of each data to be clustered relative to each cluster under the first clustering result;

[0015] A second clustering module, configured to cluster the plurality of data to be clustered based on a plurality of membership degree matrices to obtain a second clustering result of the target clustering event, so as to determine the categories of the plurality of data to be clustered.

[0016] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory; the memory stores at least one program code, and the at least one program code is used to be executed by the processor to implement the method for clustering data as described in the above aspect.

[0017] On the other hand, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one program code, and the at least one program code is used to be executed by a processor to implement the method for clustering data as described in the above aspect.

[0018] On the other hand, a computer program product is provided. The computer program product stores at least one program code, and the at least one program code is used to be executed by a processor to implement the method for clustering data as described in the above aspect.

[0019] In the embodiments of the present application, multiple clustering algorithms are used for clustering to obtain multiple clustering results, and membership degrees of a plurality of data to be clustered under different clusters are determined based on the multiple clustering results. In this way, the partitioning information of the plurality of data to be clustered is represented by multiple membership degree matrices. Therefore, in the process of re-clustering, the plurality of data to be clustered are clustered based on the membership degree matrix that fuses multiple partitioning information of the data to be clustered, retaining more partitioning information and avoiding the problem that a single clustering algorithm is not suitable for the data structure. Description of the Drawings

[0020] Figure 1 Shows a schematic structural diagram of a terminal provided by an exemplary embodiment of the present application;

[0021] Figure 2 Shows a schematic structural diagram of a server provided by an exemplary embodiment of the present application;

[0022] Figure 3 Shows a flowchart of a method for data clustering of data shown in an exemplary embodiment of the present application;

[0023] Figure 4 Shows a flowchart of a method for data clustering of data shown in an exemplary embodiment of the present application;

[0024] Figure 5 Shows a block diagram of the structure of a data clustering device for data provided by an embodiment of the present application. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0026] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0027] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the data to be clustered involved in the present application is obtained under full authorization.

[0028] In the embodiments of the present application, the electronic device may be provided as a terminal or a server. When the electronic device is provided as a terminal, please refer to Figure 1 , which shows a block diagram of the structure of a terminal 100 provided by an exemplary embodiment of the present application. The terminal 100 may be a terminal with data processing functions such as a smart phone or a tablet computer. The terminal 100 in the present application may include one or more of the following components: a processor 110 and a memory 120.

[0029] Optionally, the processor 110 includes one or more processing cores. The processor 110 connects various parts within the entire terminal 100 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by invoking data stored in the memory 120, it performs various functions of the terminal 100 and processes data. Optionally, the processor 110 is implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), a neural-network processing unit (NPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen 130; the NPU is used to implement artificial intelligence (AI) functions; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110 and can be implemented separately by a single chip.

[0030] Optionally, the memory 120 includes a random access memory (RAM) and may also include a read-only memory. Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc.; the data storage area can store data created according to the use of the terminal 100 (such as audio data, phone book, etc.).

[0031] In some embodiments, the terminal 100 further includes a display screen. The display screen is a display component for displaying a user interface. Optionally, the display screen is a display screen with a touch function, through which a user can use a finger, a stylus, or any other suitable object to perform touch operations on the display screen 130.

[0032] The display screen is typically provided on the front panel of the terminal 100. The display screen can be designed as a full screen, a curved screen, a special-shaped screen, a double-sided screen, or a foldable screen. The display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, etc., which are not limited in this embodiment.

[0033] In addition, those skilled in the art will understand that the structure of the terminal 100 shown in the above figures does not constitute a limitation of the terminal 100. The terminal 100 may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the terminal 100 also includes a microphone, a speaker, a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (Wi-Fi) module, a power supply, a Bluetooth module, and other components, which will not be described in detail here.

[0034] In the case where the electronic device is provided as a server, please refer to Figure 2 , which shows a structural block diagram of a server 200 provided by an exemplary embodiment of the present application. The server 200 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 210 and one or more memories 220, wherein the memory 220 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 210 to implement the method for clustering data provided by the above-mentioned various method embodiments. Of course, the server 200 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 200 may also include other components for implementing device functions, which will not be described here.

[0035] The following introduces the application scenarios of this solution.

[0036] Clustering is an unsupervised learning technique for automatically finding categories, which divides multiple unlabeled data into groups (clusters) with similar characteristics. Data belonging to the same cluster are more similar than data not belonging to the same cluster. In some embodiments, clustering the dataset X means finding k clusters, where the data within each cluster are as similar as possible, and the data in different clusters are as different as possible. Here, the dataset X includes multiple data to be clustered. Clustering has been successfully applied in different fields. For example, in market segmentation, customers with similar behaviors or attributes are found through clustering; or, in image processing, similar image regions are grouped together through clustering; or, in document management, documents with the same theme are classified through clustering.

[0037] With the maturity of data processing technology, a large number of clustering algorithms have emerged. Different clustering algorithms can cluster data based on different criteria, thus being able to be divided into various clustering methods. For example, clustering based on partitioning, clustering based on hierarchy, clustering based on density, clustering based on model, clustering based on fuzzy, etc. These different clustering algorithms have different principles and adaptabilities. For example, the k-means partitioning clustering algorithm is sensitive to noise and outliers, but cannot solve non-convex data; the model-based clustering algorithm cannot handle irregularly distributed data; the fuzzy clustering algorithm works well on data that satisfies the normal distribution, but is sensitive to isolated points.

[0038] In addition, choosing different parameters and initializations for clustering algorithms will also have a great impact on the clustering algorithms. For example, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm for applications with noise is not sensitive to noise and can discover clusters of any shape. However, when the sparsity of the clusters is different, using fixed parameters for identification will destroy the natural structure of the clusters; the performance of the fuzzy clustering algorithm depends on the initial clustering centers.

[0039] Therefore, no clustering algorithm can be generally used to solve multiple problems. It is difficult to select a suitable clustering algorithm and initialization parameters when clustering data. The embodiments of this application propose a clustering combination method that combines the clustering results of multiple clustering algorithms to obtain a suitable clustering result.

[0040] Please refer to Figure 3 , which shows the flowchart of the method for clustering data shown in an exemplary embodiment of this application. The method includes:

[0041] Step S301: The electronic device obtains multiple data to be clustered for the target clustering event.

[0042] The target clustering event is any event that requires classifying data. The multiple data to be clustered are the data corresponding to the target event. In some embodiments, the target clustering event is a user behavior analysis event, and the data to be clustered is user behavior data. For example, in the case of determining the characteristics of the user group using a target application by analyzing user behavior, it is necessary to cluster the user behavior data, and the data to be clustered includes at least one of the user's age, gender, and the time period of using the target application. For another example, in the case of determining the user's video viewing interest by analyzing user behavior, the data to be clustered includes the video data that the user has watched historically. In some embodiments, the target clustering event is a multimedia data analysis event, and the data to be clustered is multimedia data. For example, in the case of segmenting an image, it is necessary to cluster the image data, and the data to be clustered includes the image data. For another example, in the case of classifying a video, it is necessary to cluster the video data, and the data to be clustered includes the image data.

[0043] In some embodiments, the electronic device receives the data to be clustered input by the user. For example, the electronic device receives the image data input by the user. In some embodiments, the electronic device reads multiple data to be clustered from a database. For example, in the case of analyzing the behavior of a certain user, in response to receiving a clustering instruction, the user behavior data corresponding to the user indicated by the clustering instruction is read from the database. Or, in the case of analyzing the behavior data of multiple users to obtain the behavior characteristics of the multiple users, the user behavior data generated within a specified time period can be obtained. Correspondingly, in response to receiving a clustering instruction, the user behavior data generated within the time period corresponding to the clustering instruction is read from the database.

[0044] Step S302: The electronic device clusters the multiple data to be clustered respectively through multiple clustering algorithms to obtain multiple first clustering results.

[0045] The multiple clustering algorithms are algorithms that cluster based on different bases. The first clustering result is a clustering result obtained by clustering the multiple data to be clustered based on any clustering algorithm. Each first clustering result includes multiple clusters, and each cluster includes at least one data to be clustered.

[0046] In some embodiments, the electronic device performs clustering on multiple data to be clustered through different clustering algorithms respectively, and obtains a first clustering result corresponding to each clustering algorithm. For example, in the case of determining the characteristics of the user group of the target application by analyzing user behavior, the user behavior data is clustered through the k-means partitioning clustering algorithm respectively, and multiple groupings of user data under different user characteristics are obtained. For another example, when performing image segmentation, based on the model-based clustering algorithm, pixel points with the same pixel characteristics are divided into the same image region, so as to divide the image into multiple image regions. In the embodiments of the present application, different types of clustering algorithms are used for clustering, so as to fully identify and explore various data structures of the multiple data to be clustered, and rich clustering results are obtained.

[0047] In some embodiments, for each clustering algorithm, the electronic device determines different initial parameters, performs clustering based on the different initial parameters, and obtains a first clustering result of each clustering algorithm under different initial parameters. Correspondingly, the electronic device determines the initial parameters of the multiple clustering algorithms based on the multiple data to be clustered; based on the initial parameters, the multiple data to be clustered are clustered respectively through the multiple clustering algorithms, and the multiple first clustering results are obtained. For example, in the case of determining the characteristics of the user group of the target application by analyzing user behavior, the user behavior data is clustered respectively using different initial parameters, and multiple groupings of user data under different user characteristics are obtained. In the embodiments of the present application, by adjusting different initial parameters to cluster the multiple data to be clustered, clustering results with large differences but all reasonable can be obtained. By combining the clustering results, the uncertainty of the clustering algorithm in parameters is alleviated, making the clustering results more reasonable and rich, and thus making the clustering results more adaptable.

[0048] It should be noted that in the embodiments of the present application, neither the number of data to be clustered nor the number of clustering algorithms is specifically limited.

[0049] Step S303: For each first clustering result, the electronic device determines the membership matrix of the multiple data to be clustered under the first clustering result, and the membership matrix represents the membership of each data to be clustered relative to each cluster under the first clustering result.

[0050] In this step, the electronic device determines the membership matrix of each data to be clustered under each first clustering result respectively.

[0051] Step S304: The electronic device clusters the multiple data to be clustered based on the multiple membership matrices, and obtains a second clustering result of the target clustering event to determine the categories of the multiple data to be clustered.

[0052] In this step, the electronic device uses multiple membership matrices as clustering parameters based on the target clustering event to cluster multiple data to be clustered. Among them, the electronic device determines a target clustering algorithm that matches the target clustering event based on the target clustering event, and continues to cluster the multiple data to be clustered based on the multiple membership matrices through the target clustering algorithm to obtain a second clustering result. In the embodiments of the present application, the clustering algorithm used for reclustering is not specifically limited. For example, the algorithm used in the reclustering process is the K-means partitioning clustering algorithm or the Euclidean distance algorithm.

[0053] The multiple membership matrices represent the partitioning information of the multiple data to be clustered in the multiple first clustering results and the similarity of the multiple clusters in the multiple first clustering results. For example, for two clusters that are far apart, the membership degree differences of the same data to be clustered are generally large, and for clusters that are close, the membership degree differences of the same data to be clustered are generally small. Example: For clusters A, B, and C, A and B are close and far from C. In the constructed membership matrix, the membership degree values representing A and B have small differences, while the membership degree values representing C have large differences. In this way, the relationships between multiple first clustering results are represented by the membership matrix, so that when clustering the data to be clustered, the clustering results corresponding to different clustering algorithms can be combined.

[0054] For example, in the case of determining the characteristics of the user group using the target application program by analyzing user behavior, based on the multiple membership matrices of multiple user data, the membership degree of the user behavior data corresponding in the membership matrix is determined. If the membership degrees of the user behavior data in different membership matrices are similar, the characteristic parameters of the cluster where the user data is located are strengthened, and the user behavior data with strengthened characteristic parameters is reclustered to obtain groupings of multiple user data under different user characteristics.

[0055] In some embodiments, before reclustering the multiple membership matrices, the electronic device also fuses the multiple membership matrices. Accordingly, the electronic device fuses the membership matrices under the multiple first clustering results based on the number of clusters in each first clustering result to obtain a fusion matrix; and clusters the multiple data to be clustered based on the fusion matrix to obtain the second clustering result. In some embodiments, the way for the electronic device to fuse the membership matrices under the multiple first clustering results is that the electronic device horizontally splices the membership matrices under the multiple first clustering results.

[0056] In the embodiments of the present application, the electronic device horizontally splices the membership matrices under the multiple first clustering results to form a size of Membership matrix. Where z is the number of the first clustering results, k is the number of clusters corresponding to each first clustering result. When the number of clusters in each clustering result is k, horizontal splicing refers to splicing into a matrix with a length of m and a width of k*z. When the clustering results are different each time, horizontal splicing refers to splicing into a matrix with a length of m and a width of matrix, where k n is the number of clusters in the nth clustering result. For example, information collected by 100 devices is used for K-means clustering to obtain 5 clusters, and DBSCAN clustering is also performed to obtain 10 clusters. Membership matrices with a length of 100 and a width of 5, and a length of 100 and a width of 10 are respectively constructed. The spliced matrix has a length of 100 and a width of 15.

[0057] In the embodiments of the present application, by splicing the membership matrices corresponding to multiple first clustering results into one membership matrix, the spliced membership matrix contains clustering information corresponding to multiple clustering algorithms, as well as the similarity relationships between multiple clustering results. Therefore, when re-clustering, it is possible to fuse the membership matrix that integrates multiple features of the data, retain more clustering information, and avoid the problem that a single clustering algorithm is not suitable for some data structures.

[0058] In some embodiments, the electronic device determines the weights of the multiple membership matrices based on the target event. Correspondingly, based on the number of clusters in each first clustering result, the membership matrices under the multiple first clustering results are fused to obtain a fusion matrix, including: based on the weights of the multiple membership matrices and the number of clusters in each first clustering result, the membership matrices under the multiple first clustering results are fused to obtain a fusion matrix.

[0059] In some embodiments, the weights of the membership matrix are determined based on the prior experience corresponding to the target event. For example, the prior experience is determined based on at least one of the shape, dimension, and sample size of the data to be clustered. Correspondingly, if some clustering algorithms perform well on the current multiple data to be clustered, and some initialization methods or parameters have good robustness, then these clustering algorithms or these initialization methods or the clustering algorithms corresponding to the adopted numbers are set with higher weights.

[0060] In the embodiments of the present application, the weights of different membership matrices are determined according to the target event to reflect the importance of different membership matrices, so that when re-clustering the membership matrix, it can conform to the target event, thereby making the clustering result of higher quality.

[0061] In the embodiments of the present application, multiple clustering algorithms are used for clustering to obtain multiple clustering results. Based on the multiple clustering results, the membership degrees of multiple data to be clustered under different clusters are determined. In this way, the partitioning information of multiple data to be clustered is represented by multiple membership matrices. Therefore, in the re-clustering process, based on the membership matrix that integrates various partitioning information of the data to be clustered, multiple data to be clustered are clustered, retaining more partitioning information and avoiding the problem that a single clustering algorithm is not suitable for the data structure.

[0062] Please refer to Figure 4 , which shows a flowchart of a method for clustering data shown in an exemplary embodiment of the present application. The method includes:

[0063] Step S401: The electronic device obtains multiple data to be clustered for a target clustering event.

[0064] The principle of this step is the same as that of step S301 and will not be elaborated here.

[0065] Step S402: The electronic device respectively clusters the multiple data to be clustered through multiple clustering algorithms to obtain multiple first clustering results.

[0066] The principle of this step is the same as that of step S302 and will not be elaborated here.

[0067] Step S403: For each of the first clustering results, the electronic device determines the cluster center of each cluster of the first clustering result.

[0068] In this step, for each cluster, the electronic device determines the average value of the data in the cluster as the cluster center of the cluster. See Formula 1:

[0069] Formula 1:

[0070] where j represents the identifier of the cluster in the first clustering result, c j represents the cluster center of cluster j in the first clustering result, N is the number of data in cluster j in the first clustering result, and x i represents the i-th data.

[0071] Step S404: The electronic device determines the membership degree of the data to be clustered relative to each cluster based on the distances from each data to be clustered to multiple cluster centers.

[0072] The membership degree represents the degree of approximation between the data and each cluster. The sum of the membership degrees of a data belonging to all clusters of the first clustering result is 1.

[0073] In some embodiments, for each piece of data, the electronic device determines the membership degree of the data relative to the target cluster based on the ratio of the first distance to the second distance. The first distance is the distance between the data and the cluster center of the target cluster, and the second distance is the distances between the data and the cluster centers of the multiple clusters in the first clustering result. The target cluster is any one of the clusters in the first clustering result.

[0074] Correspondingly, in this step, the electronic device determines the distances between each piece of data and the cluster centers c of each cluster respectively. j In some embodiments, the distance is determined by the difference between the data and the cluster center. For example, the distance between the data x i and the cluster center c j is: ||x i - c j ||. Where ||·|| represents taking the modulus value, c j represents the cluster center of cluster j, and x i represents the i-th piece of data. For any piece of data, the membership degree of the data relative to any target cluster is determined by the sum of the ratios of the distance between the data and the target cluster to the distances between the data and other clusters. See Formula Two:

[0075] Formula Two:

[0076] Where c j represents the cluster center of target cluster j, x i represents the i-th piece of data, u ij represents the membership degree of data x i relative to cluster c j , C represents the number of clusters, that is, the number of clusters in the first clustering result, c k represents the cluster center of other cluster k, and m is a factor of the membership degree, and its value is set as needed. For example, the membership degree factor is 2.

[0077] Step S405: The electronic device constructs a membership degree matrix of the multiple pieces of data to be clustered under the first clustering result based on the membership degrees of each piece of data to be clustered relative to each cluster.

[0078] In this step, using the membership degree as the feature of each piece of data, a membership degree matrix of size M*k is constructed. Where M is the number of samples and k is the number of clusters in the first clustering result. The u ij in the membership degree feature matrix u represents the membership degree of any piece of data x i to be clustered among the multiple pieces of data to be clustered relative to cluster c j . The value range of u ij is [0, 1], and that is, the sum of the data in each row is 1.

[0079] Repeat steps S404 - S405 to obtain multiple membership matrices corresponding to multiple first clustering results.

[0080] In the embodiments of the present application, by the distance between the data and the clusters, the similarity of the data with respect to each cluster is determined, thereby representing the membership degree of the data with respect to each cluster. In this way, the similarity information of each data with respect to different clusters is retained through the membership matrix, thus enriching the clustering information.

[0081] Step S406: The electronic device clusters the multiple data to be clustered based on the multiple membership matrices to obtain a second clustering result of the target clustering event, so as to determine the categories of the multiple data to be clustered.

[0082] The principle of this step is the same as that of step S303, and will not be elaborated here.

[0083] In the embodiments of the present application, multiple clustering results are obtained by multiple clustering algorithms, and the membership degrees of the multiple data to be clustered under different clusters are determined based on the multiple clustering results. In this way, the partitioning information of the multiple data to be clustered is represented by multiple membership matrices. Therefore, during the re - clustering process, the multiple data to be clustered are clustered based on the membership matrix that fuses the multiple partitioning information of the data to be clustered, retaining more partitioning information and avoiding the problem that a single clustering algorithm is not suitable for the data structure.

[0084] Please refer to Figure 5 , which shows a structural block diagram of a device for clustering data provided by an embodiment of the present application. The device for clustering data can be implemented as all or part of a processor through software, hardware, or a combination of both. The device includes:

[0085] An acquisition module 501, configured to acquire multiple data to be clustered for a target clustering event;

[0086] A first clustering module 502, configured to cluster the multiple data to be clustered respectively by multiple clustering algorithms to obtain multiple first clustering results;

[0087] A determination module 503, configured to, for each first clustering result, determine a membership matrix of the multiple data to be clustered under the first clustering result, where the membership matrix represents the membership degree of each data to be clustered with respect to each cluster of the first clustering result;

[0088] A second clustering module 504, configured to cluster the multiple data to be clustered based on the multiple membership matrices to obtain a second clustering result of the target clustering event, so as to determine the categories of the multiple data to be clustered.

[0089] In some embodiments, the determination module 503 includes:

[0090] A first determination unit, configured to determine, for each of the first clustering results, the cluster center of each cluster in the first clustering result;

[0091] A second determination unit, configured to determine the membership degree of the data to be clustered with respect to each cluster based on the distances from each data to be clustered to multiple cluster centers;

[0092] A construction unit, configured to construct a membership degree matrix of the multiple data to be clustered under the first clustering result based on the membership degree of each data to be clustered with respect to each cluster.

[0093] In some embodiments, the second determination unit is configured to, for each data to be clustered, respectively determine the ratio of a first distance to multiple second distances, where the first distance is the distance between the data to be clustered and the cluster center of any cluster, and the second distance is the distance between the data to be clustered and the cluster centers of the multiple clusters in the first clustering result; and determine the membership degree of the data to be clustered with respect to the any cluster based on the sum of the ratios of the first distance to the multiple second distances and a membership degree factor.

[0094] In some embodiments, the first determination unit is configured to, for each cluster, determine the cluster center of the cluster based on the average value of the data to be clustered in the cluster.

[0095] In some embodiments, the first clustering module 502 includes:

[0096] A third determination unit, configured to determine initial parameters of the multiple clustering algorithms based on the multiple data to be clustered;

[0097] A first clustering unit, configured to perform clustering on the multiple data to be clustered respectively through the multiple clustering algorithms based on the initial parameters to obtain the multiple first clustering results.

[0098] In some embodiments, the second clustering module 504 includes:

[0099] A fusion unit, configured to fuse the membership degree matrices under the multiple first clustering results based on the number of clusters in each first clustering result to obtain a fusion matrix;

[0100] A second clustering unit, configured to perform clustering on the multiple data to be clustered based on the fusion matrix to obtain the second clustering result.

[0101] In some embodiments, the apparatus further includes:

[0102] A fourth determination unit, configured to determine weights of the multiple membership degree matrices based on the target event;

[0103] The fusion unit is configured to fuse the membership matrices under the multiple first clustering results based on the weights of the multiple membership matrices and the number of clusters in each first clustering result, so as to obtain a fusion matrix.

[0104] In the embodiments of the present application, multiple clustering results are obtained through multiple clustering algorithms, and the membership degrees of multiple data to be clustered under different clusters are determined based on the multiple clustering results. In this way, the division information of the multiple data to be clustered is represented by multiple membership matrices. Therefore, during the re-clustering process, the multiple data to be clustered are clustered based on the membership matrix that fuses the multiple division information of the data to be clustered, retaining more division information and avoiding the problem that a single clustering algorithm is not suitable for the data structure.

[0105] The embodiments of the present application also provide a computer-readable medium, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method for clustering data shown in the above various embodiments.

[0106] The embodiments of the present application also provide a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method for clustering data shown in the above various embodiments.

[0107] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0108] The above are only the optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for clustering data, characterized in that, The method includes: Obtaining a plurality of data to be clustered for a target clustering event; Respectively clustering the plurality of data to be clustered through a variety of clustering algorithms to obtain a plurality of first clustering results; For each first clustering result, determining a membership degree matrix of the plurality of data to be clustered under the first clustering result, where the membership degree matrix represents the membership degree of each data to be clustered relative to each cluster under the first clustering result; Based on the prior experience corresponding to the target clustering event, determining the weights of the plurality of membership degree matrices, where the prior experience is determined based on at least one of the shape, dimension, and sample size of the plurality of data to be clustered; Based on the weights of the plurality of membership degree matrices and the number of clusters in each first clustering result, horizontally splicing the membership degree matrices under the plurality of first clustering results to obtain a fusion matrix; Based on the fusion matrix, clustering the plurality of data to be clustered to obtain a second clustering result of the target clustering event to determine the categories of the plurality of data to be clustered.

2. The method according to claim 1, wherein The step of, for each first clustering result, determining a membership degree matrix of the plurality of data to be clustered under the first clustering result includes: For each first clustering result, determining the cluster center of each cluster in the first clustering result; Based on the distances from each data to be clustered to a plurality of cluster centers, determining the membership degree of the data to be clustered relative to each cluster; Based on the membership degree of each data to be clustered relative to each cluster, constructing a membership degree matrix of the plurality of data to be clustered under the first clustering result.

3. The method according to claim 2, wherein The step of, based on the distances from each data to be clustered to a plurality of cluster centers, determining the membership degree of the data to be clustered relative to each cluster includes: For each data to be clustered, respectively determining the ratio of a first distance to a plurality of second distances, where the first distance is the distance between the data to be clustered and the cluster center of any cluster, and the second distance is the distance between the data to be clustered and the cluster centers of the plurality of clusters in the first clustering result; Based on the sum of the ratios of the first distance to the plurality of second distances and a membership degree factor, determining the membership degree of the data to be clustered relative to any cluster.

4. The method according to claim 2, characterized in that, The step of, for each first clustering result, determining the cluster center of each cluster in the first clustering result includes: For each cluster, based on the average value of the data to be clustered in the cluster, determining the cluster center of the cluster.

5. The method according to claim 1, characterized in that, The step of respectively clustering the plurality of data to be clustered through a variety of clustering algorithms to obtain a plurality of first clustering results includes: Based on the plurality of data to be clustered, determining the initial parameters of the variety of clustering algorithms; Based on the initial parameters, respectively clustering the plurality of data to be clustered through the plurality of clustering algorithms to obtain the plurality of first clustering results.

6. An apparatus for clustering data, characterized in that, The apparatus includes: An obtaining module, configured to obtain a plurality of data to be clustered for a target clustering event; A first clustering module, configured to respectively cluster the plurality of data to be clustered through a variety of clustering algorithms to obtain a plurality of first clustering results; A determination module, configured to determine, for each first clustering result, a membership degree matrix of the multiple data to be clustered under the first clustering result, where the membership degree matrix represents the membership degree of each data to be clustered relative to each cluster of the first clustering result under the first clustering result; A fourth determination module, configured to determine weights of the multiple membership degree matrices based on prior experience corresponding to the target clustering event, where the prior experience is determined based on at least one of the shape, dimension, and sample size of the multiple data to be clustered; A fusion unit, configured to horizontally splice the membership degree matrices under the multiple first clustering results based on the weights of the multiple membership degree matrices and the number of clusters in each first clustering result to obtain a fusion matrix; A second clustering unit, configured to cluster the multiple data to be clustered based on the fusion matrix to obtain a second clustering result of the target clustering event, so as to determine the categories of the multiple data to be clustered.

7. An electronic device, characterized in that, The electronic device includes a processor and a memory; the memory stores at least one program code, and the at least one program code is used to be executed by the processor to implement the method for clustering data according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one program code, and the at least one program code is used to be executed by a processor to implement the method for clustering data according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • A color image segmentation method

    CN101216890A

  • Big data classification method, device and equipment based on hard clustering algorithm

    CN109447103A