Self-organizing mapping weight particle swarm mean clustering method, device, equipment and storage medium
Coarse and fine clustering are performed through self-organized mapping (SOM) clustering algorithm to determine the number of cluster clusters and the initial cluster center, which solves the difficulties of traditional clustering algorithms in determining these parameters and improves the accuracy and applicability of clustering.
Patent Information
- Application Number
- CN202111128532.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-26
AI Technical Summary
Traditional clustering algorithms are difficult to determine the number of cluster clusters and the initial cluster center, which affects the accuracy of clustering results. Different algorithms are suitable for specific data and have poor applicability.
Coarse clustering is performed using self-organized mapping (SOM) clustering algorithm to determine the number of cluster clusters and the initial cluster center, and then fine clustering is performed based on the coarse clustering results to obtain the target cluster center.
It improves the accuracy and effect of clustering, solves the problem of traditional clustering algorithms being sensitive to initial values, and enhances the applicability of the algorithm.
Smart Images

Figure CN113850327B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data mining, and in particular, to a self-organizing mapping weight particle swarm mean clustering method, device, equipment and storage medium. Background Art
[0002] As an important branch in the field of data mining, clustering algorithms have been widely applied in many fields, including machine learning, pattern recognition, image analysis, information retrieval, computer vision, etc. An efficient clustering algorithm can improve work efficiency and work quality.
[0003] However, there are some problems with traditional clustering algorithms. First of all, the selection of initial values is random, and it is difficult to determine the number of clustering clusters and the initial clustering centers, which will directly affect the final clustering results and the accuracy of clustering. Moreover, different clustering algorithms are applicable to specific data, and there is no clustering algorithm that can be applicable to all data, so the applicability is not good. Summary of the Invention
[0004] The embodiments of the present invention provide a self-organizing mapping weight particle swarm mean clustering method, device, equipment and storage medium to realize clustering analysis of sample data and improve the clustering effect and accuracy.
[0005] In the first aspect, the embodiments of the present invention provide a self-organizing mapping weight particle swarm mean clustering method, and the self-organizing mapping weight particle swarm mean clustering method includes:
[0006] Obtain original sample data, where the original sample data is software portrait data;
[0007] Use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; where K is a natural number;
[0008] Based on the rough clustering clusters and the rough clustering centers, perform fine clustering on the original sample data to obtain target clustering centers.
[0009] In the second aspect, the embodiments of the present invention further provide a self-organizing mapping weight particle swarm mean clustering device, and the self-organizing mapping weight particle swarm mean clustering device includes:
[0010] A data acquisition module for acquiring original sample data;
[0011] A rough clustering module for using the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; where K is a natural number;
[0012] A fine clustering module, configured to perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers.
[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes:
[0014] One or more processors;
[0015] A storage device, configured to store one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the self-organizing mapping weighted particle swarm mean clustering method as described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the self-organizing mapping weighted particle swarm mean clustering method as described in the first aspect is implemented.
[0018] By providing a self-organizing mapping weighted particle swarm mean clustering method, device, equipment and storage medium, the method of the embodiment of the present invention includes: obtaining original sample data; performing rough clustering on the original sample data by using a SOM clustering algorithm to obtain K rough clustering clusters and rough clustering centers; performing fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers. It can be seen that through this method, clustering of the original sample data can be realized. First, performing rough clustering on the original sample data by using the SOM clustering algorithm can obtain K rough clustering clusters and rough clustering centers, thereby determining the number of clustering clusters and the initial clustering centers. Then, based on the determined rough clustering clusters and rough clustering centers, fine clustering is performed on the original sample data to obtain target clustering centers. Since the initial number of clustering clusters and the clustering centers for fine clustering are determined, compared with the situation where the initial values of traditional clustering algorithms are random, the present application can improve the accuracy and effect of clustering, and solve the problem that traditional clustering algorithms are sensitive to initial values. Description of the Drawings
[0019] Figure 1 is a flowchart of a self-organizing mapping weighted particle swarm mean clustering method in Embodiment 1 of the present invention;
[0020] Figure 2 is a flowchart of a self-organizing mapping weighted particle swarm mean clustering method in Embodiment 2 of the present invention;
[0021] Figure 3 is a flowchart of a self-organizing mapping weighted particle swarm mean clustering method in Embodiment 3 of the present invention;
[0022] Figure 4 It is a flowchart of a weight particle swarm mean clustering method based on self-organizing mapping in the fourth embodiment of the present invention;
[0023] Figure 5 It is a comparison diagram of algorithm purity in the fourth embodiment of the present invention;
[0024] Figure 6 It is a convergence effect diagram of the SOM&WPSM clustering algorithm in the fourth embodiment of the present invention;
[0025] Figure 7 It is a convergence effect diagram of the traditional PSO / K-means clustering algorithm in the fourth embodiment of the present invention;
[0026] Figure 8 It is a structural block diagram of a weight particle swarm mean clustering device based on self-organizing mapping in the fifth embodiment of the present invention;
[0027] Figure 9 It is a schematic structural diagram of an electronic device in the sixth embodiment of the present invention. Detailed implementation manners
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.
[0029] Through the research of the inventor, it is found that in recent years, big data and artificial intelligence have been highly pursued and are extremely popular, and information has also shown an explosive growth. What we are facing will be massive amounts of text, video, picture, and audio data. How to mine information with practical value from the huge-scale data has gradually become an important topic in the field of computer research. Traditional data mining mainly relies on manual experience and teamwork to complete, but this method undoubtedly increases the time and labor costs, and even the final results are not satisfactory. Therefore, data mining technology was born, aiming to help people extract potential and valuable information from massive and disordered data. As an important branch in the field of data mining, clustering algorithms have been widely applied in many fields, including machine learning, pattern recognition, image analysis, information retrieval, computer vision, etc. An efficient clustering algorithm can improve work efficiency and work quality.
[0030] Clustering algorithms divide samples into different distributed clusters according to the similarity of sample features. Samples within a class have high text feature similarity, while samples between classes are far apart in distribution and have large feature differences. As a branch of data mining, clustering analysis has been studied for many years, and a large number of clustering algorithms have emerged. However, since different data have their own forms and dimensions, no algorithm is universal for data, and various algorithms adopt specific clustering schemes for specific data. For example, some clustering algorithms have significant clustering results for medium and low-dimensional data but perform poorly on high-dimensional data; some clustering algorithms can only target data with special distribution structures and cannot handle data with other distributions well. These characteristics require the algorithm to be scalable, capable of processing different types of data, and also able to discover clusters of various shapes, and solve the problems of "noise" and outliers. Traditional clustering algorithms can no longer solve the above problems. Some scholars have applied swarm intelligence optimization algorithms to traditional algorithms and found that clustering algorithms based on swarm intelligence optimization have better clustering effects than traditional algorithms.
[0031] Partition-based clustering algorithms divide the sample set into multiple non-overlapping clusters according to distance rules and stop clustering until the objective function is minimized through iteration. This method is easy to implement and has a fast convergence speed, but its complexity is linearly related to the sample size, sample dimension, and cluster center. The K-means algorithm proposed by MacQueen is a classic partition-based clustering algorithm. This algorithm combines simple calculation and fast convergence and uses distance as an evaluation index for sample similarity. However, the K-means clustering algorithm has inherent drawbacks: (1) the determination of the number of clusters K, and the choice of the index for selecting the value of K will directly affect the clustering accuracy; (2) the selection of the initial cluster center, and the clustering effect depends on the initialization of the cluster center. The fused clustering algorithm is sometimes better than a single algorithm in terms of accuracy and convergence. Some scholars have also combined clustering algorithms, and this process called clustering inheritance can provide stronger robustness and stability solutions in different neighborhoods and data. Some scholars have integrated the particle swarm optimization and K-means data clustering algorithms, combining the advantages of both to improve the quality of the clustering algorithm, but still have not solved the problem of the value of K.
[0032] In summary, the existing technical problems are as follows: It is difficult for traditional clustering algorithms to determine the number of clustering clusters. The selection of the number of clustering clusters will directly affect the accuracy of the clustering result. If the selected number of clustering clusters is too large, it will lead to redundant clustering results. It is possible that the samples of one class are divided into two classes, or the samples of two different classes are cross-divided. The selection of the initial clustering center will also affect the final clustering result. If the initial clustering centers are selected too densely, it will lead to the crossing of samples of multiple classes in the clustering division, and it is impossible to label the divided classes well. If the initial clustering centers are selected too far apart, it is possible that the selected center points are edge points or abnormal points, resulting in a worthless clustering result. Therefore, the selection of the clustering center and the number of clustering clusters should be reasonable. The combined clustering algorithms generally have the disadvantage of being easily trapped in local optima. If a good fitness value is reached, the iteration of the clustering algorithm will stop, and thus the global optimum cannot be found.
[0033] Reasons for the ineffective solution of technical problems: Different clustering algorithms are applicable to specific data, and there is no clustering algorithm that can be applied to all data. Since it belongs to unsupervised learning, the training of word vectors directly affects the clustering result. Without good corpus, the clustering algorithm will become meaningless. In addition, a good clustering algorithm can effectively divide the sample distribution and make the final clustering result meaningful.
[0034] The difficulty in solving lies in: Since the objective function of K-means is not a convex function and may contain many local minima, therefore, the K-means algorithm mainly has two inherent disadvantages: (1) Random selection of initial values may lead to different clustering results, and there may even be a situation of no solution; (2) This algorithm is an algorithm based on the objective function and usually uses the gradient method to solve the extreme value. Since the search direction of the gradient method is along the direction of energy reduction, the algorithm is easily trapped in local extrema. These two major defects greatly limit its application scope.
[0035] Based on this, the embodiments of the present invention provide a weight particle swarm mean clustering method for self-organizing mapping, which can realize clustering analysis of sample data. By using the self-organizing mapping clustering algorithm (Self Organizing FeatureMaps, SOM) for rough clustering, the number of clustering clusters and the initial clustering center can be determined. Then, based on the determined rough clustering clusters and rough clustering centers, the original sample data is finely clustered to obtain the target clustering center, which can improve the accuracy and clustering effect of clustering and solve the problem that traditional clustering algorithms are sensitive to initial values.
[0036] Embodiment 1
[0037] Figure 1The following is a flowchart of a weighted particle swarm mean clustering method based on self-organizing mapping provided in Embodiment 1 of the present invention. This embodiment is applicable to a method for accurately and efficiently clustering a large amount of sample data in a data mining processing platform. This method can be executed by a weighted particle swarm mean clustering device based on self-organizing mapping. The device can be implemented in software and / or hardware, and can be configured in the server of the processing platform. The specific steps are as follows:
[0038] Step 110: Obtain the original sample data; where the original sample data is software portrait data.
[0039] Among them, the original sample data can be pre-stored in the server and can be obtained from the server when clustering is required. The original sample data is the sample data that the user needs to perform clustering analysis on. For example, when the user needs to cluster a large amount of software data, since different software requirements correspond to different software types, the original sample data is a set of software portrait data containing various functional types of software. For example, it is a set of software data containing various functional types of software such as security software, drawing software, development software, and teaching software. Among them, the original sample data is all in vector form and can be vector data of any dimension. The software portrait data is data describing software features, and can describe software features from multiple dimensions. For example, security software can include multiple dimensions such as application industry, usage, data type, download times, access volume, and security level; drawing software can include multiple dimensions such as the industry it belongs to, drawing type, usage, and download times; development software can include multiple dimensions such as development scenario, data type, download times, and access volume; teaching software can include multiple dimensions such as the teaching type it belongs to, function, teaching display type, and play times.
[0040] Step 120: Use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; where K is a natural number;
[0041] Among them, the number of rough clustering clusters is the same as the number of rough clustering centers, and each rough clustering cluster corresponds to a rough clustering center. Among them, K is a natural number, and the specific value of K is related to the actual original sample data and the specific rough clustering result.
[0042] Among them, the SOM clustering algorithm uses an unsupervised learning method for iterative training. It can map the original sample data of any dimension to a low-dimensional space, which not only reduces the vector dimension, but also reduces the computational complexity of iterative training, and at the same time maintains the original topological structure of the original sample data. Therefore, after the original sample data is roughly clustered by the SOM clustering algorithm, on the one hand, the determined number of rough clustering clusters and the number of rough clustering centers can be obtained; on the other hand, the original topological structure of the original sample data remains unchanged, and the samples when entering fine clustering are still the original sample data, so as to ensure the consistency and stability of the data.
[0043] Step 130: Perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain the target clustering center.
[0044] Among them, the target clustering center is the optimal clustering center finally obtained according to the rough clustering and fine clustering processes. Specifically, perform fine clustering on the original sample data according to the rough clustering clusters and the rough clustering centers determined after rough clustering. As the fine clustering progresses, the clustering clusters where the original sample data is located may change, but the number of clustering clusters in fine clustering remains unchanged, which is the same as the number of rough clustering clusters. Thus, using the determined rough clustering clusters and rough clustering centers as the initial number of clustering clusters and clustering centers in fine clustering can improve the clustering accuracy and clustering effect of fine clustering.
[0045] In the technical solution of this embodiment, the working principle of the self-organizing mapping weighted particle swarm mean clustering method is as follows: First, obtain the original sample data; then, use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers, where K is a natural number; finally, perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain the target clustering center. It can be seen that when a user needs to perform clustering analysis on data, such as clustering software-related data, only need to store the software-related data samples in the server. The server obtains the software-related data samples, and then uses the SOM clustering algorithm to perform rough clustering on the software-related data samples. After rough clustering, a determined number of rough clustering clusters and rough clustering centers will be obtained. For example, K rough clustering clusters and rough clustering centers are obtained; finally, perform fine clustering on the original sample data based on the determined rough clustering clusters and rough clustering centers to obtain the target clustering center. Since the initial number of clustering clusters and clustering centers in fine clustering are determined, compared with the situation where the initial values of traditional clustering algorithms are random, this application can improve the clustering accuracy and clustering effect and solve the problem that traditional clustering algorithms are sensitive to initial values.
[0046] Generally, traditional clustering algorithms are designed to cluster specific samples using specific clustering algorithms. For example, the K-means clustering algorithm is mainly used to process spherical cluster data and cannot handle data with non-spherical clusters, different sizes, and different densities. Some clustering algorithms have significant clustering results for medium and low-dimensional data but perform poorly on high-dimensional data. Some clustering algorithms can only handle data with special distribution structures and cannot handle other distributed data well. The original sample data in the embodiments of this application can be spherical data, non-spherical data, etc. The self-organizing mapping weighted particle swarm mean clustering method provided in this embodiment can be applied to medium and low-dimensional data, high-dimensional data, and data with special distribution structures such as spherical data and non-spherical data. Compared with traditional clustering algorithms, the applicability to samples is improved.
[0047] The technical solution of this embodiment provides a self-organizing mapping weighted particle swarm mean clustering method, which includes: obtaining original sample data; using the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; and performing fine clustering on the original sample data based on the rough clustering clusters and rough clustering centers to obtain target clustering centers. It can be seen that through this method, clustering of the original sample data can be achieved. First, using the SOM clustering algorithm to perform rough clustering on the original sample data can obtain K rough clustering clusters and rough clustering centers, thereby determining the number of clustering clusters and the initial clustering centers. Then, based on the determined rough clustering clusters and rough clustering centers, fine clustering is performed on the original sample data to obtain the target clustering centers. Since the number of initial clustering clusters and the clustering centers for fine clustering are determined, compared with the situation where the initial values of traditional clustering algorithms are random, this application improves the accuracy and effect of clustering and solves the problem of the sensitivity of traditional clustering algorithms to initial values.
[0048] Embodiment 2
[0049] Figure 2 is a flowchart of a self-organizing mapping weighted particle swarm mean clustering method provided in Embodiment 2 of the present invention. On the basis of the above Embodiment 1, optionally, using the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers, refer to Figure 2 , specifically including the following steps:
[0050] Step 210, obtain the original sample data;
[0051] Step 220, use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters;
[0052] Step 230, calculate the sample mean of the original sample data in each rough clustering cluster according to the K rough clustering clusters;
[0053] Among them, the original sample data contains multiple sample data. After rough clustering the original sample data, K rough clustering clusters can be obtained. Each rough clustering cluster contains several original sample data. According to the original sample data in each rough clustering cluster, calculate the sample mean of all the original sample data contained in the corresponding rough clustering cluster.
[0054] Step 240: According to the sample mean of each rough clustering cluster, use the nearest neighbor principle to find the sample in each rough clustering cluster that is closest to its sample mean as the rough clustering center of the corresponding rough clustering cluster.
[0055] Specifically, after rough clustering the original sample data, K rough clustering clusters can be obtained. Then, based on the determined K rough clustering clusters, find the rough clustering center of each rough clustering cluster. Among them, the method for finding the rough clustering center of each rough clustering cluster is the same. For example, taking the search for the rough clustering center of the Kth rough clustering cluster as an example, in order to find the clustering center of the Kth rough clustering cluster, first, calculate the sample mean of the Kth rough clustering cluster (the Kth rough clustering cluster contains several original sample data, and the specific number is determined by the actual rough clustering situation), and then use the nearest neighbor principle to find the original sample data in the Kth rough clustering cluster that is closest to its sample mean as the clustering center of the Kth rough clustering cluster. According to this method, the rough clustering centers of the other K - 1 rough clustering clusters can be found.
[0056] Step 250: Based on the rough clustering clusters and the rough clustering centers, perform fine clustering on the original sample data to obtain the target clustering center.
[0057] On the basis of the above technical solution, optionally, using the nearest neighbor principle to find the original sample data in each rough clustering cluster that is closest to its sample mean as the rough clustering center of the corresponding rough clustering cluster includes:
[0058] Calculate the Euclidean distance between the sample mean of each rough clustering cluster and each original sample data in its corresponding rough clustering cluster, and use the original sample data with the smallest Euclidean distance value as the rough clustering center of the corresponding rough clustering cluster.
[0059] Among them, the method for finding the rough clustering center of each rough clustering cluster is the same. Exemplarily, taking the search for the rough clustering center of the Kth rough clustering cluster as an example, first, calculate the sample mean of the Kth rough clustering cluster, and then calculate the Euclidean distances between the sample mean of the Kth rough clustering cluster and all the original sample data in the Kth rough clustering cluster respectively. Use the original sample data with the smallest Euclidean distance value in the Kth rough clustering cluster as the rough clustering center of the Kth rough clustering cluster. According to this method, the rough clustering centers of the other K - 1 rough clustering clusters are found.
[0060] Example Three
[0061] Figure 3 This is the flowchart of a weight particle swarm mean clustering method for self-organizing mapping provided in Embodiment 3 of the present invention. On the basis of the above embodiments, optionally, the SOM clustering algorithm is used to perform rough clustering on the original sample data to obtain K rough clustering clusters. Refer to Figure 3 , and specifically includes the following steps:
[0062] Step 310: Obtain the original sample data;
[0063] Step 320: Input the original sample data into the rough clustering model based on the SOM clustering algorithm for iterative processing;
[0064] Exemplarily, assume that the original sample data is the training set x i , where i = 1, 2, 3,..., m, the dimension of each sample is n, the competitive layer adopts the form of a matrix neuron array, the output matrix size is n*k, and the dimension of each competitive layer neuron w j is k, where j = 1, 2,..., n. Among them, the rough clustering model of the SOM clustering algorithm is specifically as follows:
[0065] Step S11: First, perform initialization, that is, set the training set x i , the first preset number of iterations interation, and the competitive layer matrix n*k, and randomly generate the weight w of the competitive layer neurons according to the training set x i , the first preset number of iterations interation, and the competitive layer matrix n*k j .
[0066] Step S12: Perform normalization processing, that is, normalize the training set x i and the weight w of the competitive layer neurons j , and calculate the inner product of the normalized input training set x i and the weight w of the competitive layer neurons j . Among them, the index of the dimension with the largest inner product is the winning neuron index winner.
[0067] Among them, the normalization equations of the training set x i and the weight w of the competitive layer neurons j are respectively:
[0068]
[0069]
[0070] Among them, the winning neuron index winner is:
[0071] winner = argmax||train_x i *wj ||
[0072] Step S13: Define the neighborhood function N j * (t), and the winning neighborhood gradually shrinks as time decreases. The winning neighborhood is represented by the neighborhood radius r:
[0073]
[0074] where C 1 is a positive constant related to the neuron output node, and the specific value is determined by the minimum value of the actual neuron output.
[0075] Step S14: Adjust the neuron weights, that is, adjust the weights of all neurons within the winning domain N j * (t).
[0076] The adjustment formula is as follows:
[0077] w ij (t + 1) = w ij (t) + σ(t, N)[x i - w ij (t)]
[0078] where σ(t, N) represents the learning rate, and the learning rate also decreases with time. σ(t, N) is:
[0079]
[0080] where t is the time, n = 1, 2, 3, ……, r.
[0081] Step S15: Perform iterative loop. If the first preset iteration number is reached, then execute Step 330; otherwise, repeat Steps S12, S13, and S14.
[0082] Step 330: When the iteration number reaches the first preset iteration number interation, output the weights of each neuron of the rough clustering model when the first preset iteration number interation is reached;
[0083] By setting the first preset iteration number interation, the original sample data stops iterating when it loops to the first preset iteration number interation in the rough clustering model, so as to achieve rough clustering of the original sample data, which can avoid problems such as too long training time, unsaturated clustering of data, and the occurrence of dead neurons due to excessive loop iteration times.
[0084] Among them, the first preset number of iterations interation can be set according to the actual situation and will not be specifically limited here.
[0085] Step 340: Calculate the inner product of each original sample data and each neuron weight in sequence to form an inner product matrix, and use the element position of the maximum inner product of each original sample data in the inner product matrix as the winning index of each original sample data.
[0086] Specifically, calculate the inner product of each original sample data and each neuron weight. For example, assume there are j neuron weights. Taking the calculation of the inner product of the i-th original sample data as an example, calculate the inner product of the i-th original sample data and the j neuron weights respectively, and then in the same way, calculate the inner product of the other i - 1 original sample data and the j neuron weights respectively, and form an inner product matrix with all the calculated inner products.
[0087] Use the element position of the maximum inner product of each original sample data in the inner product matrix as the winning index of each original sample data. For example, taking the i-th original sample data as an example, calculating the inner product of the i-th original sample data and the j neuron weights can obtain j inner product values, and use the element position corresponding to the maximum inner product value among these j inner product values in the inner product matrix as the winning index of the i-th original sample data. In the same way, the winning indices of all original sample data can be obtained.
[0088] Step 350: Classify the original sample data corresponding to the winning indices with the same inner product value to obtain K rough clustering clusters belonging to K classes.
[0089] Specifically, classify the original sample data corresponding to all winning indices with the same inner product value into the same class. From this, K rough clustering clusters can be planned (the number of K is related to the actual rough clustering situation), thereby realizing the preliminary sample division of the original sample data.
[0090] Step 360: Perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain the target clustering centers.
[0091] Embodiment Four
[0092] Figure 4 It is a flowchart of a weight particle swarm mean clustering method for self-organizing mapping provided in Embodiment Four of the present invention. On the basis of the above embodiments, optionally, perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain the target clustering centers. Refer to Figure 4 , specifically including the following steps:
[0093] Step 410: Obtain the original sample data.
[0094] Step 420: Coarsely cluster the original sample data using the SOM clustering algorithm to obtain K coarse clustering clusters and coarse clustering centers; where K is a natural number;
[0095] Step 430: Initialize the positions and velocities of multiple particles using the original sample data;
[0096] Specifically, since the SOM clustering algorithm has the advantage of maintaining the original topological structure of the original sample data, after the original sample data is coarsely clustered by the SOM clustering algorithm, its topological structure remains unchanged. That is, when performing fine clustering, the original sample data does not change. Because the original sample data are all in vector form, the original sample data can be used to initialize the positions and velocities of multiple particles, that is, the vector representation of the original sample data is used to represent the initial positions and initial velocities of multiple particles.
[0097] Among them, the position of each particle is represented by the sample weight of the original sample data, and a coding format based on the sample weight can be used to represent it. In addition to the position, there are also velocity and fitness value. The dimension of the velocity is the same as the dimension of the original sample data.
[0098] Exemplarily, the coding method of the particle is as follows:
[0099] <![CDATA[x 1 ,x 2 ,x 3 ,…,x m > <![CDATA[v 1 ,v 2 ,v 3 ,…,v m > fitness(x,c)
[0100] Among them, after initializing the positions and velocities of multiple particles using the original sample data, it also includes: calculating the fitness value of each current particle and the fitness value of the particle population according to the position and velocity of each initialized particle, and then, according to the fitness value of each particle, selecting the local optimal position vector pbest of each particle, and selecting the global optimal position vector gbest of the particle population according to the fitness value of the particle population.
[0101] Step 440: Initialize the coarse clustering clusters and coarse clustering centers as the current clustering clusters and current clustering centers;
[0102] Specifically, the rough clustering clusters and rough clustering centers obtained by performing rough clustering on the original sample data according to the SOM clustering algorithm are initialized as the current clustering clusters and the current clustering centers, that is, the rough clustering clusters are used as the initial clustering clusters for fine clustering, and the rough clustering centers are used as the initial clustering centers of the respective initial clustering clusters for fine clustering. Compared with the traditional clustering method where it is difficult to determine the number of clustering clusters, the initial values are set randomly, and the clustering effect is prone to be poor, the present application can achieve: determining the initial number of clustering clusters for fine clustering after rough clustering by the SOM clustering algorithm, thereby avoiding the problem of sensitive selection of initial values, improving the accuracy and clustering effect of clustering, and since the SOM clustering algorithm performs a certain degree of rough clustering on the original sample data, making the division of samples and clustering centers develop towards a better clustering direction, thus reducing the occurrence of empty clusters.
[0103] Step 450: Calculate the distances from each particle to the K current clustering clusters in sequence according to the Euclidean distance formula, and divide each particle into the nearest current clustering cluster according to the nearest neighbor principle;
[0104] Exemplarily, taking the i-th particle as an example, calculate the distances from the i-th particle to the K current clustering clusters according to the Euclidean distance formula, obtaining K distance values. Then, divide the i-th particle into the current clustering cluster corresponding to the minimum distance value among the K distance values according to the nearest neighbor principle, thereby completing the re-division of the i-th particle. In this way, all particles are re-sampled, that is, the clustering clusters to which the particles belong are re-divided, and the current clustering clusters are updated. It should be noted that no matter how many times the update iteration is performed, the number of clusters of the current clustering clusters is the same as the number of rough clustering clusters obtained after rough clustering by the SOM clustering algorithm.
[0105] Among them, the Euclidean distance formula is:
[0106] distance(x i ,C j )
[0107] Among them, x i is a particle, i = 1, 2, 3,..., m; C j is the current clustering center;
[0108] Among them, dividing each particle into the nearest current clustering cluster according to the nearest neighbor principle means satisfying:
[0109] distance(x i ,C j )=min{distance(x i ,C j ),j=1,2,3…,k}
[0110] Then x i ∈yj , where y j is the set of samples belonging to the same cluster after sample division.
[0111] Among them, by calculating the distances from each particle to the K current clusters in turn according to the Euclidean distance formula and dividing each particle into the nearest current cluster according to the nearest neighbor principle, the advantages are as follows: on the one hand, the particles can be re-divided to obtain new clusters, that is, new current clusters; on the other hand, the influence of noise points can be reduced, the cluster means can be mapped to sample points, which is beneficial to class label annotation, not only accelerating the convergence speed but also preventing the occurrence of empty clusters.
[0112] Step 460: Update the current cluster centers according to the cluster center calculation formula;
[0113] Specifically, after updating the current clusters, since the particles included in the current clusters have changed, it is necessary to update the cluster centers of each current cluster. Specifically, update the K current cluster centers in turn according to the cluster center calculation formula. For example, taking the Kth current cluster updated in step 450 as an example, update the cluster center of the Kth current cluster according to the cluster center calculation formula. In this way, update the cluster centers of all the current clusters updated in step 450 according to the cluster center calculation formula.
[0114] Among them, C j ′ represents the cluster mean. The sample closest to C j ′ is taken as the final cluster center through Euclidean distance calculation. The calculation formula of the cluster center C j is:
[0115]
[0116] C j = data[argmin(distance(C j ′ , x i ))] x i ∈ y i
[0117] In addition, the calculation formula of the within-class cohesion E is:
[0118]
[0119] Among them, the within-class cohesion E is used to evaluate the quality of the within-class cohesion of each cluster.
[0120] Step 470: Update the positions and velocities of each particle, and return to perform the operation of dividing the particles into the current clustering cluster until the iteration stops when the second preset number of iterations is reached, and use the current clustering center when reaching the second preset number of iterations as the target clustering center.
[0121] Specifically, according to the fitness value of each particle, select the local optimal position vector pbest of each particle, select the global optimal position vector gbest of the particle population according to the fitness value of the particle population, then update the positions and velocities of each particle according to the local optimal position vector pbest of the particle and the global optimal position vector gbest of the particle population, and after each update of the positions and velocities of the particles, calculate the fitness value of the particles after updating the positions and velocities and the fitness value of the particle population, and then, according to the current fitness value of each particle, update the local optimal position vector pbest of the current particle, and update the global optimal position vector gbest of the particle population according to the current fitness value of the particle population. In addition, calculating the fitness value of the particles and the fitness value of the particle population after each update of the positions and velocities of the particles can also facilitate the update of the positions and velocities of the particles during subsequent iterations.
[0122] After updating the positions and velocities of the particles, as well as the local optimal position vector pbest of the current particle and the global optimal position vector gbest of the population, return to Step 450 to perform the operation of dividing the particles into the current clustering cluster until the iteration stops when the second preset number of iterations is reached, and use the current clustering center when reaching the second preset number of iterations as the target clustering center.
[0123] Exemplarily, assume that the original sample data is the training set x i , where i = 1, 2, 3,..., m, and the dimension of each sample is n. Then, in the n-dimensional space, there are m particles. At a certain moment, the position of the particle can be expressed as:
[0124] x i =(x i1 , x i2 , x i3 , …, x in )
[0125] where i = 1, 2, 3,..., m;
[0126] The velocity of the particle is a random distribution between 0 and 1:
[0127] v i =(v i1 , v i2 , v i3 , ……, v in )
[0128] where \(i = 1, 2, 3, \cdots, m\);
[0129] Each particle represents a feasible solution in an \(N\)-dimensional space. All particles have two attributes: position and velocity. The position represents the direction of particle movement, and the velocity represents the speed of particle movement. In each iteration, the particle determines its next updated position and velocity based on its historical local optimal position vector \(pbest\) and the global optimal position vector \(gbest\) of the population. The position and velocity update formulas for the particle are as follows:
[0130] Among them, the velocity update formula for the particle is as follows:
[0131] v i (t + 1)=w\times v i (t)+c 1 r 1 [pbest - x i (t)]+c 2 r 2 [gbest - x i (t)]
[0132] Among them, the position update formula for the particle is as follows:
[0133] x i (t + 1)=x i (t)+v i (t + 1)
[0134] Among them, v i (t + 1) represents the velocity of the \(i\)-th particle at time \(t + 1\), \(pbest\) represents the local optimal position vector of the individual particle, \(gbest\) represents the global optimal position vector of the population, that is, the minimum value of the best position of all particles. \(c1\) and \(c2\) represent learning factors, generally taking the value of 2, \(w\) represents the inertia factor, and \(r1\) and \(r2\) are two random numbers between \((0, 1)\).
[0135] Among them, the calculation formula for the fitness value \(pf\) of each particle is:
[0136] pf=\(\|x i - centroids i \|\)
[0137] Among them, x i - centroids i represents the minimum Euclidean distance between particle \(i\) and its cluster center.
[0138] The local optimal position vector \(pbest\) of the particle is expressed as:
[0139] pbest=(p i1 , p i2 , pi3 , ……, p in )
[0140] Among them, p in represents the particle vector, i = 1, 2, 3, ……, m;
[0141] Among them, the selection of the fitness function of the particle swarm directly affects the convergence speed of the clustering algorithm and whether the optimal solution can be found. Having an overall understanding of the clustering division and judging the accuracy rate of the clustering result, the mean clustering algorithm can use the criterion function for evaluating the clustering quality as the fitness function of the particle swarm. The within-class compactness MSE is selected to represent the quality of the clustering. The smaller the mse, the better the clustering effect.
[0142]
[0143] Among them, C j represents the clustering center; D i represents the set of particles, that is, the original sample data set.
[0144] The fitness value of the particle swarm represents the similarity between the data objects within each class. The smaller the fitness value, the closer the combination degree of the data objects within the class, and the better the clustering effect. The calculation formula for the fitness value of the particle swarm is:
[0145] fitness = mse
[0146] On the basis of the above technical solution, optionally, the second preset number of iterations is the number of iterations when the fitness value of the particle and the fitness value of the population where the particle is located are the smallest.
[0147] Among them, the smaller the fitness value of the particle, the better the position of the particle, and the particle tends to optimize in the direction of the clustering center; the smaller the fitness value of the population where the particle is located, the closer the combination degree of the original sample data within the clustering cluster, the better the clustering effect, and the better the global convergence effect of the algorithm. Therefore, when the fitness value of the particle and the fitness value of the population where the particle is located are the smallest, the clustering effect is the best, and the global convergence of the algorithm is also the best. At this time, the iteration stops, and the clustering center of the clustering cluster where the particle is located at this time is the optimal clustering center.
[0148] Optionally, the number of particles is the same as the number of original sample data.
[0149] Since the SOM clustering algorithm has the advantages of being able to reduce the dimension of samples, having simple calculations for iterative training, and maintaining the original topological structure of the original sample data unchanged after the iterative training is completed, when performing fine clustering on particles, the K coarse clustering clusters obtained by performing coarse clustering on the original sample data through the SOM clustering algorithm are used as the initial clustering clusters for fine clustering, and multiple particles are initialized using the original sample data, so the number of particles is the same as the number of original sample data.
[0150] Thus, by determining the initial clustering clusters for fine clustering through the SOM example algorithm and initializing the initial positions and velocities of the particles using the original sample data, the convergence speed of the algorithm can be accelerated, the accuracy and clustering effect of clustering can be improved, and the occurrence of empty clusters can be prevented.
[0151] Exemplarily, in order to verify the effectiveness and feasibility of the algorithm, in this embodiment, the Iris dataset, Wine dataset, and Glass dataset in the UCI database are used for experiments, and the samples of these three datasets are all in vector form.
[0152] Among them, the Iris dataset contains 150 data samples, divided into 3 categories, with 50 data in each category. Each data contains 4 attributes, and it is predicted which category among (Setosa, Versicolour, Virginica) the iris flower belongs to through 4 attributes: sepal length, sepal width, petal length, and petal width. The Wine dataset contains a total of 178 records from 3 different origins of wine, and 13 attributes are 13 chemical components of the wine, and the origin of the wine can be inferred through chemical analysis. The Glass dataset contains 214 data samples, divided into 3 categories, and each data contains 9 attributes.
[0153] To verify the feasibility of the algorithm, exemplarily, taking the original sample data using purity, fitness value, convergence curve, Davies - Bouldin index, Dunn's index, and Silhouette coefficient as an example, they are used as the evaluation criteria for the clustering results.
[0154] Table 1 shows a comparison of the clustering accuracies of K - means, pso - km, and the Self Organizing Feature Maps and Weight Particle Swarm Means (SOM&WPSM) clustering algorithm of the present application. Refer to Table 1.
[0155] Table 1 Comparison of Clustering Accuracies
[0156]
[0157] Table 2 Comparison of fitness values
[0158]
[0159] Table 2 shows the comparison of the fitness values of the traditional clustering algorithms K-means, pso-km and SOM&WPSM provided in this application. For details, please refer to Table 2.
[0160] The verification experiment also makes comparisons in terms of the classification fitness index (Davies-Bouldin index, DBI), Dunn Validity Index (DVI) and Silhouette coefficient (SC), as shown in Table 3, Table 4 and Table 5.
[0161] Among them, the smaller the DBI, the smaller the intra-class distance and the larger the inter-class distance, that is, the smaller the DBI, the better the clustering effect. The larger the DVI, the smaller the intra-class distance and the larger the inter-class distance, and the better the clustering effect. The value of SC ranges from [-1, 1], and the closer it is to 1, the relatively better the cohesion and separation degree are.
[0162] Among them, Table 3 shows the comparison of the K-means clustering algorithm, Pso-km clustering algorithm and SOM&WPSM clustering algorithm when the original sample data is the Iris dataset. For details, please refer to Table 3.
[0163] Table 3 Comparison of Iris datasets
[0164]
[0165] Table 4 shows the comparison of the K-means clustering algorithm, Pso-km clustering algorithm and SOM&WPSM clustering algorithm when the original sample data is the Wine dataset. For details, please refer to Table 4.
[0166] Table 4 Comparison of Wine datasets
[0167]
[0168] Table 5 shows the comparison of the K-means clustering algorithm, Pso-km clustering algorithm and SOM&WPSM clustering algorithm when the original sample data is the Glass dataset. For details, please refer to Table 5.
[0169] Table 5 Comparison of Glass datasets
[0170]
[0171] Figure 5 It is a comparison graph of algorithm purity provided in the fourth embodiment of the present invention;Figure 6 It is the convergence effect diagram of the SOM&WPSM clustering algorithm provided in the fourth embodiment of the present invention; Figure 7 It is the convergence effect diagram of the traditional PSO / K-means clustering algorithm provided in the fourth embodiment of the present invention. To verify the clustering purity of the algorithm, exemplarily, the traditional SOM clustering algorithm, PSO / K-means clustering algorithm, K-means clustering algorithm and the SOM&WPSM of the present application are compared in terms of purity. The experiment is repeated 20 times on each dataset, and the highest accuracy value is taken for comparison, as Figure 5 shown. As can be seen from Figure 5 it, the purity of the SOM&WPSM algorithm of the present application is the best.
[0172] To verify the convergence degree of the algorithm, the traditional PSO / K-means clustering algorithm and the SOM&WPSM algorithm of the present application are compared based on the Iris dataset as the original sample data, as Figure 6 and Figure 7 shown. As can be seen from Figure 6 and Figure 7 it, the convergence of the SOM&WPSM algorithm of the present application is the best.
[0173] Embodiment Five
[0174] Figure 8 It is the structural block diagram of a self-organizing mapping weight particle swarm mean clustering device provided in the fifth embodiment of the present invention. Referring to Figure 8 , the self-organizing mapping weight particle swarm mean clustering device 100 includes:
[0175] A data acquisition module 10 for acquiring original sample data;
[0176] A rough clustering module 20 for roughly clustering the original sample data using the SOM clustering algorithm to obtain K rough clustering clusters and rough clustering centers; where K is a natural number;
[0177] A fine clustering module 30 for finely clustering the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers.
[0178] In the technical solution of this embodiment, by providing a weight particle swarm mean clustering device for self-organizing mapping, the device includes: a data acquisition module for acquiring original sample data; a rough clustering module for performing rough clustering on the original sample data using the SOM clustering algorithm to obtain K rough clustering clusters and rough clustering centers, where K is a natural number; a fine clustering module for performing fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers. It can be seen that through this device, clustering of the original sample data can be achieved. First, using the SOM clustering algorithm to perform rough clustering on the original sample data can obtain K rough clustering clusters and rough clustering centers, thereby determining the number of clustering clusters and the initial clustering centers. Then, based on the determined rough clustering clusters and rough clustering centers, fine clustering is performed on the original sample data to obtain the target clustering centers. Since the initial number of clustering clusters and the clustering centers for fine clustering are determined, compared with the situation where the initial values of traditional clustering algorithms are random, the present application can improve the accuracy and effect of clustering, solve the problem that traditional clustering algorithms are sensitive to initial values, and there is no need to label the samples during clustering, which can be applicable to the clustering of all samples and can improve the applicability to samples.
[0179] Optionally, the rough clustering module 20 includes:
[0180] A rough clustering cluster obtaining unit for performing rough clustering on the original sample data using the SOM clustering algorithm to obtain K rough clustering clusters;
[0181] A mean value calculation unit for calculating the sample mean of the original sample data in each rough clustering cluster according to the K rough clustering clusters;
[0182] A rough clustering center obtaining unit for finding, according to the sample mean of each rough clustering cluster, the sample closest to its sample mean in each rough clustering cluster as the rough clustering center of the corresponding rough clustering cluster using the nearest neighbor principle.
[0183] Optionally, the rough clustering center obtaining unit is further configured to calculate the Euclidean distance between the sample mean of each rough clustering cluster and each original sample data in the corresponding rough clustering cluster, and use the original sample data with the smallest Euclidean distance value as the rough clustering center of the corresponding rough clustering cluster.
[0184] Optionally, the rough clustering module 20 further includes:
[0185] A loop iteration processing unit for inputting the original sample data into a rough clustering model based on the SOM clustering algorithm for loop iteration processing;
[0186] A neuron weight output unit for outputting the weights of each neuron of the rough clustering model when the iteration number reaches the first preset iteration number;
[0187] An inner product calculation unit for sequentially calculating the inner products of each original sample data and each neuron weight to form an inner product matrix;
[0188] A winning index acquisition unit for using the element position of the maximum inner product of each original sample data in the inner product matrix as the winning index of each original sample data;
[0189] A rough clustering cluster obtaining unit is further configured to classify the original sample data corresponding to the winning indices with the same inner product value to obtain K rough clustering clusters belonging to K classes.
[0190] Optionally, the fine clustering module 30 includes:
[0191] A particle position and velocity initialization unit for initializing the positions and velocities of multiple particles using the original sample data;
[0192] A current clustering cluster and current clustering center initialization unit for initializing the rough clustering clusters and rough clustering centers as the current clustering clusters and current clustering centers;
[0193] A particle partitioning unit for sequentially calculating the distances of each particle to the K current clustering clusters according to the Euclidean distance formula, and partitioning each particle into the nearest current clustering cluster according to the nearest neighbor principle;
[0194] A clustering center calculation and update unit for updating the current clustering center according to the clustering center calculation formula;
[0195] A particle position and velocity update unit for updating the positions and velocities of each particle;
[0196] A return execution particle partitioning and iteration unit for returning to execute the operation of partitioning particles into the current clustering clusters until the iteration stops when the second preset number of iterations is reached;
[0197] A target clustering center determination unit for using the current clustering center when the second preset number of iterations is reached as the target clustering center.
[0198] Optionally, the second preset number of iterations is the number of iterations when the fitness value of the particle and the fitness value of the population where the particle is located are the smallest.
[0199] Optionally, the number of particles is the same as the number of original sample data.
[0200] The self-organizing mapping weight particle swarm mean clustering device provided by the embodiments of the present invention can execute the self-organizing mapping weight particle swarm mean clustering method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0201] Embodiment Six
[0202] Figure 9 A schematic structural diagram of a device provided in Embodiment 6 of the present invention is shown as Figure 9 shown. The device includes a processor 70, a memory 71, an input device 72, and an output device 73. The number of processors 70 in the device can be one or more. Figure 9 Here, one processor 70 is taken as an example. The processor 70, the memory 71, the input device 72, and the output device 73 in the device can be connected through a bus or other means. Figure 9 Here, connection through a bus is taken as an example.
[0203] The memory 71, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the self-organizing mapping weight particle swarm mean clustering method in the embodiments of the present invention (for example, the data acquisition module 10, the rough clustering module 20, and the fine clustering module 30 in the self-organizing mapping weight particle swarm mean clustering device). The processor 70 executes various functional applications and data processing of the device / terminal / server by running the software programs, instructions, and modules stored in the memory 71, that is, implements the above-mentioned self-organizing mapping weight particle swarm mean clustering method.
[0204] The memory 71 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the terminal, etc. In addition, the memory 71 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 71 can further include a memory remotely set relative to the processor 70, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0205] The input device 72 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 73 can include a display device such as a display screen.
[0206] Embodiment 7
[0207] Embodiment 7 of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a self-organizing mapping weight particle swarm mean clustering method when executed by a computer processor. The method includes:
[0208] Obtain original sample data;
[0209] Use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; where K is a natural number;
[0210] Perform fine clustering on the original sample data based on the rough clustering clusters and rough clustering centers to obtain the target clustering centers.
[0211] Of course, for a storage medium containing computer-executable instructions provided in an embodiment of the present invention, the computer-executable instructions are not limited to the method operations described above, and can also execute related operations in the weight particle swarm mean clustering method of self-organizing mapping provided in any embodiment of the present invention.
[0212] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0213] It should be noted that in the embodiments of the above search device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0214] Note that the above is only a preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A weight particle swarm mean clustering method for self-organizing mapping, characterized in that, it includes: Obtain the original sample data, where the original sample data is software portrait data; Use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers; where K is a natural number; Perform fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers, where the number of clustering clusters in fine clustering is the same as the number of rough clustering clusters, and the clustering centers in fine clustering are the rough clustering centers; Among them, the step of using the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters includes: Input the original sample data into a rough clustering model based on the SOM clustering algorithm for iterative processing; When the number of iterations reaches the first preset number of iterations, output the weight values of each neuron of the rough clustering model when reaching the first preset number of iterations; Calculate the inner product of each original sample data and each neuron weight value in turn to form an inner product matrix, and use the element position of the maximum inner product of each original sample data in the inner product matrix as the winning index of each original sample data; Classify the original sample data corresponding to the winning indexes with the same inner product value to obtain the K rough clustering clusters belonging to K classes; The step of performing fine clustering on the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers includes: Use the original sample data to initialize the positions and velocities of multiple particles; Calculate the fitness value of each particle and the fitness value of the particle population according to the position and velocity of each initialized particle; Determine the local optimal position vector of each particle according to the fitness value of each particle, and determine the global optimal position vector of the particle population according to the fitness value of the particle population; Initialize the rough clustering clusters and the rough clustering centers as the current clustering clusters and the current clustering centers; Calculate the distances of each particle to the K current clustering clusters in turn according to the Euclidean distance formula, and divide each particle into the nearest current clustering cluster according to the nearest neighbor principle; Update the current clustering center according to the clustering center calculation formula; Update the position and velocity of each particle according to the local optimal position vector of the particle and the global optimal position vector of the particle population, and calculate the fitness value of the particle and the fitness value of the particle population after updating the position and velocity; Return to perform the operation of dividing the particles into the current clustering clusters until the iteration stops when reaching the second preset number of iterations, and use the current clustering center when reaching the second preset number of iterations as the target clustering center, where the second preset number of iterations is the number of iterations when the fitness value of the particle and the fitness value of the population where the particle is located are the smallest.
2. The method according to claim 1, characterized in that, the step of using the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters and rough clustering centers includes: Use the SOM clustering algorithm to perform rough clustering on the original sample data to obtain K rough clustering clusters; According to the K rough clustering clusters, calculate the sample mean of the original sample data in each rough clustering cluster; According to the sample mean of each rough clustering cluster, use the nearest neighbor principle to find the sample closest to its sample mean in each rough clustering cluster as the rough clustering center of the corresponding rough clustering cluster.
3. The method according to claim 2, characterized in that, the step of using the nearest neighbor principle to find the original sample data closest to its sample mean in each rough clustering cluster as the rough clustering center of the corresponding rough clustering cluster according to the sample mean of each rough clustering cluster includes: Calculate the Euclidean distance between the sample mean of each rough clustering cluster and each original sample data in its corresponding rough clustering cluster, and use the original sample data with the smallest Euclidean distance value as the rough clustering center of the corresponding rough clustering cluster.
4. The method according to claim 1, characterized in that, the number of the particles is the same as the number of the original sample data.
5. A weight particle swarm mean clustering device for self-organizing mapping, characterized in that, comprising: a data acquisition module for acquiring original sample data; a rough clustering module for roughly clustering the original sample data using a SOM clustering algorithm to obtain K rough clustering clusters and rough clustering centers; where K is a natural number; a fine clustering module for finely clustering the original sample data based on the rough clustering clusters and the rough clustering centers to obtain target clustering centers, where the number of clustering clusters for fine clustering is the same as the number of rough clustering clusters, and the clustering centers for fine clustering are the rough clustering centers; wherein, the rough clustering module includes: a loop iteration processing unit for inputting the original sample data into a rough clustering model based on a SOM clustering algorithm for loop iteration processing; a neuron weight output unit for outputting the weights of each neuron of the rough clustering model when the iteration number reaches a first preset iteration number; an inner product calculation unit for sequentially calculating the inner product of each original sample data and each neuron weight to form an inner product matrix; a winning index acquisition unit for using the element position of the maximum inner product of each original sample data in the inner product matrix as the winning index of each original sample data; a rough clustering cluster obtaining unit for classifying the original sample data corresponding to the winning indexes with the same inner product value to obtain K rough clustering clusters belonging to K classes; the fine clustering module includes: a particle position and velocity initialization unit for initializing the positions and velocities of a plurality of particles using the original sample data; calculating the fitness value of each particle and the fitness value of the particle population according to the positions and velocities of each initialized particle; determining the local optimal position vector of each particle according to the fitness value of each particle, and determining the global optimal position vector of the particle population according to the fitness value of the particle population; a current clustering cluster and current clustering center initialization unit for initializing the rough clustering clusters and the rough clustering centers as the current clustering clusters and the current clustering centers; A particle partitioning unit, which is used to calculate the distances from each particle to K current clustering clusters in sequence according to the Euclidean distance formula, and partition each particle into the nearest current clustering cluster according to the nearest neighbor principle; A clustering center calculation and update unit, which is used to update the current clustering center according to the clustering center calculation formula; A particle position and velocity update unit, which is used to update the position and velocity of each particle according to the local optimal position vector of the particle and the global optimal position vector of the particle population, and calculate the fitness value of the particle after updating the position and velocity and the fitness value of the particle population; A return execution particle partitioning and iteration unit, which is used to return and execute the operation of partitioning particles into the current clustering cluster until the iteration stops when the second preset iteration number is reached; A target clustering center determination unit, which is used to use the current clustering center when the second preset iteration number is reached as the target clustering center; Wherein, the second preset iteration number is the iteration number when the fitness value of the particle and the fitness value of the population where the particle is located are the smallest.
6. An electronic device, characterized in that, the electronic device includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the self-organizing mapping weighted particle swarm mean clustering method according to any one of claims 1-4.
7. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the self-organizing mapping weighted particle swarm mean clustering method according to any one of claims 1-4.
Citation Information
Patent Citations
Hoisting machinery energy efficiency evaluation method based on SOM-entropy algorithm
CN113627731A