Self-adaptive neighborhood grain clustering method suitable for mixed type attribute data
By combining the K-mean clustering algorithm and particle calculation, an adaptive neighborhood particle clustering method is proposed, which solves the problem that traditional algorithms are difficult to deal with mixed attribute data, and realizes the stability and efficiency of clustering results.
Patent Information
- Application Number
- CN202510039810.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The traditional K-mean clustering algorithm can only process numerical attribute data, and the clustering results are unstable, making it difficult to effectively process mixed attribute data.
An adaptive neighborhood particle clustering method is proposed, combining the K-mean clustering algorithm and particle calculation, and by calculating the neighborhood particle vector of mixed attribute data, selecting the initial clustering center, and using the neighborhood particle K-mean clustering algorithm to update the clustering center.
Effective processing of mixed attribute data improves the stability and applicability of clustering results, and avoids the problem of inaccurate measurement of data differentiation caused by different dimensions of numerical attributes and symbol attributes.
Smart Images

Figure CN119961708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and more specifically to an adaptive neighborhood granular clustering method suitable for mixed attribute data. Background Art
[0002] Clustering is an unsupervised analysis method that can be used to mine the internal structure or pattern of data. As one of the research hotspots in various disciplines, clustering is widely used in marketing analysis, natural language processing, image processing, bioinformatics, computer vision and other fields. In practical applications, mixed attribute data with both numerical and symbolic attributes are common. The traditional K-means clustering algorithm can only process data with numerical attributes, and the initial clustering center of the K-means clustering algorithm is randomly selected, which will lead to unstable clustering results. Summary of the invention
[0003] 1. Technical issues to be solved
[0004] In view of the problem that the traditional K-means clustering algorithm in the prior art can only process numerical attribute data and the clustering results are unstable, the present invention provides an adaptive neighborhood granular clustering method suitable for mixed attribute data, which combines the K-means clustering algorithm with granular computing to realize granulation of mixed attribute data, and has the excellent characteristics of strong applicability and high clustering performance.
[0005] 2. Technical solution
[0006] The purpose of the present invention is achieved through the following technical solutions.
[0007] An adaptive neighborhood granular clustering method suitable for mixed attribute data includes the following steps:
[0008] S1. Obtain and input the data to be clustered;
[0009] S2, calculating the neighborhood particle vector of the data to be clustered by using an adaptive neighborhood particle clustering method for mixed attribute data;
[0010] S3, selecting the initial clustering center of the data to be clustered based on the data dissimilarity measurement method;
[0011] S4, using the neighborhood particle K-means clustering algorithm to update the cluster center;
[0012] S5. Output clustering results.
[0013] As a further improvement of the present invention, in step S1, the data to be clustered includes a sample set and an attribute set, and the attribute set includes numerical attributes and symbolic attributes.
[0014] As a further improvement of the present invention, in step S2, the neighborhood particle vector of the data to be clustered is calculated, and the steps include:
[0015] Calculate the average value of attribute differences of numerical attributes, and identify the neighborhood of numerical attributes based on the average value of attribute differences;
[0016] Identify the neighborhood of symbolic attributes;
[0017] The neighborhood particle vectors of different mixed attribute data are calculated through the neighborhood of numerical attributes and the neighborhood of symbolic attributes.
[0018] As a further improvement of the present invention, the neighborhood of the numerical attribute is determined based on the average value of the attribute difference, and the determination formula is:
[0019]
[0020] Among them, i and j are natural numbers, N is the number of samples, n is the numerical attribute feature, r is the hyperparameter, a is the attribute, x is i represents the i-th sample, x ia represents the value of the i-th sample on attribute a, x ja represents the value of the jth sample on attribute a, represents the neighborhood particle vector of the i-th sample on attribute a, represents the neighborhood discriminant value of the jth sample, represents the neighborhood discriminant of a numerical attribute, It represents the Euclidean distance between any samples on attribute a, and dis(a) represents the average attribute difference.
[0021] As a further improvement of the present invention, the neighborhood of the discriminant symbolic attribute is discriminated by the following formula:
[0022]
[0023] Among them, c represents the symbolic attribute feature, b represents the attribute, and x ib represents the value of the i-th sample on attribute b, x jb represents the value of the jth sample on attribute b, Represents x i The neighborhood particle vector on attribute b, represents the neighborhood discrimination value of the jth sample, Represents the neighborhood discriminant for a symbolic attribute.
[0024] As a further improvement of the present invention, the neighborhood particle vectors of the different mixed attribute data are expressed as:
[0025] G={G A (x1),GA (x2),...,G A (x N )}
[0026]
[0027] Among them, G represents the neighborhood particle vector of the entire data to be clustered, G A (x N ) represents x N Neighborhood particle vector on attribute A, T represents matrix transpose.
[0028] As a further improvement of the present invention, in step S3, the data-based dissimilarity measurement method is used to select the initial cluster center of the data to be clustered, and the steps include:
[0029] S31: Based on the neighborhood particle vectors, calculate the dissimilarity matrix and overall dissimilarity of the data to be clustered;
[0030] S32: Calculate the number of neighborhood particle vectors whose median value in each row of data in the dissimilarity matrix is less than and equal to m times the overall dissimilarity, and construct a matrix;
[0031] S33: Select the first initial cluster center point, that is, the neighborhood particle vector corresponding to the maximum value in the matrix, and update the matrix;
[0032] S34: Determine whether the difference between the neighborhood particle vector corresponding to the maximum value of the matrix and the initial cluster center meets the condition. If the condition is met, select the neighborhood particle vector corresponding to the maximum value as the cluster center and update the matrix; if the condition is not met, update the matrix and repeat this step; the condition is expressed as: D(G A (x j ),u p )≥Tdis,(p=1,2,...,s),G A (x j ) represents x j The neighborhood particle vector on attribute A, p represents a natural number, u p represents the cluster center vector, s represents the number of cluster centers, and Tdis represents the overall dissimilarity.
[0033] S35: Set the number of clusters, and determine whether the number of cluster centers meets the number of clusters. If so, obtain the initial cluster center point.
[0034] As a further improvement of the present invention, the overall dissimilarity of the data to be clustered is expressed as:
[0035]
[0036] Among them, Tdis represents the overall dissimilarity, N represents the total number of neighborhood particle vectors, x and y represent any two samples in the data set to be clustered, and G A (x) represents the neighborhood particle vector of x on attribute A, G A (y) represents the neighborhood particle vector of y on attribute A, and ρ represents the hyperparameter.
[0037] As a further improvement of the present invention, in step S4, the step of running the neighborhood particle K-means clustering algorithm to update the cluster center includes:
[0038] S41, calculating the distance between all neighborhood particle vectors and the selected cluster center point, classifying the i-th sample into the category corresponding to the minimum distance, and updating the category;
[0039] S42, calculating a new cluster center point of the category;
[0040] S43: Determine whether the difference between the objective functions in two iterations is less than a set threshold or whether the maximum number of iterations is reached. If the conditions are met, output the clustering result. If the conditions are not met, repeat steps S41-S43.
[0041] As a further improvement of the present invention, the objective function is expressed as:
[0042]
[0043] Among them, J e represents the objective function, j represents a natural number, k represents the number of clusters, C j Indicates the category, u j Represents the mean grain vector of the class.
[0044] 3. Beneficial effects
[0045] Compared with the prior art, the advantages of the present invention are:
[0046] The invention discloses an adaptive neighborhood granular clustering method suitable for mixed attribute data, which combines a K-means clustering algorithm with granular computing, takes into account different densities of different numerical attribute data, and realizes granulation of mixed attribute data through an adaptive neighborhood granulation method of mixed attribute data, which effectively avoids the problem of inaccurate data dissimilarity measurement caused by different dimensions of numerical attribute data and symbolic attribute data. Furthermore, based on the neighborhood granular vector of data granulation, a newly defined data dissimilarity measurement method is used to select initial clustering centers. Finally, the clustering result of the data is obtained by combining the neighborhood granular K-means clustering algorithm, which effectively improves the stability of the algorithm and has the excellent characteristics of strong applicability and high clustering performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1The figure is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0049] Example
[0050] like Figure 1 As shown, an adaptive neighborhood particle clustering method suitable for mixed attribute data provided by this embodiment includes the following steps: S1, obtaining and inputting the data to be clustered; S2, calculating the neighborhood particle vectors of the data to be clustered by the adaptive neighborhood particle clustering method for mixed attribute data; S3, selecting the initial clustering center of the data to be clustered based on the data dissimilarity measurement method; S4, updating the clustering center by using the neighborhood particle K-means clustering algorithm; S5, outputting the clustering result.
[0051] Specifically in this embodiment, step S1, obtain and input the data to be clustered. In this embodiment, the data set to be clustered N = (X, A), where X represents a sample set, X = {x1, x2, ..., x N}, A represents the attribute set, A={a1,a2,...,a p , a p+1 , ..., a m In this embodiment, the attribute set includes numerical attributes and symbolic attributes. For attribute set A, {a1, a2, ..., a p} is a numeric attribute, {a p+1 , ..., a m} is a symbolic attribute.
[0052] It should be noted that in this embodiment, the algorithm parameters need to be initialized. That is, the number of clusters k, hyperparameters r and ρ, the number of iterations t = 0, the maximum number of iterations Iters = 50, and the loss function J are set in advance. e The change threshold VCJT=0.001.
[0053] Step S2, calculates the neighborhood particle vectors of the data to be clustered by an adaptive neighborhood particle clustering method for mixed attribute data, and the steps include: calculating the attribute difference average value of the numerical attribute, and judging the neighborhood of the numerical attribute based on the attribute difference average value; judging the neighborhood of the symbolic attribute; and calculating the neighborhood particle vectors of different mixed attribute data through the neighborhood of the numerical attribute and the neighborhood of the symbolic attribute.
[0054] Specifically, in this embodiment, for a numerical attribute a∈{a1, a2, ..., a p}, calculate the average attribute difference dis(a) of the numeric attribute on attribute a, and the calculation formula is:
[0055]
[0056] Where a represents the attribute, dis(a) represents the average attribute difference, N represents the number of samples, i and j both represent natural numbers, and x ia represents the value of the i-th sample on attribute a, x ja Represents the value of the jth sample on attribute a.
[0057] Furthermore, the neighborhood of the numerical attribute is determined based on the average value of the attribute difference. The determination formula is:
[0058]
[0059] Among them, i and j are natural numbers, N is the number of samples, n is the numerical attribute feature, r is the hyperparameter, a is the attribute, a∈{a1,a2,...,a m}, x i represents the i-th sample, x ia represents the value of the i-th sample on attribute a, x ja represents the value of the jth sample on attribute a, Represents x i The neighborhood particle vector on attribute a, represents the neighborhood discrimination value of the jth sample, Represents the neighborhood discriminant of the numerical attribute, used to discriminate x ia and x ja Whether they are adjacent in attribute a, dis(a) represents the average value of attribute difference. In this embodiment, if Then it means x ia and x ja Adjacent, otherwise not adjacent.
[0060] To determine the neighborhood of symbolic attributes, the determination formula is:
[0061]
[0062] Among them, c represents the symbolic attribute feature, b represents the attribute, b∈{a1,a2,...,a m}, x ib represents the value of the i-th sample on attribute b, x jb represents the value of the jth sample on attribute b, Represents x i The neighborhood particle vector on attribute b, represents the neighborhood discrimination value of the jth sample, Represents the neighborhood discriminant of the symbolic attribute, used to discriminate x ib and x jbWhether they are adjacent in attribute b. In this embodiment, if Then it means x ib and x jb Adjacent, otherwise not adjacent.
[0063] Furthermore, the neighborhood particle vectors of different mixed attribute data are calculated through the neighborhood of the numerical attribute and the neighborhood of the symbolic attribute. In this embodiment, the neighborhood particle vectors of different mixed attribute data are expressed as:
[0064] G={G A (x1),G A (x2),...,G A (x N )}
[0065]
[0066] Among them, G represents the neighborhood particle vector of the entire data to be clustered, G A (x N ) represents x N Neighborhood particle vector on attribute A, T represents matrix transpose.
[0067] Step S3, selecting the initial cluster center of the data to be clustered based on the data dissimilarity measurement method, the steps include:
[0068] S31: Based on the neighborhood particle vectors, the dissimilarity matrix and overall dissimilarity of the data to be clustered are calculated. In this embodiment, the calculation formula of the dissimilarity matrix of the data to be clustered is:
[0069]
[0070] Where disM represents the dissimilarity matrix, N represents a natural number, and D(G A (x i ), G A (x j )) represents the neighborhood particle vector x i and x j The difference.
[0071] Furthermore, the overall dissimilarity of the data to be clustered is calculated using the following formula:
[0072]
[0073] Where Tdis represents the overall dissimilarity. At this time, the number of cluster centers has been selected for initialization, and the number of cluster centers s=0.
[0074] S32: Calculate the number of neighborhood particle vectors whose median value in each row of data in the dissimilarity matrix is less than or equal to m times the overall dissimilarity, and construct a matrix. Specifically, select m=0.25, calculate the number of neighborhood particle vectors whose median value in each row of data in the dissimilarity matrix disM is ≤0.25*Tdis, and construct a matrix Counts, Counts=[counts(x1), counts(x2), ..., counts(x N )],counts(x i ) means that the data in the i-th row of the dissimilarity matrix disM satisfies the condition: D(G A (x i ), G A (x j ))≤0.25*Tdis, the number of data (j=1,…,N).
[0075] S33: Select the first initial cluster center point u i , that is, the maximum value counts(x i ) corresponds to the neighborhood particle vector G A (x i ), update the matrix Counts, counts(x i )=0, let the number of cluster centers s=s+1.
[0076] S34: Determine the maximum value of the matrix Counts counts(x j ) corresponds to the neighborhood particle vector G A (x j ) and the initial cluster center to meet the conditions. If so, select the maximum value counts(x j ) corresponds to the neighborhood particle vector G A (x j ) as the s+1th cluster center, update the matrix Counts, counts(x j )=0, the number of cluster centers s=s+1; if the condition is not met, update the matrix Counts and repeat the step; the condition is expressed as: D(G A (x j ),u p )≥Tdis, (p=1, 2,...,s), G A (x j ) represents x j The neighborhood particle vector on attribute A, p represents a natural number, u p represents the cluster center vector, s represents the number of cluster centers, and Tdis represents the overall dissimilarity.
[0077] S35: Set the number of clusters k, and determine whether the cluster center s satisfies the number of clusters k. If so, the initial cluster center point {u1, u2, ..., u s}.
[0078] It is worth noting that, in this embodiment, for the overall dissimilarity Tdis of the data to be clustered, there is:
[0079]
[0080] in,
[0081]
[0082] G A (x)∨G A (y)=(g i (x)∨g1(y), g2(x)∨g2(y),...,g m (x)∨g m (y))
[0083] G A (x)∧G A (y)=(g1(x)∧g1(y),g2(x)∧g2(y),...,g m (x)∧g m (y))
[0084] G A (x)-G A (y)=(g1(x)-g1(y), g2(x)-g2(y),...,g m (x)-g m (y))
[0085]
[0086] Step S4, using the neighborhood particle K-means clustering algorithm to update the cluster center, the steps include:
[0087] S41. Calculate all neighborhood particle vectors G A (x i ) and the selected cluster center u j The distance d of (j=1, 2, ..., s) ij , d ij =D(G A (x i ),u j ), and x i Stroke to minimum distance d ij In the corresponding category, update category C j , C j =C j∪x i .
[0088] S42, Calculation Category C j The new cluster centers, Among them, |C j | indicates category C j The number of data points in G A (x) represents the neighborhood particle vector of the data point x to be clustered.
[0089] S43, determine the objective function J during the two iterations e Whether the difference is less than the set threshold VCJT or whether the maximum number of iterations Iters is reached, if the conditions are met, the clustering result is output, if not, steps S41-S43 are repeated.
[0090] In this embodiment, the objective function is expressed as:
[0091]
[0092] Among them, J e represents the objective function, j represents a natural number, k represents the number of clusters, C j Indicates the category, u j Indicates category C j The mean particle vector, also the particle mass center, D(G A (x),u j ) represents the difference between the particle vector and the particle centroid of the data point x to be clustered.
[0093] Step S5, outputting the clustering result, that is, forming the final clustering result according to the clustering result obtained by training in step S4, C = (C1, C2, ..., C k ).
[0094] Therefore, the present embodiment provides an adaptive neighborhood particle clustering method suitable for mixed attribute data, which takes into account the different densities of different numerical attribute data, and realizes the granulation of mixed attribute data by proposing an adaptive neighborhood particle method for mixed attribute data, thereby effectively avoiding the problem of inaccurate data dissimilarity measurement caused by different dimensions of numerical attribute data and symbolic attribute data; based on the neighborhood particle vector of data granulation, the newly defined data dissimilarity measurement method is used to select the initial clustering center, and the clustering result of the data is obtained by combining the neighborhood particle K-means clustering algorithm, which effectively improves the stability of the algorithm and has the excellent characteristics of strong applicability and high clustering performance.
[0095] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited thereto, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all belong to the protection scope of the present invention. In addition, the word "including" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate the name, and do not indicate any specific order.
Claims
1. An adaptive neighborhood particle clustering method suitable for mixed attribute data, comprising the following steps: S1. Obtain and input the data to be clustered; S2, calculating the neighborhood particle vector of the data to be clustered by using an adaptive neighborhood particle clustering method for mixed attribute data; S3, selecting the initial clustering center of the data to be clustered based on the data dissimilarity measurement method; S4, using the neighborhood particle K-means clustering algorithm to update the cluster center; S5. Output clustering results.
2. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 1, characterized in that: In step S1, the data to be clustered includes a sample set and an attribute set, and the attribute set includes numerical attributes and symbolic attributes.
3. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 2, characterized in that: In step S2, the neighborhood particle vector of the data to be clustered is calculated, and the steps include: Calculate the average value of attribute differences of numerical attributes, and identify the neighborhood of numerical attributes based on the average value of attribute differences; Identify the neighborhood of symbolic attributes; The neighborhood particle vectors of different mixed attribute data are calculated through the neighborhood of numerical attributes and the neighborhood of symbolic attributes.
4. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 3, characterized in that: The neighborhood of the numerical attribute is determined based on the average value of the attribute difference, and the determination formula is: Among them, i and j are natural numbers, N is the number of samples, n is the numerical attribute feature, r is the hyperparameter, a is the attribute, x is i represents the i-th sample, x ia represents the value of the i-th sample on attribute a, x ja represents the value of the jth sample on attribute a, represents the neighborhood particle vector of the i-th sample on attribute a, represents the neighborhood discriminant value of the jth sample, represents the neighborhood discriminant of a numerical attribute, It represents the Euclidean distance between any samples on attribute a, and dis(a) represents the average attribute difference.
5. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 4, characterized in that: The neighborhood of the discriminant symbolic attribute is determined by the following formula: Among them, c represents the symbolic attribute feature, b represents the attribute, and x ib represents the value of the i-th sample on attribute b, x jb represents the value of the jth sample on attribute b, Represents x i The neighborhood particle vector on attribute b, represents the neighborhood discrimination value of the jth sample, Represents the neighborhood discriminant for a symbolic attribute.
6. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 5, characterized in that: The neighborhood particle vectors of different mixed attribute data are expressed as: G={G A (x1),G A (x2),...,G A (x N )} Among them, G represents the neighborhood particle vector of the entire data to be clustered, G A (x N ) represents x N Neighborhood particle vector on attribute A, T represents matrix transpose.
7. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 6, characterized in that: In step S3, the data-based dissimilarity measurement method is used to select the initial cluster center of the data to be clustered, and the steps include: S31: Based on the neighborhood particle vectors, calculate the dissimilarity matrix and overall dissimilarity of the data to be clustered; S32: Calculate the number of neighborhood particle vectors whose median value in each row of data in the dissimilarity matrix is less than and equal to m times the overall dissimilarity, and construct a matrix; S33: Select the first initial cluster center point, that is, the neighborhood particle vector corresponding to the maximum value in the matrix, and update the matrix; S34: Determine whether the difference between the neighborhood particle vector corresponding to the maximum value of the matrix and the initial cluster center meets the condition. If the condition is met, select the neighborhood particle vector corresponding to the maximum value as the cluster center and update the matrix; if the condition is not met, update the matrix and repeat this step; the condition is expressed as: D(G A (x j ),u p )≥Tdis,(p=1,2,...,s),G A (x j ) represents x j The neighborhood particle vector on attribute A, p represents a natural number, u p represents the cluster center vector, s represents the number of cluster centers, and Tdis represents the overall dissimilarity. S35: Set the number of clusters, and determine whether the number of cluster centers meets the number of clusters. If so, obtain the initial cluster center point.
8. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 7, characterized in that: The overall dissimilarity of the data to be clustered is expressed as: Among them, Tdis represents the overall dissimilarity, N represents the total number of neighborhood particle vectors, x and y represent any two samples in the data set to be clustered, and G A (x) represents the neighborhood particle vector of x on attribute A, G A (y) represents the neighborhood particle vector of y on attribute A, and ρ represents the hyperparameter.
9. The adaptive neighborhood granular clustering method for mixed attribute data according to claim 1, characterized in that: In step S4, the operation of the neighborhood particle K-means clustering algorithm to update the cluster center includes: S41, calculating the distance between all neighborhood particle vectors and the selected cluster center point, classifying the i-th sample into the category corresponding to the minimum distance, and updating the category; S42, calculating a new cluster center point of the category; S43: Determine whether the difference between the objective functions in two iterations is less than a set threshold or whether the maximum number of iterations is reached. If the conditions are met, output the clustering result. If the conditions are not met, repeat steps S41-S43.
10. The adaptive neighborhood granular clustering method suitable for mixed attribute data according to claim 9, characterized in that: The objective function is expressed as: Among them, J e represents the objective function, j represents a natural number, k represents the number of clusters, C j Indicates the category, u j Represents the mean grain vector of the class.
Citation Information
Patent Citations
Automatic topology identification method for photovoltaic grid-connected low-voltage transformer area power distribution network
CN114818890A
Fast K-nearest neighbor classifier method for large-scale data
CN116363420A