Population census clustering method oriented to fairness and balance constraint optimization
By combining the cuckoo search algorithm and weighted Euclidean distance, and combining fairness and equilibrium constraints, the problem of insufficient fairness and balance in the census is solved, and a more accurate, fair and stable clustering result is achieved.
Patent Information
- Application Number
- CN202510253348.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Traditional K-mean clustering algorithms are difficult to ensure the fairness and balance of clustering results in areas such as census, which leads to data bias and inaccuracy of analysis results.
Combining the cuckoo search algorithm, the K-mean clustering algorithm with weighted Euclidean distance and fairness and equilibrium constraints, the fairness and equilibrium constraints in the initial cluster center selection, K-value selection and clustering process are optimized.
It significantly improves the clustering accuracy, fairness and stability of census data analysis, ensuring the fairness of clustering results and the balance of sample cluster distribution.
Smart Images

Figure CN120180164A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data mining and machine learning, and particularly to a census clustering method for fairness and balance constraint optimization. Background Art
[0002] The K-means Clustering Algorithm (K-means) is a very common clustering algorithm in data analysis and is widely used because of its high computational efficiency and simple implementation. However, there is always a problem in the operation of the traditional K-means clustering algorithm that it is sensitive to the selection of the initial clustering center and the value of K. In practical applications, the traditional K-means clustering algorithm often fails to fully consider the influence of sensitive attributes (such as gender, age, etc.) on the clustering results. Especially in fields such as census and sociological research, the balanced distribution of sensitive attributes is crucial. If the clustering algorithm does not consider the balance of these sensitive attributes, it may lead to some groups being over-clustered or ignored, resulting in data bias and affecting the fairness of the analysis results. For example, when conducting a census, failure to ensure the balanced distribution of different genders, age groups or groups will affect the fairness and accuracy of decision-making. In addition, sample imbalance is also a major challenge faced by the traditional K-means clustering algorithm. The K-means clustering algorithm does not have a built-in mechanism to solve the sample imbalance problem between clustering clusters, which often leads to too many samples in some clusters and too few samples in other clusters, which not only affects the clustering effect, but also may affect subsequent data analysis and decision support.
[0003] To solve the above problems, in recent years, many researchers have improved the traditional K-means clustering algorithm. For example, some researchers have proposed an algorithm that combines the particle swarm algorithm and the K-means clustering algorithm, and solved the problem of initializing the clustering center of the K-means clustering algorithm through the particle swarm algorithm; other researchers have used the FairClustering with Sensitive Attributes (FCSN) algorithm to process the sensitive attributes of samples, and introduced fairness constraints in the optimization process to ensure that the clustering results will not be biased due to sensitive attributes. In addition, for the problems of sensitive attributes and sample imbalance in data, some research has also proposed the Balanced Fair K-Means (BFKM) algorithm to ensure that each cluster in the clustering results has a relatively balanced number of samples and eliminate the bias in the data.
[0004] Although these improvement schemes have enhanced the clustering effect of census data to a certain extent, there are still many deficiencies. Especially in the fields of complex datasets and sensitive data analysis, how to better optimize the initial clustering centers, reasonably select the number of clusters K, and effectively handle fairness and balance constraints remains a hot topic and a difficult point in current research. Especially in data analysis involving sensitive group characteristics such as the census, the fairness, accuracy, and stability of clustering results are crucial for decision-making support. Therefore, innovative algorithms are urgently needed to overcome the limitations of existing methods, improve the fairness guarantee and result reliability in the census data clustering process, so as to provide a more accurate and fair basis for policy-making and socio-economic analysis. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a census clustering method optimized for fairness and balance constraints in view of the deficiencies of the prior art. Specifically, it is a clustering optimization method of the K-means clustering algorithm combining the cuckoo search algorithm, weighted Euclidean distance, and fairness and balance constraints, aiming to improve the clustering accuracy, fairness, and stability of census data analysis.
[0006] The method of the present invention includes the following steps:
[0007] Step 1, obtain the census sampling dataset X;
[0008] Step 2, use the cuckoo search algorithm for optimization and select the initial population clustering centers;
[0009] Step 3, perform the silhouette coefficient comparison method within the K value selection range to select the best K value for the population survey clustering;
[0010] Step 4, perform weighted Euclidean distance optimization iteration;
[0011] Step 5, introduce fairness and balance constraints;
[0012] Step 6, iterate until the convergence condition is reached.
[0013] Step 2 includes:
[0014] Step 2-1, initialize num_nest nests, and randomly generate K candidate initial population clustering centers in each nest. The formula for num_nest is:
[0015]
[0016] where n is the number of samples in the census sampling dataset;
[0017] Calculate the sum of squared errors SSE (Sum of Squared Errors, SSE) of each nest as the fitness:
[0018]
[0019] where x j represents the census sample point in the i-th cluster, and μ i is the population clustering center of the i-th cluster, and C i is the i-th cluster;
[0020] Step 2-2: First, randomly select a nest a, and then randomly select two different nests C i and C j . Linearly combine the two different nests C i and C j to form a new nest C new1 . Compare the fitness of the new nest C new1 with that of the nest a. If the fitness of the new nest C new1 is less than the fitness of the nest a, then replace a with the new nest C new1 ; otherwise, do not replace it;
[0021] The formula for the linear combination is:
[0022] C new1 =αC i +(1 - α)C j (3)
[0023] where α is a random step factor with a value ranging from 0 to 1;
[0024] Step 2-3: Perform Levy flight on each nest. Each nest has a probability of 0.25 to perform Levy random search to increase diversity. In the solution of the present invention, a probability value of 0.25 is selected to perform Levy random search to increase the diversity during the search process. The selection of this probability value has been verified through experiments, which can balance the exploration and exploitation efficiency while ensuring the global search ability, and avoid excessive computational overhead caused by frequent Levy flights. By controlling the occurrence probability of Levy flight, this method effectively improves the convergence speed and computational efficiency of the algorithm, and is particularly suitable for high-dimensional optimization problems. Of course, the probability value of 0.25 in this technical solution is not the only choice, and other probability values can also be used to achieve similar effects. The specific choice should be optimized according to the actual application scenario and experimental results:
[0025] C new2 =C old +L(β) (4)
[0026] where β is the Levy distribution parameter, C old is the initial nest of Levy flight, and C new2 is C oldThe newly synthesized nest after Levy flight, L(β) is the step size of Levy flight;
[0027] The step size calculation formula of Levy distribution is:
[0028]
[0029] where u and v are random variables generated from a normal distribution, u ∼ N(0, δ 2 ), v ∼ N(0, 1), δ is a parameter of Levy distribution, usually used to adjust the distribution characteristics of the step size; Γ is the gamma function, used to represent the factorial in mathematics;
[0030] To prevent the newly generated solution from exceeding the data range, the following boundary handling strategy is used:
[0031] C new2 = max(min(C new2 , max(X)), min(X)) (7);
[0032] where max(X) is the upper limit of the census sampling data set X, and min(X) is the lower limit of the census sampling data set X;
[0033] Step 2-4, repeat Steps 2-1 to 2-3 for iteration until the maximum number of iterations is reached, and output the best initial clustering center. After multiple experimental comparisons, the maximum number of iterations in the present invention is set to 150 times.
[0034] Step 3 includes:
[0035] Step 3-1, calculate the within-class distance a(i) of the i-th census sample:
[0036]
[0037] where l i is the cluster number to which the i-th census sample belongs, is the cluster to which the i-th census sample belongs, is the number of samples in the cluster to which the i-th census sample belongs, X i and X j respectively represent the i-th census sample and the j-th census sample; X i and X j belong to the same cluster;
[0038] Step 3-2, calculate the between-class distance b(i) of the nearest cluster of census data analysis:
[0039]
[0040] where Cc Denote the cluster with the number equal to c;
[0041] Step 3-3: Calculate the silhouette coefficient of each census sample;
[0042] Step 3-4: For the current K value, calculate the average silhouette coefficient S k ;
[0043] Step 3-5: Select the K value with the largest average silhouette coefficient as the number of clusters, and proceed to Step 4.
[0044] Step 3-3 includes: The calculation formula for the silhouette coefficient s(i) of the i-th census sample is:
[0045]
[0046] Step 3-4 includes: Calculate the average silhouette coefficient S using the following formula k :
[0047]
[0048] Step 4 includes:
[0049] Step 4-1: Calculate the standard deviation σ of the m-th feature m :
[0050]
[0051] where X im is the m-th feature value of the i-th census sample in the census sampling dataset X, n is the total number of census samples, and μ m represents the mean of all samples of the m-th feature;
[0052] Step 4-2: When calculating the weight w, determine the weight by the standard deviation of each feature:
[0053] w m = 1 / σ m (13)
[0054] where w m represents the weight of the m-th feature;
[0055] Step 4-3: Calculate the weighted Euclidean distance.
[0056] In Step 4-3, calculate the weighted Euclidean distance using the following formula:
[0057]
[0058] where D ij is the weighted Euclidean distance from the i-th census sample to the j-th census sample, Cjm is the m-th feature of the j-th cluster center, and M is the total number of features of the census samples.
[0059] Step 5 includes:
[0060] By using the indicator matrix, the fairness of the population-sensitive attribute is expressed as:
[0061]
[0062] where F represents the indicator matrix of the sensitive attributes of the population sampling data, and Y represents the indicator matrix of the cluster labels. In the present invention, the indicator matrix is used to represent the existence of certain features, attributes or categories. The elements in the matrix only take 0 or 1, where 1 means that a certain specific condition holds or a certain feature exists, and 0 means that the condition does not hold or the feature does not exist. In the prior art, the indicator matrix can be generated by automated data processing techniques;
[0063] When considering the balance constraint of population data analysis, in order to ensure that the size of each cluster is about the same as much as possible, this method uses the arithmetic mean inequality. When E h=n / k , E1 + E2 + E3 +... + E k = n, we get:
[0064]
[0065] where E h represents the size of the h-th cluster, and k represents the total number of clusters;
[0066] Using the indicator matrix, formula (16) is expressed as:
[0067] tr((Y T Y) -1 )(17)
[0068] where tr represents the trace of the matrix;
[0069] Using the Lagrange multiplier method to obtain the objective cost function:
[0070]
[0071] where cost c represents the objective cost when distributed in the c-th cluster, and D ic represents the weighted Euclidean distance from the i-th population prevalence sample to the centroid of the c-th cluster. c and h take values from 1 to k.
[0072] Finally, the appropriate penalty terms p1 and p2 are manually selected according to the actual situation, and the cost function is continuously iterated and minimized until the change value in the iteration is less than the preset threshold (set to 0.001 in this method) or the number of iterations exceeds the set maximum number of times (set to 250 times in this method). The settings of the threshold and the maximum number of iterations will vary for different datasets and need to be determined according to the actual situation.
[0073] The present invention also provides an electronic device, including a processor and a memory. The memory stores program codes. When the program codes are executed by the processor, the processor is caused to execute the steps of the method.
[0074] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, the steps of the method are executed.
[0075] Beneficial effects: The method of the present invention overcomes the problem of selecting the initial clustering center by combining the cuckoo search algorithm, and significantly improves the clustering effect. At the same time, the method combines the weighted Euclidean distance, fairness, balance constraint with the traditional K-means clustering algorithm, effectively improving the balance of the distribution of population data sensitive data and the balance of the distribution of sample clusters. Description of the Drawings
[0076] Figure 1 is the overall flow chart of the method of the present invention.
[0077] Figure 2 is the clustering effect diagram of the Lloyd algorithm on the 2d-4c-no0 synthetic dataset.
[0078] Figure 3 is the fair clustering effect diagram considering sensitive attributes on the 2d-4c-no0 synthetic dataset.
[0079] Figure 4 is the clustering effect diagram of balanced fair K-means on the 2d-4c-no0 synthetic dataset.
[0080] Figure 5 is the clustering effect diagram of the present invention on the 2d-4c-no0 synthetic dataset. Detailed Embodiments
[0081] The following further specific descriptions of the present invention are made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0082] The embodiment of the present invention provides a census clustering method for fairness and balance constraint optimization, including the following steps:
[0083] Step 1: Obtain the census sampling data set X;
[0084] Step 2: Use the cuckoo search algorithm for optimization and select the initial population clustering centers;
[0085] Step 3: Conduct the silhouette coefficient comparison method within the K-value selection range to select the optimal K-value for population survey clustering;
[0086] Step 4: Perform weighted Euclidean distance optimization iteration;
[0087] Step 5: Introduce fairness and balance constraints;
[0088] Step 6: Iterate until the convergence condition is reached.
[0089] Step 2 includes:
[0090] Step 2-1: Initialize num_nest nests, and randomly generate K candidate initial population clustering centers in each nest. The formula for num_nest is:
[0091]
[0092] where n is the number of samples in the census sampling data set;
[0093] Calculate the sum of squared errors SSE (Sum of Squared Errors, SSE) of each nest as the fitness:
[0094]
[0095] where x j represents the population survey sample point in the i-th cluster, μ i is the population clustering center of the i-th cluster, and C i is the i-th cluster;
[0096] Step 2-2: First, randomly select a nest a, and then randomly select two different nests C i and C j , linearly combine the two different nests C i and C j into a new nest C new1 , compare the fitness of the new nest C new1 with that of nest a. If the fitness of the new nest C new1 is less than that of nest a, then replace a with the new nest C new1 , otherwise do not replace;
[0097] The formula for the linear combination is:
[0098] C new1 = αC i +(1 - α)Cj (3)
[0099] Among them, α is a random step factor, and its value ranges from 0 to 1;
[0100] Step 2-3: Perform Levy Flight on each nest. Each nest has a probability of 0.25 to perform Levy random search to increase diversity. In the solution of the present invention, a probability value of 0.25 is selected to perform Levy random search to increase the diversity during the search process. The selection of this probability value has been verified by experiments, which can balance the exploration and exploitation efficiency while ensuring the global search ability, and avoid excessive computational overhead caused by frequent Levy flights. By controlling the occurrence probability of Levy flight, this method effectively improves the convergence speed and computational efficiency of the algorithm, and is particularly suitable for high-dimensional optimization problems. Of course, the probability value of 0.25 in this technical solution is not the only choice, and other probability values can also be used to achieve similar effects. The specific choice should be optimized according to the actual application scenario and experimental results:
[0101] C new2 = C old + L(β) (4)
[0102] Among them, β is the Levy distribution parameter, C old is the initial nest of Levy flight, C new2 is C old the newly synthesized nest after Levy flight, and L(β) is the step size of Levy flight;
[0103] The step size calculation formula of Levy distribution is:
[0104]
[0105] Among them, u and v are random variables generated from the normal distribution, u ~ N(0,δ 2 ), v ~ N(0,1), δ is a parameter of Levy distribution, usually used to adjust the distribution characteristics of the step size; Γ is the gamma function, used to represent the factorial in mathematics;
[0106] To prevent the newly generated solution from exceeding the data range, the following boundary processing strategy is used:
[0107] C new2 = max(min(C new2 , max(X)), min(X)) (7);
[0108] Among them, max(X) is the upper limit of the census sampling data set X, and min(X) is the lower limit of the census sampling data set X;
[0109] Step 2-4: Repeat Steps 2-1 to 2-3 for iteration until the maximum number of iterations is reached, and output the optimal initial clustering centers. After multiple experimental comparisons, the maximum number of iterations in the present invention is set to 150 times.
[0110] Step 3 includes:
[0111] Step 3-1: Calculate the within-class distance a(i) of the i-th census sample:
[0112]
[0113] where l i is the number of the cluster to which the i-th census sample belongs, is the cluster to which the i-th census sample belongs, is the number of samples in the cluster to which the i-th census sample belongs, X i and X j respectively represent the i-th census sample and the j-th census sample; X i and X j belong to the same cluster;
[0114] Step 3-2: Calculate the between-class distance b(i) of the nearest cluster for the census data analysis:
[0115]
[0116] where C c represents the cluster with the number equal to c;
[0117] Step 3-3: Calculate the silhouette coefficient of each census sample;
[0118] Step 3-4: For the current K value, calculate the average silhouette coefficient S k ;
[0119] Step 3-5: Select the K value with the largest average silhouette coefficient as the number of clusters, and proceed to Step 4.
[0120] Step 3-3 includes: The calculation formula for the silhouette coefficient s(i) of the i-th census sample is:
[0121]
[0122] Step 3-4 includes: Calculate the average silhouette coefficient S using the following formula k :
[0123]
[0124] Step 4 includes:
[0125] Step 4-1: Calculate the standard deviation σ of the m-th featurem :
[0126]
[0127] where X im is the m-th eigenvalue of the i-th census sample in the census sampling dataset X, n is the total number of census samples, and μ m represents the mean of all samples of the m-th feature;
[0128] Step 4-2, when calculating the weight w, determine the weight by the standard deviation of each feature:
[0129] w m = 1 / σ m (13)
[0130] where w m represents the weight of the m-th feature;
[0131] Step 4-3, calculate the weighted Euclidean distance.
[0132] In Step 4-3, use the following formula to calculate the weighted Euclidean distance:
[0133]
[0134] where D ij is the weighted Euclidean distance from the i-th census sample to the j-th census sample, C jm is the m-th feature of the j-th cluster center, and M is the total number of features of the census samples.
[0135] Step 5 includes:
[0136] By using the indicator matrix, the fairness of the population-sensitive attribute is expressed as:
[0137]
[0138] where F represents the indicator matrix of the sensitive attributes of the population sampling data, and Y represents the indicator matrix of the cluster labels. In the present invention, the indicator matrix is used to represent the presence of certain features, attributes, or categories. The elements in the matrix can only take 0 or 1, where 1 indicates that a certain specific condition holds or a certain feature exists, and 0 indicates that the condition does not hold or the feature does not exist. In the prior art, the indicator matrix can be generated by automated data processing techniques;
[0139] When considering the balance constraint of population data analysis, in order to ensure that the size of each cluster is approximately the same as much as possible, this method uses the arithmetic mean inequality. When E h=b / k , E1 + E2 + E3 +... + E k = n, we get:
[0140]
[0141] Among them, E h represents the h-th cluster size, and k represents the total number of clusters;
[0142] The formula (16) is expressed using an indicator matrix as:
[0143] tr((Y T Y) -1 )(17)
[0144] Among them, tr represents the trace of a matrix;
[0145] The Lagrange multiplier method is used to obtain the objective cost function:
[0146]
[0147] Among them, cost c represents the objective cost when distributed in the c-th cluster, and D ic represents the weighted Euclidean distance from the i-th population prevalence sample to the centroid of the c-th cluster. The values of c and h range from 1 to k.
[0148] Finally, appropriate penalty terms p1 and p2 are manually selected according to the actual situation, and the cost function is continuously iterated and minimized until the change value in the iteration is less than the preset threshold (set to 0.001 in this method) or the number of iterations exceeds the set maximum number of iterations (set to 250 times in this method). The settings of the threshold and the maximum number of iterations will vary for different data sets and need to be determined according to the actual situation.
[0149] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the above method.
[0150] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the above method.
[0151] As Figure 1 shown, in this embodiment, the experimental effect is verified by using the synthetic datasets Elliptical, DS-577, 2d-4c-no0 commonly used in clustering research and the sampling data Census1990 of the 1990 US census.
[0152] Table 1
[0153]
[0154] As shown in Table 1, the table lists the parameter sizes of the penalty terms p1 and p2 set in the experiment and the scale of each dataset. In the face of different datasets, the penalty terms for fairness and balance constraints need to be adjusted according to the dataset scale and actual situation to ensure that the algorithm can achieve an optimized clustering effect on different datasets.
[0155] Such as Figure 2 , Figure 3 , Figure 4 and Figure 5 shown, respectively represent the clustering effects of different algorithms on the 2d-4c-no0 synthetic dataset. Among them, different colors represent different clusters, and different shapes represent sensitive attributes. By comparing with the other three algorithms, it can be clearly seen that the algorithm of the present invention achieves good clustering fairness while ensuring the clustering effect.
[0156] In this experiment, the algorithm of the present invention was compared with three clustering algorithms: the Lloyd algorithm, fairness clustering considering sensitive attributes, and balanced fairness K-means clustering. Among them, the Lloyd algorithm is the algorithm of the traditional K-means clustering algorithm, and fairness clustering considering sensitive attributes is a fairness optimization method based on Spectral Clustering. Balanced fairness K-means clustering introduces fairness and cluster size balance constraints on the basis of the standard K-means clustering algorithm. This experiment uses three evaluation indicators: sum of squared errors, weighted fairness gap, and fairness ratio to measure the clustering performance. The sum of squared errors measures the sum of the squares of the Euclidean distances between the clustering centers and the sample points. The lower the value, the closer the sample distribution within the cluster, and the better the clustering effect. The weighted fairness gap reflects the distribution balance degree of sensitive attributes in each cluster. The lower the value, the more uniform the distribution of sensitive attributes among the clusters. The fairness ratio is used to measure whether the distribution of sensitive attributes in the dataset meets the fairness goal. A fairness ratio close to 1 represents the best fairness, while 0 means that fairness constraints are not considered. Generally speaking, the sum of squared errors focuses on the tightness of clustering, the weighted fairness gap evaluates fairness deviation, and the fairness ratio measures the overall fairness level. The three together constitute a comprehensive evaluation system for clustering algorithms in this experiment.
[0157] Table 2
[0158] Sum of Squared Errors ↓ Weighted Fairness Gap ↓ Fairness Ratio ↑ Lloyd's algorithm 206.2982 0.494 0 Fair Clustering Considering Sensitive Attributes 343.964 0.048 0.8845 Balanced Fair K-Means Clustering 351.109 0.0472 0.9005 The algorithm of the present invention 329.0549 0 1
[0159] Table 3
[0160] Sum of Squared Errors ↓ Weighted Fairness Gap ↓ Fairness Ratio ↑ Lloyd's algorithm 71.0134 0.4183 0 Fair Clustering Considering Sensitive Attributes 361.4299 0.1496 0 Balanced Fair K-Means Clustering 516.0655 0.0319 0.8042 The algorithm of the present invention 514.4775 0.0290 0.8405
[0161] Table 4
[0162]
[0163]
[0164] Table 5
[0165] Sum of Squared Errors ↓ Weighted Fairness Gap ↓ Fairness Ratio ↑ Lloyd's algorithm 1760.4 0.0671 0.5129 Fair Clustering Considering Sensitive Attributes 1821.9 0.047 0.7418 Balanced Fair K-Means Clustering 1852 0.0141 0.9169 The algorithm of the present invention 1722 0.0135 0.9516
[0166] As shown by the experimental results in Table 2, Table 3, Table 4 and Table 5, the algorithm of the present invention achieves the best balance between fairness and clustering quality, demonstrating significant advantages. First, the algorithm of the present invention performs best in terms of fairness. In all datasets, the fairness ratio of the algorithm of the present invention reaches or is close to 1.0, far exceeding fair clustering considering sensitive attributes and balanced fair K-means clustering. For example, on the Elliptical dataset, the algorithm of the present invention achieves a fairness ratio of 1.0, that is, a completely fair cluster partition, while on the Census1990 dataset, the algorithm of the present invention also reaches a fairness ratio of 0.9516, demonstrating extremely high fairness optimization ability. In contrast, the Lloyd's algorithm does not consider fairness, and the fairness ratio is always 0, while fair clustering considering sensitive attributes and balanced fair K-means clustering show some improvement on some datasets, but still fall short of the optimization effect of the algorithm of the present invention. Second, the algorithm of the present invention performs optimally in terms of weighted fairness gap. The weighted fairness gap reflects the degree of distribution balance of sensitive attributes in different clusters in the dataset, and the lower the value, the higher the fairness. The weighted fairness gap values of the algorithm of the present invention on all datasets are significantly lower than those of other methods. This result indicates that the algorithm of the present invention can effectively reduce the uneven distribution of sensitive attributes on different datasets and achieve optimal fairness control. Finally, the algorithm of the present invention takes into account the clustering compactness and maintains a reasonable sum of squared errors. Although the Lloyd's algorithm performs best in terms of the sum of squared errors, it completely ignores fairness, resulting in serious data bias. In contrast, the algorithm of the present invention can still maintain a low sum of squared errors while optimizing fairness, avoiding the excessive impact of fairness optimization on clustering quality. For example, on the Elliptical dataset, the sum of squared errors of the algorithm of the present invention is 329.0549, lower than that of balanced fair K-means clustering (351.109) and fair clustering considering sensitive attributes (343.964), indicating that its clustering quality is still relatively good; on the Census1990 dataset, the sum of squared errors of the algorithm of the present invention is 1722.0, also better than that of balanced fair K-means clustering (1852.0) and fair clustering considering sensitive attributes (1821.9), further demonstrating its stability on large-scale data.
[0167] Generally speaking, the algorithm of the present invention demonstrates the optimal fairness performance on all datasets and maintains a relatively high clustering quality on the premise of ensuring fairness. This makes the algorithm of the present invention have great application significance in census data analysis with high fairness requirements.
[0168] The present invention provides a census clustering method for fairness and balance constraint optimization. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A census clustering method for fairness and balance constraint optimization, characterized in that: The following steps are involved: Step 1, obtain the population census sampling data set X; Step 2, using the cuckoo search algorithm for optimization and selecting the initial population cluster center; Step 3, perform the silhouette coefficient comparison method within the K value selection interval to select the best K value for population survey clustering; Step 4, perform weighted Euclidean distance optimization iteration; Step 5, introduce fairness and balance constraints; Step 6, iterate until convergence condition is reached.
2. The method according to claim 1, characterized in that Step 2 includes: Step 2-1, initialize num_nest nests, randomly generate K candidate initial population cluster centers in each nest, and the num_nest formula is: Where n is the sample size of the census sampling data set; Calculate the squared error SSE for each nest as the fitness: Among them, x j represents the population survey sample point in the i-th cluster, μ i is the population cluster center of the ith cluster, C i is the i-th cluster; Step 2-2, first randomly select a nest a, then randomly select two different nests C i and C j , two different nests C i and X j Linear combination into a new nest C new1 , compared with Xinchao C new1 With the fitness of nest a, if the new nest C new1 The fitness of nest C is less than that of nest a, so the new nest C is used. new1 Replace a, otherwise do not replace; The formula of the linear combination is: C new1 =αC i +(1-α)C j (3) Among them, α is the random step factor, ranging from 0 to 1; Steps 2-3, perform Levy flights on each nest and Levy random search on each nest to increase diversity: C new2 =C old +L(β) (4) Where β is the Levy distribution parameter, X old It is the initial nest of Levy flight, X new2 It is C old The newly synthesized nest after Levy flight, L(β) is the step length of Levy flight; The step length calculation formula of Levy distribution is: Among them, u and v are random variables generated from normal distribution, u~N(0,δ 2 ), v~N(0,1), δ is a parameter of Levy distribution; Γ is the gamma function; Use the following boundary handling strategy: C new2 =max(min(C new2 ,max(X)),min(X)) (7); Among them, max(X) is the upper limit of the census sampling data set X, and min(X) is the lower limit of the census sampling data set X; Step 2-4, repeat steps 2-1 to 2-3 for iteration until the maximum number of iterations is reached, and output the best initialized cluster center.
3. The method according to claim 2, characterized in that Step 3 includes: Step 3-1, calculate the intra-class distance a(i) of the i-th census sample: Among them, l i is the number of the cluster to which the i-th census sample belongs, is the cluster to which the i-th census sample belongs, is the number of samples in the cluster to which the ith census sample belongs, X i and X j represent the i-th census sample and the j-th census sample respectively; X i and X j Belong to the same cluster; Step 3-2, calculate the inter-cluster distance b(i) of the nearest cluster for census data analysis: Among them, C c represents the cluster with number equal to c; Step 3-3, calculate the silhouette coefficient for each census sample; Step 3-4, for the current K value, calculate the average silhouette coefficient S k ; Step 3-5, select the K value with the largest average silhouette coefficient as the number of clusters and proceed to step 4.
4. The method according to claim 3, characterized in that Step 3-3 includes: The calculation formula of the silhouette coefficient s(i) of the i-th population census sample is:
5. The method according to claim 4, characterized in that Step 3-4 includes: calculating the average silhouette coefficient S using the following formula k :
6. The method according to claim 5, characterized in that Step 4 includes: Step 4-1, calculate the standard deviation σ of the mth feature m : Among them, X im is the mth eigenvalue of the i-th census sample in the census sampling data set X, n is the total number of census samples, μ m Represents the mean of all samples of the mth feature; Step 4-2, when calculating the weight w, the weight is determined by the standard deviation of each feature: w m =1 / s m (13) where w m represents the weight of the mth feature; Step 4-3, calculate the weighted Euclidean distance.
7. The method according to claim 6, characterized in that In step 4-3, the weighted Euclidean distance is calculated using the following formula: Among them, D ij is the weighted Euclidean distance from the i-th census sample to the j-th census sample, C jm is the mth feature of the jth cluster center, and M is the total number of features of the census sample.
8. The method according to claim 7, characterized in that Step 5 includes: By using the indicator matrix, the fairness of the population-sensitive attribute is expressed as: Among them, F represents the indicator matrix of sensitive attributes of population sampling data, and Y represents the indicator matrix of cluster labels; By the arithmetic mean inequality, when E h=n / k , E1+E2+E3+...+E k =n, we get: Among them, E h represents the size of the hth cluster, and k represents the total number of clusters; Using the indicator matrix, formula (16) is expressed as: tr((And T AND) -1 ) (17) Where tr represents the trace of the matrix; The target cost function is obtained using the Lagrange multiplier method: Among them, cost c represents the target cost when distributed in the cth cluster, D ic It represents the weighted Euclidean distance from the i-th population popular sample to the centroid of the c-th cluster.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
New energy output scene analysis method and system based on improved k-means clustering algorithm
CN113239503A
Method, system and device for clustering analysis of crowd based on spatio-temporal data
CN119513636A