Intelligent community demand division method based on hierarchical clustering and K-means clustering
By using hierarchical clustering and K-means clustering methods, a multi-subject demand system is constructed, and the optimal number of clusters is adaptively determined, which solves the problem of supply and demand mismatch in smart community construction and achieves precise demand segmentation and resource optimization.
Patent Information
- Application Number
- CN202511435004.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies lack accurate quantification of the needs of diverse stakeholders in the construction of smart communities, leading to supply and demand mismatch, and a disconnect between resource allocation and community development. Traditional clustering algorithms are easily affected by initial values and ignore the dynamic evolution of needs.
We employ hierarchical clustering and K-means clustering methods, construct a multi-subject demand system, collect data using questionnaires, calculate Euclidean distance and clustering coefficient, adaptively determine the optimal number of clusters, and perform final optimization partitioning.
It enables precise segmentation of the needs of diverse stakeholders, improves the interpretability and comprehensiveness of demand assessment, reduces data noise sensitivity, ensures the scientific nature and repeatability of the segmentation results, and supports differentiated service strategies and resource optimization.
Smart Images

Figure CN120893802A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a demand division method for a smart community, in particular to a demand division method for a smart community based on hierarchical clustering and K-means clustering. BACKGROUND
[0002] With the continuous development of Internet communication technology, smart community construction has emerged in many parts of the world. In the process of practice, the participants in the construction of smart communities are increasingly diverse, and the needs of all parties are slightly different. Due to the lack of accurate quantification of the needs of multiple subjects, the supply and demand of smart community construction are mismatched. Therefore, how to meet the needs of multiple subjects for smart community construction, clarify the demand types of multiple subjects of smart community, and improve the adaptability of the supply system of smart community construction are the key and difficult points of smart community construction in the new era.
[0003] At present, there are significant technical bottlenecks in the construction of smart communities. Traditional demand identification methods mainly rely on scattered research of a single subject (such as residents or the government), and lack systematic quantitative analysis of the collaborative needs of four core subjects such as residents, property service enterprises, management personnel and social organizations, resulting in incomplete coverage of the demand index system (such as the fragmentation of three dimensions of community safety, livelihood services and community governance). Existing technologies mostly use basic clustering algorithms (such as single K-means clustering algorithm) to process high-dimensional demand data, which is easily disturbed by initial values and produces classification bias, making it difficult to accurately classify the complex demand types in the community, and ignoring the demand dynamic evolution law. This causes a serious disconnection between resource allocation and the actual development stage of the community, and the supply and demand mismatch problems such as misallocation of service facilities in old communities and lagging safety construction in emerging communities are prominent, which restricts the effectiveness of smart community construction and resource utilization efficiency. SUMMARY
[0004] In order to overcome the defects of the prior art, the present application provides a demand division method for a smart community based on hierarchical clustering and K-means clustering.
[0005] The technical scheme adopted by the present application is as follows: a demand division method for a smart community based on hierarchical clustering and K-means clustering, comprising the following steps:
[0006] S1: Establishing a demand system for multiple subjects of a smart community, and subdividing the demand system into primary demand indicators and secondary demand indicators under the primary demand indicators.
[0007] S2: Preparing a questionnaire based on the demand system for multiple subjects, collecting samples, and calculating the primary demand indicator score and the secondary demand indicator score of each sample.
[0008] S3: Based on hierarchical clustering analysis, iteratively updating the category by calculating the average Euclidean distance, determining the best cluster number according to the convergence of the agglomeration coefficient, and calculating the clustering category under the best cluster number.
[0009] S4: updating the samples in the clustering categories of S3 based on the K-means clustering analysis method, performing final optimization division on samples, and achieving the need division of the multiple subjects in the smart community.
[0010] Further, step S2 specifically includes the following steps:
[0011] S2.1: compiling the questionnaire based on the multiple subject demand system.
[0012] S2.2: distributing the questionnaire, performing reliability and validity analysis on the collected data, and excluding invalid data.
[0013] S2.3: obtaining the samples collected, calculating the secondary demand index score of the samples, and calculating the primary demand index score based on the secondary demand index score.
[0014] Further, the primary demand index score is obtained by averaging the secondary demand index score.
[0015] Further, step S3 specifically includes:
[0016] S3.1: constructing n initial categories corresponding to the number of samples, and setting 1 sample in each initial category.
[0017] S3.2: calculating the average Euclidean distance between the initial categories, merging the initial categories according to the nearest average Euclidean distance to obtain updated categories, and the calculation formula of the average Euclidean distance is as follows:
[0018]
[0019] where i and j are the i-th and j-th samples, respectively, corresponding to the initial category i and the initial category j; m represents the number of primary demand indexes, represents the average Euclidean distance between the initial category i and the initial category j; and are the i-th and j-th primary demand index scores in the initial category i and the initial category j, respectively.
[0020] S3.3: calculating the difference deviation sum of the categories, and the calculation formula of the difference deviation sum is as follows:
[0021]
[0022] wherein, represents the deviation sum of the category p, represents the number of samples in category p, , , These represent the first in category p. The scores of the 1st, 2nd, and mth primary demand indicators for each sample.
[0023] S3.4: Calculate the sum of squared differences of the new category after merging any two categories:
[0024]
[0025] in, and Let these represent the sum of squared deviations of category p and category q, respectively. This indicates that the sum of squared deviations of category p and category q are obtained by merging them to obtain category l. The average Euclidean distance between category p and category q.
[0026] S3.5: Calculate the average Euclidean distance between the stated category l and any other category r. :
[0027]
[0028] in, and Let r represent the number of samples in category r and l, respectively. and These represent the average Euclidean distance between category p and category r, and the average Euclidean distance between category q and category r, respectively.
[0029] S3.6: Merge each pair of clusters according to the most recent average Euclidean distance to obtain the updated category. Repeat S3.3-S3.5 to determine the optimal number of clusters s based on the convergence of the clustering coefficient, and obtain the optimal number of clusters s.
[0030] Furthermore, the average Euclidean distance between category p and category q The calculation formula is:
[0031]
[0032] in, The difference between the sum of squared deviations of category p and category q. and Let p and q represent the number of samples for category p and category q, respectively. This represents the centroid distance between category p and category q.
[0033] Furthermore, the aforementioned The calculation formula of the first demand index score is:
[0034]
[0035] wherein, and respectively represent the first demand index score of the i-th sample in the category p and the i-th sample in the category q.
[0036] Further, the optimal cluster number s represents the number of categories required by the multi-subjects of the wisdom community construction. According to the corresponding cohesion coefficient of different cluster numbers, a curve of the cohesion coefficient changing with the cluster number is drawn, and the optimal cluster number is determined according to the convergence of the cohesion coefficient.
[0037] Further, step S4 specifically comprises:
[0038] S4.1: calculating the cluster center of the cluster category, and calculating the distance of each sample to each cluster center:
[0039]
[0040] wherein, represents the i-th sample, represents the cluster center of the t-th cluster category.
[0041] S4.2: assigning each sample to the cluster category where the cluster center closest to the sample is located, and updating the sample of the cluster category.
[0042] S4.3: repeating S4.1-S4.2 until the cluster center of the cluster category after updating the sample is equal to the cluster center of the cluster category before updating or the error sum of squares does not change.
[0043] Further, when the cluster center of the cluster category after updating the sample is equal to the cluster center of the cluster category before updating:
[0044] .
[0045] Further, the calculation formula of the first demand index score is:
[0046]
[0047] wherein, represents the number of samples in the cluster category. representing the sample contained in the cluster category.
[0048] Further, if is not true or the error sum of squares changes, then , return to S4.1; if is true or the error sum of squares does not change, then end.
[0049] Compared with the prior art, the present application has at least the following advantages: (1) The present application combines the dual advantages of hierarchical clustering analysis method and K-means clustering analysis, and determines the optimal number of clusters adaptively through hierarchical clustering analysis, thereby overcoming the problem that the traditional K-means clustering is sensitive to the initial number of clusters and is prone to local optimization; on this basis, K-means clustering is used to efficiently adjust the sample attribution, thereby significantly improving the accuracy and stability of demand grouping, and making the final division result more reasonable and reliable. (2) The present application constructs a multi-element demand index system covering residents, property service enterprises, management personnel and social organizations, constructs a demand system including primary and secondary indexes, decomposes the abstract community demand into specific indexes which are operable and measurable layer by layer, and realizes standardized collection of demand data by means of questionnaire survey, thereby providing a solid data foundation for subsequent clustering analysis, and greatly improving the interpretability and comprehensiveness of demand evaluation. (3) The present application uses average Euclidean distance to calculate the distance between classes, which can reduce the sensitivity to abnormal values and data noise compared with single connection or full connection, thereby laying a solid foundation for subsequent K-means clustering, avoiding the problem of slow convergence or poor effect caused by random initialization, and obtaining a more optimal final division result. (4) The present application determines the optimal number of clusters by calculating and comparing the agglomeration coefficients under different numbers of clusters, and determines the optimal number of clusters according to the convergence of the agglomeration coefficients, thereby providing a quantitative and objective basis for stopping iteration and selecting the final classification scheme. This overcomes the limitation of relying on prior knowledge or subjective experience to determine the number of clusters in clustering analysis, realizes adaptive determination of the key parameters of the model, and guarantees the scientificity and repeatability of the method. (5) The present application realizes accurate grouping of various demands of the community by identifying the potential differences in demand structure and intensity of different groups, thereby providing reliable basis for subsequent differentiated service strategy formulation and resource optimization, and avoiding resource waste or service deviation caused by "one-size-fits-all" construction. BRIEF DESCRIPTION OF DRAWINGS
[0050] The present application will be described in detail below with reference to the accompanying drawings:
[0051] Figure 1 is a step flowchart of the demand division method of the smart community based on hierarchical clustering and K-means clustering according to the embodiment 1 of the present application.
[0052] Figure 2 is the average score of the primary demand index of each party subject in embodiment 1 of the present application.
[0053] Figure 3 is a schematic diagram of the relationship between the aggregation coefficient and the number of aggregations in embodiment 1 of the present application.
[0054] Figure 4 is the final demand division of the smart community in embodiment 1 of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0056] Embodiment 1: A smart community demand division method based on hierarchical clustering and K-means clustering, comprising the following steps:
[0057] S1: Establishing a smart community multi-subject demand system, and subdividing the demand system into primary demand indexes and secondary demand indexes under each primary demand index.
[0058] The demand system of the present embodiment fully considers the demands of four parties, including residents, property service enterprise managers and social organizations.
[0059] In the present embodiment, the demand system of the smart community multi-subject is constructed based on the ERG (Existence Relatedness Growth) theory, Maslow's demand hierarchy theory and 4C (Consumer, Cost, Convenience, Communication) theory. The smart community multi-subject demand is initially divided into three types of community safety, community service and community governance, and subdivided into 45 secondary demand indexes, and a smart community construction four-party demand system is initially constructed, as shown in the following table. According to the expert opinions, 3 primary demand indexes and 45 secondary demand indexes are constructed in the table, in which there are 20 secondary demand indexes under the primary demand index of “community safety”, 8 secondary demand indexes under the primary demand index of “livelihood service”, and 17 secondary demand indexes under the primary demand index of “community governance”.
[0060] Table 1 Smart community construction four-party demand system
[0061]
[0062]
[0063] S2: compiling a questionnaire based on the multi-subject demand system, collecting the first-level demand index score and the second-level demand index score.
[0064] S2 specifically includes:
[0065] S2.1: compiling a questionnaire based on the multi-subject demand system;
[0066] Based on the demand system of residents, property service enterprises, management personnel and social organizations in the construction of smart communities, a demand questionnaire for the construction of smart communities for residents, property service enterprises, management personnel and social organizations is designed. The questionnaire is divided into three parts: first, basic personal information, including community location, gender, age, education level, understanding of smart community services, and work experience; second, smart community construction demand survey, in which respondents score the demand degree of community safety, livelihood services and community governance; third, other suggestions, aiming to understand the respondents' other ideas about the construction of smart communities.
[0067] S2.2: distributing the questionnaire, analyzing the collected data for reliability and validity, and excluding invalid data;
[0068] S2.3: based on the collected results, obtaining samples and obtaining the second-level demand index score of each sample, and calculating the first-level demand index score based on the second-level demand index score. The first-level demand index score is obtained by averaging the second-level demand index score.
[0069] Considering the funding support, subject cognition, and construction effectiveness of smart communities, 32 typical smart communities in Zhengzhou, Luoyang, Putian, Shenzhen, Dongguan, and Huizhou were selected as survey objects, and large sample survey data of the 32 typical smart communities in the six cities were collected. After analyzing the collected data for reliability and validity and excluding invalid data, a total of 1606 valid questionnaire data were collected. According to the 1606 survey data, the first-level demand index scores of the parties in the six cities were obtained. Figure 2 ).
[0070] S3: based on hierarchical clustering analysis, updating the categories by calculating the average Euclidean distance iteratively, determining the best cluster number according to the convergence of the agglomeration coefficient with the number of clusters, and calculating the clustering categories under the best cluster number;
[0071] Step S3 specifically includes:
[0072] S3.1: constructing n initial categories corresponding to the number of samples, and setting 1 sample in each initial category;
[0073] The number of samples in this embodiment is the number of valid questionnaires collected 1606, and 1606 initial categories are constructed according to the number of samples.
[0074] S3.2: Calculate the average Euclidean distance between the initial categories , and the new categories are obtained by merging the initial categories in pairs according to the nearest distance. The calculation formula of the average Euclidean distance is as follows:
[0075]
[0076] In the above formula, i and j are the i-th sample and the j-th sample in the n samples, respectively, corresponding to the i-th initial category and the j-th initial category; m represents the number of primary demand indicators, and the number of m in this embodiment is 3; is the average Euclidean distance between the initial category i and the initial category j; and are the scores of the i-th primary demand indicator in the i-th sample and the j-th sample, respectively.
[0077] S3.3: Calculate the difference deviation sum of squares of the new categories, and the calculation formula is as follows:
[0078]
[0079] In the above formula, represents the deviation sum of squares of the category p, represents the number of samples of the category p, , , represent the scores of the 1st, 2nd, and m-th primary demand indicators of the i-th sample in the category p, respectively; S3.4: Calculate the deviation sum of squares of the new category after merging any two categories:
[0080]
[0081] In the above formula,
[0082] and represent the deviation sum of squares of the category p and the category q, respectively, represents the deviation sum of squares of the category l obtained by merging the category p and the category q. wherein the average Euclidean distance between the category p and the category q
[0083] is the difference between the deviation sum of squares of the two categories , that is:
[0084]
[0085]
[0086] In the above formula, and respectively represent the number of samples of category p and category q, represents the centroid distance of category p and category q, and respectively represent the first i sample of category p and the first i sample of category q.
[0087] S3.5: Calculate the average Euclidean distance between the new category l and any other category r:
[0088]
[0089] In the above formula, and respectively represent the number of samples of category r and category l, and respectively represent the average Euclidean distance of category p and category r and the average Euclidean distance of category q and category r.
[0090] S3.6: Merge all categories by the nearest distance to obtain new categories, repeat S3.3-S3.5 until the optimal number of clusters s clusters are obtained.
[0091] Wherein, the optimal number of clusters s represents the number of categories required by the multi-subjects of the smart community construction, according to the corresponding cohesion coefficient of different cluster numbers, the curve of the cohesion coefficient changing with the cluster number is drawn, and the optimal cluster number is determined according to the convergence of the cohesion coefficient. In this embodiment, the change of the cohesion coefficient with the cluster number is calculated by Ward's method and Squared Euclidean distance in SPSS 27 software, and the optimal cluster number is finally determined according to the convergence of the cohesion coefficient. In this embodiment, the number of categories of multi-subjects demand determined by calculating the relationship of the cohesion coefficients of each cluster solution is 4. Figure 3 ).
[0092] S4: Based on the K-means clustering analysis method, update the samples in the cluster categories of S3, and finally optimize the division of the samples to realize the division of the needs of the multi-subjects in the smart community.
[0093] Wherein, S4 specifically includes:
[0094] S4.1: Calculate the cluster center of the cluster category, and calculate the distance of each sample to each cluster center:
[0095]
[0096] In the above formula, denotes the i-th sample, denotes the t-th cluster center of the cluster class.
[0097] In the above formula, denotes the t-th cluster center of the cluster class The calculation formula is:
[0098]
[0099] In the above formula, denotes the number of samples in the cluster class denotes the sample contained in the cluster class
[0100] S4.2: Assign each sample to the cluster class with the smallest distance to its cluster center, and update the samples in the cluster class;
[0101] S4.3: Repeat S4.1-S4.2 until the cluster center of the cluster class after updating the samples is equal to the cluster center of the cluster class before updating or the error sum of squares does not change.
[0102] If the cluster center of the cluster class after updating the samples is equal to the cluster center of the cluster class before updating , then:
[0103]
[0104] In the above formula, denotes the cluster center of the cluster class after updating in S4.2.
[0105] If the equation is not established or the error sum of squares changes, then , return to S4.1; if the equation is established or the error sum of squares does not change, then end. Finally, the s types of demand of multiple subjects in the construction of smart community can be obtained, and the demand degree of multiple subjects for community safety, livelihood services and community governance under each type of demand can be obtained.
[0106] According to Figure 4 The cluster analysis results show that the diverse needs of various stakeholders in smart community construction can be divided into four categories. These four categories exhibit significant differences in cluster center scores, and the emphasis of construction content varies across each category. Based on the clustering characteristics of each need type, the first to fourth clustered needs can be categorized as follows: "Primarily focused on public services," "Primarily focused on community safety and governance," "Comprehensive needs," and "Insufficient needs."
[0107] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0108] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the invention is illustrative rather than restrictive, and the scope of the invention is defined by the appended claims.
Claims
1. A method for segmenting smart community needs based on hierarchical clustering and K-means clustering, characterized in that, Includes the following steps: S1: Establish a multi-stakeholder demand system for smart communities, and subdivide the demand system into primary demand indicators and secondary demand indicators under the primary demand indicators; S2: Based on the aforementioned multi-subject demand system, develop a questionnaire, collect samples, and calculate the primary demand index score and secondary demand index score for each sample. S3: Based on the hierarchical clustering analysis method, the categories are updated iteratively by calculating the average Euclidean distance, the optimal number of clusters is determined according to the convergence of the clustering coefficient, and the cluster categories under the optimal number of clusters are obtained. S4: Based on the K-means clustering analysis method, update the samples of the cluster categories in S3, perform final optimization division of the samples, and realize the required division of the multiple subjects in the smart community.
2. The smart community demand segmentation method based on hierarchical clustering and K-means clustering according to claim 1, characterized in that, S2 specifically includes: S2.1: The questionnaire is developed based on the aforementioned multi-subject needs system; S2.2: Distribute the questionnaire, conduct reliability and validity analysis on the collected data, and exclude invalid data; S2.3: Obtain the collected samples, calculate the scores of the secondary demand indicators for the samples, and calculate the scores of the primary demand indicators based on the scores of the secondary demand indicators.
3. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 2, characterized in that, The score of the primary demand indicator is obtained by averaging the scores of the secondary demand indicators.
4. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 1, characterized in that, S3 specifically includes: S3.1: Construct n initial categories corresponding to the number of samples, with each initial category containing 1 sample; S3.2: Calculate the average Euclidean distance between the initial categories. The initial categories are merged pairwise according to the nearest average Euclidean distance to obtain the updated categories; ; Where i and j are the i-th and j-th samples, respectively, corresponding to the initial category i and the initial category j; m represents the number of the primary demand indicators. This represents the average Euclidean distance between the initial category i and the initial category j; and They are respectively the first of the initial categories i and j. The scores of the first-level demand indicators mentioned above; S3.3: Calculate the sum of squared differences for the categories, using the following formula: ; in, This represents the sum of squared deviations of the category p. This represents the number of samples in category p. , , These represent the first and second elements in category p, respectively. The scores of the 1st, 2nd, and mth primary demand indicators for each sample; S3.4: Calculate the sum of squared differences between any two categories after merging them: ; in, and Let these represent the sum of squared deviations of category p and category q, respectively. This indicates that the sum of squared deviations of category p and category q are obtained by merging them to obtain category l. The average Euclidean distance between category p and category q; S3.5: Calculate the average Euclidean distance between category l and any other category r. : ; in, and Let r represent the number of samples in category r and l, respectively. and These represent the average Euclidean distance between category p and category r, and the average Euclidean distance between category q and category r, respectively. S3.6: Merge each pair of clusters according to the most recent average Euclidean distance to obtain the updated category. Repeat S3.3-S3.5 until the optimal number of clusters, s, is obtained based on the convergence of the clustering coefficient.
5. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 4, characterized in that, The average Euclidean distance between category p and category q The calculation formula is: ; in, The difference between the sum of squared deviations of category p and category q. and Let p and q represent the number of samples for category p and category q, respectively. This represents the centroid distance between category p and category q.
6. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 5, characterized in that, The The calculation formula is: ; in, and These represent the i-th sample of category p and the i-th sample of category q, respectively. The scores for the first-level demand indicators are as follows.
7. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 1, characterized in that, The optimal number of clusters s is the number of categories required by the diverse stakeholders in the construction of a smart community. A curve showing the change of the clustering coefficient with the number of clusters is plotted, and the optimal number of clusters is determined based on the convergence of the clustering coefficient.
8. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 1, characterized in that, S4 specifically includes: S4.1: Calculate the cluster centers of the cluster categories and calculate the distance from each sample to each cluster center: ; in, Let i represent the i-th sample. Represents the t-th cluster category Cluster centers; S4.2: Assign each sample to the cluster category containing the nearest cluster center, and update the samples in the cluster category; S4.3: Repeat S4.1-S4.2 until the cluster centers of the cluster categories after updating the samples are obtained. The cluster centers of the cluster categories before the update The sum of squares of the errors remains unchanged if they are equal.
9. The method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 8, characterized in that, The The calculation formula is: ; in, This indicates the number of samples in the cluster category; This indicates the samples included in the cluster category.
10. A method for segmenting smart community needs based on hierarchical clustering and K-means clustering according to claim 8, characterized in that, The cluster centers of the cluster categories after updating the samples. The cluster centers of the cluster categories before the update When they are equal: ; If the above If it does not hold true or the sum of squared errors changes, then Return to S4.1; if the If the condition is met or the sum of squared errors remains unchanged, the process ends.
Citation Information
Patent Citations
Modified K-means clustering algorithm based on hierarchical clustering
CN104102726A
A county power grid development demand hierarchy dividing method based on the Maslow's theory
CN105262085A
A gastroesophageal reflux disease dangerous factor extraction method and system based on accurate clustering
CN109685139A
Active power distribution network element self-organizing division and evaluation method
CN115062999A