Data management method of cross-domain data automatic clustering algorithm
Through the combination of distributed collaborative clustering method, multi-domain kernel density estimation and adaptive nearest neighbor search algorithm, the problem of inefficient data clustering across domains is solved, and accurate clustering and efficient utilization of large-scale data is achieved.
Patent Information
- Application Number
- CN202411997371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing cross-domain data clustering technologies are inefficient when processing large-scale distributed data and cannot meet the high requirements of data management and analysis.
The distributed collaborative clustering method is adopted to optimize the communication mechanism and data transmission strategy between nodes, and combine the clustering algorithm based on multi-domain kernel density estimation and adaptive neighbor search to achieve collaborative optimization of local clustering and global clustering.
It significantly improves the efficiency and scalability of data processing, realizes accurate clustering and efficient utilization of large-scale cross-domain data, and ensures the accuracy and stability of clustering results.
Smart Images

Figure CN119989010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a data management method of a cross-domain data automatic clustering algorithm. Background Art
[0002] In today's digital age, data in various fields are growing explosively, and the integration and analysis of cross-domain data has become a key requirement for tapping potential value and promoting innovation and development. However, these cross-domain data often have the characteristics of wide sources, complex structures, large scale, and inconsistent semantics and formats, which brings huge challenges to data management and analysis.
[0003] At present, traditional data clustering methods have many defects when processing cross-domain data. For cross-domain data collaborative clustering in a distributed environment, the existing technology lacks effective optimization of the inter-node communication mechanism, resulting in low data transmission efficiency and high communication overhead, which seriously affects the overall efficiency of clustering. At the same time, in terms of cross-domain data processing, it mainly focuses on data mining tasks within a single field. Faced with large-scale cross-domain data distributed on multiple computing nodes, the existing centralized clustering methods seem to be unable to cope with it. Due to the huge amount of data, directly aggregating all data to a central node for processing will lead to serious network bandwidth consumption and long waiting times, which is difficult to meet the high requirements for data quality in practical applications.
[0004] In summary, the existing cross-domain data clustering technology cannot meet the growing needs of data management and analysis. There is an urgent need for an efficient data management method of cross-domain data automatic clustering algorithm to solve the above problems and achieve accurate clustering and efficient utilization of large-scale cross-domain data. Summary of the invention
[0005] The purpose of the present invention is to make up for the shortcomings of the prior art and to provide a data management method for a cross-domain data automatic clustering algorithm. It can design a distributed collaborative clustering method for scenarios where large-scale cross-domain data is distributed on multiple computing nodes. By optimizing the communication mechanism and data transmission strategy between nodes, it can achieve effective coordination between local clustering and global clustering of data on different nodes, while ensuring clustering accuracy, significantly improving the efficiency and scalability of data processing.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a data management method for a cross-domain data automatic clustering algorithm, the specific steps of the method are:
[0007] S100, build a distributed architecture: set up a central control node and multiple computing nodes. The central control node is responsible for information integration and strategy formulation, and the computing nodes are used to perform local data processing tasks and preliminary clustering;
[0008] S200, data preprocessing and data partitioning: the central control node preprocesses the input cross-domain raw data, divides the data into multiple data subsets, and distributes the data subsets to the computing nodes;
[0009] S300, local clustering: After each computing node receives the allocated data subset, the processing unit performs local clustering using an automatic clustering algorithm. During the clustering process, the node dynamically adjusts the clustering parameters according to the distribution characteristics of the local data to accurately identify the clustering structure in the data;
[0010] When the automatic clustering algorithm performs local clustering, a clustering algorithm based on multi-domain kernel density estimation and adaptive nearest neighbor search is adopted, wherein:
[0011] For the kernel density estimation of multi-domain data, we adopt the kernel density estimation based on multi-domain and introduce the domain adaptive bandwidth adjustment factor h d h d =σ d ×(n d / N) -1 / (4+d) , where h d is the bandwidth of domain d, σ d is the standard deviation of the data in field d, n d is the number of data points in domain d, N is the total number of data points, d is the data dimension, and the adjustment factor h d Adjust the bandwidth of the kernel density estimation, and then perform kernel density estimation calculation on the data points to obtain the kernel density value of each data point to reflect the distribution density of the data;
[0012] The adaptive nearest neighbor search clustering algorithm is based on the local density ρ of the data points. i and the neighbor weight function w of the relative distance ij To achieve where w ij is the neighbor weight between data points i and j, ρ i is the local density of data point i and d ij is the distance between i and j, N i is the set of neighboring points of data point i, α and β are weight parameters used to balance the influence of local density and relative distance in neighbor judgment;
[0013] During the local clustering process, the memory unit records the detailed information of the cluster in real time and uploads it to the central control node. At the same time, it records the key information of each cluster, including the coordinates of the cluster center C. ij , Member data point identification ID ij , the cluster density estimate ρ cluster , and store it in a local storage unit;
[0014] S400, inter-node communication coordination: the central control node collects the local clustering information uploaded by each computing node, and integrates the information using a global clustering coordination algorithm to obtain global clustering information, and feeds back the integrated global clustering information to each computing node;
[0015] S500, global clustering optimization: each computing node further optimizes and adjusts the local clustering result according to the global clustering information fed back by the central control node, repeats the steps of S300 and S400, and performs multiple iterations until the clustering convergence condition is met, determines that the clustering process converges, and stops iteration;
[0016] S600, clustering result storage and management: After clustering convergence, the final data clustering is completed, and the central control node integrates and summarizes the global clustering results to generate a detailed and comprehensive clustering report.
[0017] Furthermore, each computing node of S100 is equipped with an independent processing unit, a memory unit and a storage unit, wherein:
[0018] The processing unit has data parallel processing capability and is used for computing tasks of local data subsets;
[0019] The memory unit can quickly respond to the data read and write requests of the processing unit, and provide temporary data storage and fast data exchange space for the data processing process;
[0020] The storage unit is used to store local data and intermediate data and result data in the clustering process for a long time, and has a data redundancy backup function to ensure data integrity.
[0021] Furthermore, the S300 allocates a subset of data points x i , i = 1, 2, ..., n, n is the total number of data points, and the Gaussian kernel function is used for kernel density estimation. When the data dimension d remains unchanged, the Gaussian kernel function is where u=u1,u2,…,u d ), for each data point x j , j = 1, 2, ..., n, calculate where x ik is the data point x i The value of the kth dimension of x, and k = 1, 2, ..., d, jk is the data point x j The value of the kth dimension of jk Substituting into the Gaussian kernel function, we get Represents data point x j For data point x i The contribution of the kernel density estimate to the data point x iThe kernel density estimate of By j The contribution of is summed and divided by n to obtain, that is,
[0022] Furthermore, after completing the kernel density estimation and the adaptive nearest neighbor search in S300, the process of computing the node to determine the local cluster center is as follows:
[0023] Step 1: Sort the data points from large to small according to the kernel density value, and select the first m data points from the data points with high kernel density values as the candidate points of the initial cluster center;
[0024] Step 2: For each candidate cluster center, calculate the weighted average position of the surrounding data points as its new coordinate estimate. The weighting coefficient is the neighbor weight w ij , that is, for the candidate cluster center c i , its new coordinates for: in is the candidate cluster center c i The set of neighboring points of x kj is the coordinate value of the neighbor point k in the jth dimension, is the candidate cluster center c i The neighbor weight between the nearest neighbor point k;
[0025] Step 3: Repeat step 2 until the estimated value of the cluster center coordinates in each dimension changes less than the threshold ∈ in n consecutive iterations. The cluster center coordinates determined at this time are the final local cluster center coordinates C ij , and record it.
[0026] Furthermore, the density estimation value ρ of the cluster in S300 is cluster , after the data points are assigned to each cluster, the local density of all data points in the cluster is statistically analyzed, and the average value of the local density of the data points in the cluster is calculated as the density estimate of the cluster, that is, Where n cluster is the number of data points in cluster C, ρ i It is the local density of data point i in cluster C. During the whole clustering process, the local density at the data point level can be naturally transitioned to the cluster density estimate at the cluster level.
[0027] For member data point identification ID ij When the computing node assigns data points to each cluster, for each data point assigned to the cluster, the unique identification information in the original data set is recorded. The identification information is the member data point identification ID ij .
[0028] Furthermore, each computing node in S400 compresses the clustering information when sending local clustering information. After receiving the compressed clustering information sent by each computing node, the central control node uses a global clustering collaborative algorithm to integrate and optimize the information. The global clustering collaborative algorithm calculates the cosine similarity between each local cluster. and Mahalanobis distance MD ij The fusion similarity metric sim ij , determine the cluster pair set S with high similarity, where the cosine similarity for: where x ik and x jk are the eigenvectors of the kth data point in clusters i and j, respectively. and is the characteristic mean vector of the corresponding cluster, MD ij is the Mahalanobis distance between clusters i and j, and γ is the similarity weight parameter;
[0029] When MD ij ≥M th When M th is the Mahalanobis distance threshold, sim ij =0;
[0030] When MD ij <M th hour,
[0031] Furthermore, for the cluster pair set S with high similarity, the cluster center C after merging is obtained by merging based on weighted average. new for: Where n i is the number of members in cluster i, C i It is the center of cluster i. The central control node feeds back the integrated global clustering information to each computing node to achieve effective coordination between local clustering and global clustering.
[0032] Furthermore, each computing node in S500 optimizes and adjusts the local clustering results according to the global clustering information fed back by the central control node, and recalculates the membership relationship between the local data points and the global clustering center by using a classification method based on the fusion of fuzzy logic and evidence theory, and constructs a fuzzy membership function μ ij for: where d ij is the distance from data point i to the center of cluster j, δ j is the fuzzy parameter of cluster j, and the credibility index EC is introduced ij , according to the feature weight w of the subset where the data point is locatedf , the distance d from the data point to the cluster center ij And the global density information of the cluster ρ global To assign the credibility of evidence EC ij For:EC ij =w f ×μ ij ×ρ global , where w f is the domain feature weight. Based on the fuzzy membership function and credibility, the final clustering attribution of the data point is determined, and the data point is assigned to the cluster with the maximum evidence credibility. According to the new clustering attribution result, the computing node fine-tunes the parameters of the local clustering and adjusts the bandwidth factor h of the local clustering. d , and fine-tune and optimize the weight parameters α and β of the nearest neighbor search.
[0033] Furthermore, the S500 clustering convergence condition is:
[0034] After each iteration, for the local clustering of each computing node, the ratio of the displacement of each cluster center to the distance from the cluster center of the previous iteration to its farthest data point is calculated, that is, the cluster center of the current iteration is The cluster center of the last iteration is Distance between clusters The farthest data point is x max , then the displacement ratio Calculate the average displacement ratio of all local clusters Where n is the total number of data points;
[0035] Set the cluster displacement threshold to T PCD , when PCD <T PCD When , it indicates that the movement of the local cluster center has become stable relative to its clustering range, then the clustering process is judged to have converged and the iteration is stopped.
[0036] Compared with the prior art, this data management method of a cross-domain data automatic clustering algorithm has the following beneficial effects:
[0037] 1. The present invention fully considers the characteristics of cross-domain data through distributed collaborative clustering, and adopts a clustering algorithm based on multi-domain kernel density estimation and adaptive neighbor search. It can dynamically adjust the bandwidth and neighbor weight of kernel density estimation according to the distribution characteristics of data in different fields, and more accurately identify the clustering structure in the data. At the same time, in the global clustering optimization process, a classification method based on the fusion of fuzzy logic and evidence theory is used to comprehensively consider the feature weights of the field where the data point is located, the distance from the data point to the cluster center, and the global density information of the cluster, to ensure that the clustering attribution of the data point is more reasonable and accurate, and effectively avoid the clustering errors and instability caused by the inability of traditional clustering methods to adapt to the complexity of cross-domain data, thereby providing a more reliable basis for subsequent data analysis and decision-making.
[0038] 2. The present invention adopts a distributed architecture to distribute data processing tasks to multiple computing nodes for parallel execution, which greatly shortens the clustering time of large-scale cross-domain data. The central control node performs reasonable data allocation and task scheduling according to the domain attributes of the data and the resource status of each computing node, thereby improving the overall processing efficiency. Moreover, the method of the present invention has good scalability and can easily add new computing nodes to adapt to the growing data scale and increasingly complex cross-domain data processing tasks. When the amount of data increases, there is no need to make large-scale adjustments to the overall architecture and algorithm. Only a simple expansion of the computing nodes is required to maintain efficient processing performance, which provides enterprises and institutions with a more flexible and efficient solution when processing massive cross-domain data.
[0039] Other advantages, objectives and features of the present invention will be set forth in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 An operation diagram of a data management method for an automatic clustering algorithm of cross-domain data;
[0042] Figure 2 This is a flowchart of the steps of user classification in Example 2. DETAILED DESCRIPTION
[0043] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0044] Embodiment 1
[0045] like Figure 1 As shown, this embodiment shows in detail the specific application process of a data management method of a cross-domain data automatic clustering algorithm. By building a distributed architecture, the cross-domain raw data is preprocessed and divided, and local clustering is performed on the computing nodes. The clustering algorithm based on multi-domain kernel density estimation and adaptive nearest neighbor search is used, combined with inter-node communication collaboration and global clustering optimization, to ultimately achieve efficient and accurate clustering of large-scale cross-domain data.
[0046] A central control node and multiple computing nodes are established. The central control node is responsible for the scheduling and management of the entire distributed clustering process. The central node has powerful computing power and large-capacity memory, which is used to store key information and metadata of global clustering and execute complex optimization algorithms to ensure the coordinated operation and efficient performance of the entire system. The computing nodes are each equipped with an independent processing unit, memory unit and storage unit. The processing unit has powerful data parallel processing capabilities and can process multiple data threads at the same time to accelerate the data calculation process. The memory unit can quickly respond to data read and write requests from the processing unit. It uses high-speed cache and low-latency memory technology to provide temporary data storage and fast data exchange space for data processing. The storage unit is used to store local data as well as intermediate data and result data in the clustering process for a long time, and has data redundancy backup function to ensure data integrity and prevent data loss through redundant storage. Then, the original cross-domain data is preprocessed in all aspects to facilitate subsequent clustering analysis. The preprocessed data is reasonably divided according to the domain attributes of the data, the data volume distribution, and the resource status of each computing node. Through the semantic analysis and statistical characteristics of the data, the data is divided into multiple subsets with certain independence and correlation, and these subsets are assigned to the corresponding computing nodes to ensure that the amount of data processed by each node is relatively balanced and that the data subsets have certain independence in the domain, providing a good data foundation for local clustering.
[0047] After each computing node receives the assigned data subset, the processing unit performs local clustering using a clustering algorithm based on multi-domain kernel density estimation and adaptive nearest neighbor search. For the kernel density estimation of multi-domain data, a domain adaptive bandwidth adjustment factor h is introduced. d h d =σ d ×(n d / N) -1 / ( 4+d ), where hd is the bandwidth of domain d, σ d is the standard deviation of the data in field d, n d is the number of data points in domain d, N is the total number of data points, d is the data dimension, and the adjustment factor h d Adjust the bandwidth of the kernel density estimation, and then perform kernel density estimation calculations on the data points to obtain the kernel density value of each data point to reflect the distribution density of the data, and for the data points x of the allocated subset i , i = 1, 2, ..., n, n is the total number of data points, and the Gaussian kernel function is used for kernel density estimation. When the data dimension d remains unchanged, the Gaussian kernel function is where u=(u1,u2,…,u d ), for each data point x j , j = 1, 2, ..., n, calculate where x ik is the data point x i The value of the kth dimension of x, and k = 1, 2, ..., d, jk is the data point x j The value of the kth dimension of jk Substituting into the Gaussian kernel function, we get Represents data point x j For data point x i The contribution of the kernel density estimate to the data point x i The kernel density estimate of By j The contribution of is summed and divided by n to obtain, that is, In this way, the bandwidth of the kernel density estimation can be dynamically adjusted according to the distribution characteristics of data in different fields to more accurately reflect the distribution density of the data. The adaptive neighbor search is defined based on the local density ρ of the data point. i and the neighbor weight function w of the relative distance ij To achieve this, assume that data point i is in the data subset and its local density ρ i The kernel density estimation above shows that For the distance d from data point j ij , and the neighboring point set N of data point i i , the weight parameters α and β are determined by the particle swarm optimization algorithm. After multiple iterations of optimization, according to the formula Calculate the neighbor weight w ijIn this way, it can better adapt to the complex distribution of cross-domain data and accurately identify clustering structures. After completing kernel density estimation and adaptive neighbor search, the computing node determines the local cluster center. First, it sorts the data points from large to small according to the kernel density value, and selects the first m data points from the data points with high kernel density values as the candidate points of the initial cluster center. For each candidate cluster center c i , calculate the weighted average position of the surrounding data points as its new coordinate estimate, and the weighting coefficient is the neighbor weight w ij , that is, for the candidate cluster center c i , its new coordinates for: Repeat this step until the estimated value of the cluster center coordinates in each dimension changes less than the threshold ∈ in n consecutive iterations. The cluster center coordinates determined at this time are the final local cluster center coordinates C ij , and record it. At the same time, during the local clustering process, the memory unit records the detailed information of the clustering in real time and uploads it to the central control node, including the cluster center coordinates C ij , Member data point identification ID ij , the cluster density estimate ρ cluster , for the cluster density estimate ρ cluster , after the data points are assigned to each cluster, the local density of all data points in the cluster is statistically analyzed, and the average value of the local density of the data points in the cluster is calculated as the density estimate of the cluster, that is, For member data point identification ID ij When the computing node assigns data points to each cluster, for each data point assigned to the cluster, the unique identification information in the original data set is recorded. The identification information is the member data point identification ID ij ,This information is stored in the local storage unit for subsequent information interaction and collaborative processing with other nodes.
[0048] When each computing node sends local clustering information, it compresses the clustering information. For numerical data such as cluster center coordinates, it predicts and encodes the data according to the numerical change law and distribution characteristics to reduce data redundancy. For discrete data such as cluster member identifiers, it compresses the data using the frequency distribution and continuous repetition characteristics to minimize the amount of data transmission while ensuring data integrity. After the central control node receives the compressed clustering information sent by each computing node, it uses the global clustering collaborative algorithm to integrate and optimize the information. By calculating the cosine similarity between each local cluster and Mahalanobis distance MD ij The fusion similarity metric sim ij , determine the cluster pair set S with high similarity, where the similarity x ik and x jk are the eigenvectors of the kth data point in clusters i and j, respectively. and is the characteristic mean vector of the corresponding cluster, MD ij is the Mahalanobis distance between clusters i and j, γ is the similarity weight parameter, when MD ij ≥M th When ij =0; when MD ij <M th hour, Extract the feature vectors x of the data points in the two clusters respectively ik and x jk , calculate its characteristic mean vector and And Mahalanobis distance MD ij , substitute into the formula to calculate the similarity metric sim ij , thereby determining the cluster pair set S with high similarity. For the cluster pair set S with high similarity, the combined cluster center C new for: Where n i is the number of members in cluster i, C i It is the center of cluster i. The central control node feeds back the integrated global clustering information to each computing node to achieve effective coordination between local clustering and global clustering.
[0049] Each computing node optimizes and adjusts the local clustering results based on the global clustering information fed back by the central control node, and recalculates the membership relationship between the local data points and the global clustering center using a classification method based on the fusion of fuzzy logic and evidence theory to construct a fuzzy membership function. where d ij is the distance from data point i to the center of cluster j, δ j is the fuzzy parameter of cluster j, and the credibility index EC is introduced ij , according to the feature weight w of the subset where the data point is located f , the distance d from the data point to the cluster center ij And the global density information of the cluster ρ global To assign the credibility of evidence EC ij =w f ×μ ij ×ρ global Based on the fuzzy membership function and credibility, the final clustering of the data point is determined, and the data point is assigned to the cluster with the maximum evidence credibility. According to the new clustering result, the computing node fine-tunes the parameters of the local clustering and adjusts the bandwidth factor h of the local clustering.d , and the weight parameters α and β of the nearest neighbor search are fine-tuned and optimized. For example, if it is found that the data points of a cluster are relatively scattered during the optimization process, the bandwidth adjustment factor h can be appropriately increased. d , recalculate the kernel density estimation and clustering process; if the neighbor relationship of a cluster is unreasonable, the weight parameters α and β of the neighbor search can be adjusted according to the distribution and membership relationship of the data points to further improve the accuracy and stability of the clustering. After each iteration, for the local clustering of each computing node, calculate the ratio of the displacement of each cluster center to the distance from the cluster center of the previous iteration to its farthest data point, that is, the cluster center of the current iteration is The cluster center of the last iteration is Distance between clusters The farthest data point is x max , then the displacement ratio Calculate the average displacement ratio of all local clusters Where n is the total number of data points. <T PCD When , it indicates that the movement of the local cluster center has become stable relative to its clustering range, the clustering process is judged to be converged, the iteration is stopped, and the final clustering operation is completed to ensure the stability and reliability of the clustering results.
[0050] After completing the final clustering operation, the central control node integrates and summarizes the global clustering results to generate a detailed and comprehensive clustering report. The clustering report covers a detailed list of members of each cluster, a description of the characteristics of the clusters, the hierarchical relationship between clusters, and the distribution of clusters in various fields. It presents the clustering structure and internal relationship of cross-domain data in an intuitive and clear way, which is convenient for users to analyze and make decisions. The clustering results are stored in distributed computing nodes to ensure high availability, reliability and scalability of the data. At the same time, an efficient data indexing and metadata management mechanism is established to facilitate users to quickly query and retrieve information about specific clusters or data points, improve data access efficiency, and encrypt and store clustering result data to ensure data security and privacy, prevent data leakage and illegal access, and ensure data security and integrity.
[0051] To sum up, this embodiment elaborates on the implementation process of the data management method of the cross-domain data automatic clustering algorithm. Through the distributed architecture, innovative clustering algorithm and collaborative optimization mechanism, it can give full play to its advantages, improve the accuracy and stability of clustering, and greatly improve the data processing efficiency through distributed computing and efficient communication collaboration. It has good scalability and provides strong technical support for the deep mining and application of cross-domain data.
[0052] Embodiment 2
[0053] like Figure 2As shown, this embodiment focuses on applying the data management method of the cross-domain data automatic clustering algorithm to the smart city management scenario, integrating traffic, environmental and energy data to improve urban operation efficiency. By building a distributed architecture, performing data preprocessing and partitioning, performing local clustering, achieving inter-node communication coordination, completing global clustering optimization and generating a clustering report, the actual application process and remarkable results of this method in complex multi-domain data processing are fully demonstrated, providing data-driven decision support for smart city construction.
[0054] First, enter the stage of building a distributed architecture (S100). In the smart city management project, a powerful central control node is set up, which is equipped with a high-end server-level processor with multi-core and high main frequency characteristics, and can quickly process the integration and strategy calculation of massive data; the large-capacity memory uses advanced memory expansion technology to ensure data high-speed caching and fast exchange; the high-speed storage device is based on the enterprise-level storage architecture, with high read and write speed and data redundancy protection mechanism, used to store global key information and core algorithm models. At the same time, multiple computing nodes are deployed in data centers in various regions of the city. The processing unit of each computing node is configured with an appropriate number of processor cores according to local data processing requirements, and adopts advanced parallel computing architecture to accelerate data calculation; the memory unit ensures fast reading and writing of data during processing; the storage unit uses a large-capacity hard disk array with hot-swap and data backup functions to store local original data and various types of data generated by the clustering process. These computing nodes are connected through a high-speed dedicated network to ensure low latency and high reliability of data transmission, forming an efficient distributed computing environment and laying the foundation for subsequent data processing.
[0055] Then, the data preprocessing and data division stage (S200) is entered, and data such as vehicle flow, vehicle speed, and traffic accident information are collected from the urban traffic system. The environmental monitoring system obtains air quality, noise level, and meteorological data. The energy management system collects cross-domain raw data such as power consumption and energy production data and aggregates them to the central control node. In the data cleaning stage, for traffic data, based on the normal distribution law of historical vehicle flow and vehicle speed and the statistical information of common traffic accident types, obviously abnormal vehicle speed or erroneous accident records (such as time logic confusion) are identified and eliminated; in environmental data, abnormal air quality or noise values are removed; energy data is filtered out of unreasonable energy data points based on the seasonal and periodic laws of energy consumption and the rated capacity range of power generation equipment. Then, standardization processing is performed. The central control node divides and distributes data based on the geographical area attributes of the data (such as traffic data by different urban areas, environmental data by monitoring site locations, and energy data by energy supply areas), time series characteristics (such as by hour, day, month, etc.) and the real-time load of each computing node. For example, traffic and environmental data of the same urban area and similar time periods can be distributed to computing nodes close to the area, while ensuring that the data volume and processing complexity of each node are relatively balanced, reducing data transmission overhead and processing delays.
[0056] Then, the local clustering stage (S300) is entered. After receiving the data subset, each computing node starts a clustering algorithm based on multi-domain kernel density estimation and adaptive neighbor search. In the kernel density estimation, taking the traffic and environment data subset as an example, for the traffic data, the standard deviation σ of the road section traffic flow data is calculated and counted. d and the number of data points n on this road section d , the total number of data points N, according to the formula h d =σ d ×(n d / N) -1 / (4+d) Calculate the bandwidth h of the traffic area d , environmental data such as air quality data, calculate the corresponding bandwidth according to its own statistical parameters, and use the Gaussian kernel function to calculate the kernel density. Under a fixed data dimension, for each data point x j , calculated according to the formula Substitute into the Gaussian kernel function Get the kernel density contribution between data points, and then find the data point x i The kernel density estimate f(x i ). During the adaptive neighbor search, the local density ρ of the traffic congestion hotspot is calculated i The distance d between the data points ij , perform multiple iterations to optimize the weight parameters α and β, and use the formula Calculate the neighbor weight w ij, when accurately identifying the data clustering structure and determining the local cluster center, first sort by kernel density value in descending order, and select the first m high kernel density data points as candidates. For example, in traffic congestion analysis, select high kernel density points in areas with complex changes in traffic volume and speed, and select candidate centers c i , according to the formula Calculate the weighted average position of the surrounding points as the new coordinate estimate, where is the candidate cluster center c i The set of neighboring points of x kj is the coordinate value of the neighbor point k in the jth dimension, is the candidate cluster center c i The neighbor weight between the nearest neighbor point k is repeated multiple times until the cluster center coordinates are in r consecutive times to determine the final local cluster center C ij ,At the same time, the cluster center coordinates, member data point identifiers (vehicle ID in traffic data, monitoring point number in environmental data) and cluster density estimates are recorded and stored in the local storage unit, and key information is uploaded to the central control node.
[0057] Next, the inter-node communication coordination phase (S400) is entered. Before the computing node sends the local clustering information, the central control node receives the compressed information and executes the global clustering coordination algorithm to calculate the cosine similarity between each local cluster. Mahalanobis distance MD ij The fusion similarity metric sim ij For example, in the comparison of traffic and energy clustering, the traffic congestion mode feature vector and the energy consumption peak time feature vector are extracted, the feature mean vector and Mahalanobis distance are calculated, and the formula is substituted where x ik and x jk are the eigenvectors of the kth data point in clusters i and j, respectively. and is the characteristic mean vector of the corresponding cluster, MD ij is the Mahalanobis distance between clusters i and j, γ is the similarity weight parameter, and the Mahalanobis distance threshold M th Judgement, when MD ij ≥M th When M th is the Mahalanobis distance threshold, sim ij =0; when MD ij <M th hour, Determine the set S of similar clusters, and for the clusters within the set S, use the weighted average formula Merge cluster centers, where n i is the number of members in cluster i, C iIt is the center of cluster i. The central control node then feeds back the optimized global clustering information to each computing node to promote local and global clustering collaboration.
[0058] Finally, the global cluster optimization stage (S500) is entered, and the computing nodes use the classification method based on the fusion of fuzzy logic and evidence theory to optimize the clustering of data points according to the feedback information. Constructing the fuzzy membership function For example, when judging the relationship between traffic data points and energy clusters, d ij is the distance from the data point to the energy cluster center, δ j is the fuzzy parameter of energy cluster j, according to the association weight w between traffic and energy data f , the distance from the data point to the cluster center and the global cluster density information, and calculate the evidence credibility EC ij =w f ×μ ij ×ρ global , and redistribute the data points accordingly. Fine-tune the local clustering parameters according to the new attribution results. If traffic congestion clusters are found to be dispersed, increase the kernel density estimation bandwidth h d ; If the neighbor relationship of energy consumption clustering is unreasonable, adjust the weight parameters α and β to improve clustering accuracy and stability. Repeat steps S300-S500 to calculate the displacement ratio of cluster centers Ratio to average displacement The current iteration cluster center is The cluster center of the last iteration is Distance between clusters The farthest data point is x max , n is the total number of data points, and the cluster displacement threshold is T PCD , when PCD <T PCD When , the clustering is determined to have converged and the iteration is stopped.
[0059] Finally, the clustering result storage and management stage (S600) is entered. After the clustering convergence is completed, the central control node integrates the global clustering results and generates a detailed clustering report. The report covers a detailed list of cluster members (such as specific data point identifiers of traffic congested sections, high-pollution environmental areas, and high-energy consumption areas), cluster feature descriptions (such as traffic congestion time and degree, environmental quality level, energy consumption pattern), inter-cluster correlation analysis (such as the spatiotemporal correlation between traffic congestion and energy consumption peak) and cross-domain data distribution insights (such as the differences in cluster distribution of traffic, environment, and energy data in different urban areas), providing city managers with comprehensive data insights and decision-making basis, and assisting in the formulation of traffic diversion, environmental governance, and energy allocation strategies.
[0060] In summary, this embodiment successfully applies the data management method of the cross-domain data automatic clustering algorithm to smart city management. Through close collaboration of various steps, complex multi-source data can be effectively processed. The distributed architecture ensures computing efficiency and scalability. Node collaboration and global optimization enhance the accuracy and stability of the results. The final clustering report provides strong support for the refined management of the city, demonstrating the potential of this method in actual complex scenarios.
[0061] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A data management method for cross-domain data automatic clustering algorithm, characterized in that: The specific steps of this method are: S100, build a distributed architecture: set up a central control node and multiple computing nodes. The central control node is responsible for information integration and strategy formulation, and the computing nodes are used to perform local data processing tasks and preliminary clustering; S200, data preprocessing and data partitioning: the central control node preprocesses the input cross-domain raw data, divides the data into multiple data subsets, and distributes the data subsets to the computing nodes; S300, local clustering: After each computing node receives the allocated data subset, the processing unit performs local clustering using an automatic clustering algorithm. During the clustering process, the node dynamically adjusts the clustering parameters according to the distribution characteristics of the local data to accurately identify the clustering structure in the data; When the automatic clustering algorithm performs local clustering, a clustering algorithm based on multi-domain kernel density estimation and adaptive nearest neighbor search is adopted, wherein: For the kernel density estimation of multi-domain data, we adopt the kernel density estimation based on multi-domain and introduce the domain adaptive bandwidth adjustment factor h d h d =σ d ×(n d / N) -1 / (4+d) , where h d is the bandwidth of domain d, σ d is the standard deviation of the data in field d, n d is the number of data points in domain d, N is the total number of data points, d is the data dimension, and the adjustment factor h d Adjust the bandwidth of the kernel density estimation, and then perform kernel density estimation calculation on the data points to obtain the kernel density value of each data point to reflect the distribution density of the data; The adaptive nearest neighbor search clustering algorithm is based on the local density ρ of the data points. i and the neighbor weight function w of the relative distance ij To achieve where w ij is the neighbor weight between data points i and j, ρ i is the local density of data point i and d ij is the distance between i and j, N i is the set of neighboring points of data point i, α and β are weight parameters used to balance the influence of local density and relative distance in neighbor judgment; During the local clustering process, the memory unit records the detailed information of the cluster in real time and uploads it to the central control node. At the same time, it records the key information of each cluster, including the coordinates of the cluster center C. ij , Member data point identification ID ij , the cluster density estimate ρ cluster , and store it in a local storage unit; S400, inter-node communication coordination: the central control node collects the local clustering information uploaded by each computing node, and integrates the information using a global clustering coordination algorithm to obtain global clustering information, and feeds back the integrated global clustering information to each computing node; S500, global clustering optimization: each computing node further optimizes and adjusts the local clustering result according to the global clustering information fed back by the central control node, repeats the steps of S300 and S400, and performs multiple iterations until the clustering convergence condition is met, determines that the clustering process converges, and stops iteration; S600, clustering result storage and management: After clustering convergence, the final data clustering is completed, and the central control node integrates and summarizes the global clustering results to generate a detailed and comprehensive clustering report.
2. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: Each computing node of S100 is equipped with an independent processing unit, memory unit and storage unit, wherein: The processing unit has data parallel processing capability and is used for computing tasks of local data subsets; The memory unit can quickly respond to the data read and write requests of the processing unit, and provide temporary data storage and fast data exchange space for the data processing process; The storage unit is used to store local data and intermediate data and result data in the clustering process for a long time, and has a data redundancy backup function to ensure data integrity.
3. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: The S300 allocates a subset of data points x i , i = 1, 2, ..., n, n is the total number of data points, and Gaussian kernel function is used for kernel density estimation. When the data dimension d remains unchanged, the Gaussian kernel function is where u=(u1,u2,…,u d ), for each data point x j , j = 1, 2, ..., n, calculate where x ik is the data point x i The value of the kth dimension of jk is the data point x j The value of the kth dimension of jk Substituting into the Gaussian kernel function, we get Represents data point x j For data point x i The contribution of the kernel density estimate to the data point x i The kernel density estimate of By j The contribution of is summed and divided by n to obtain, that is, 4. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: After completing the kernel density estimation and the adaptive nearest neighbor search in S300, the process of computing the node to determine the local cluster center is as follows: Step 1: Sort the data points from large to small according to the kernel density value, and select the first m data points from the data points with high kernel density values as the candidate points of the initial cluster center; Step 2: For each candidate cluster center, calculate the weighted average position of the surrounding data points as its new coordinate estimate. The weighting coefficient is the neighbor weight w ij , that is, for the candidate cluster center c i , its new coordinates for: in is the candidate cluster center c i The set of neighboring points of x kj is the coordinate value of the neighbor point k in the jth dimension, is the candidate cluster center c i The neighbor weight between the nearest neighbor point k; Step 3: Repeat step 2 until the estimated value of the cluster center coordinates in each dimension changes less than the threshold ∈ in n consecutive iterations. The cluster center coordinates determined at this time are the final local cluster center coordinates C ij , and record it.
5. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: The density estimation value ρ of the cluster in S300 cluster , after the data points are assigned to each cluster, the local density of all data points in the cluster is statistically analyzed, and the average value of the local density of the data points in the cluster is calculated as the density estimate of the cluster, that is, where n cluster is the number of data points in cluster C, ρ i It is the local density of data point i in cluster C. During the entire clustering process, the local density at the data point level can naturally transition to the cluster density estimate at the cluster level. For member data point identification ID ij When the computing node assigns data points to each cluster, for each data point assigned to the cluster, the unique identification information in the original data set is recorded. The identification information is the member data point identification ID ij .
6. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: In S400, each computing node compresses the clustering information when sending local clustering information. After receiving the compressed clustering information sent by each computing node, the central control node uses a global clustering collaborative algorithm to integrate and optimize the information. The global clustering collaborative algorithm calculates the cosine similarity between each local cluster. and Mahalanobis distance MD ij The fusion similarity metric sim ij , determine the cluster pair set S with high similarity, where the cosine similarity for: where x ik and x jk are the eigenvectors of the kth data point in clusters i and j, respectively. and is the characteristic mean vector of the corresponding cluster, MD ij is the Mahalanobis distance between clusters i and j, and γ is the similarity weight parameter; When MD ij ≥M th When M th is the Mahalanobis distance threshold, sim ij =0; When MD ij <M th hour, 7. The data management method of a cross-domain data automatic clustering algorithm according to claim 6 is characterized in that: For the cluster pair set S with high similarity, the cluster center C after merging is new for: where n i is the number of members in cluster i, C i It is the center of cluster i. The central control node feeds back the integrated global clustering information to each computing node to achieve effective coordination between local clustering and global clustering.
8. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: Each computing node in S500 optimizes and adjusts the local clustering results according to the global clustering information fed back by the central control node, recalculates the membership relationship between the local data points and the global clustering center by using a classification method based on the fusion of fuzzy logic and evidence theory, and constructs a fuzzy membership function μ ij for: where d ij is the distance from data point i to the center of cluster j, δ j is the fuzzy parameter of cluster j, and the credibility index EC is introduced ij , according to the feature weight w of the subset where the data point is located f , the distance d from the data point to the cluster center ij And the global density information of the cluster ρ global To assign the credibility of evidence EC ij For:EC ij =w f ×μ ij ×ρ global , where w f is the domain feature weight. Based on the fuzzy membership function and credibility, the final clustering attribution of the data point is determined, and the data point is assigned to the cluster with the maximum evidence credibility. According to the new clustering attribution result, the computing node fine-tunes the parameters of the local clustering and adjusts the bandwidth factor h of the local clustering. d , and fine-tune and optimize the weight parameters α and β of the nearest neighbor search.
9. The data management method of a cross-domain data automatic clustering algorithm according to claim 1 is characterized in that: The S500 clustering convergence condition is: After each iteration, for the local clustering of each computing node, the ratio of the displacement of each cluster center to the distance from the cluster center of the previous iteration to its farthest data point is calculated, that is, the cluster center of the current iteration is The cluster center of the last iteration is Distance between clusters The farthest data point is x max , then the displacement ratio Calculate the average displacement ratio of all local clusters Where n is the total number of data points; Set the cluster displacement threshold to T PCD , when PCD<T PCD When , it indicates that the movement of the local cluster center has become stable relative to its clustering range, then the clustering process is judged to have converged and the iteration is stopped.
Citation Information
Cited By
High-salt-tolerance rice breeding data management method and system based on clustering processing
CN121030284A
High-salt-tolerant rice breeding data management method and system based on clustering processing
CN121030284B
Multi-stage spatial-temporal clustering method and system based on fused mahalanobis distance
CN121051489A
Method for fitting ellipse with any adaptation error through plane coordinate point cloud
CN121071282A
A method for fitting an error ellipse of any tightness to a planar coordinate point cloud
CN121071282B