Stream data hierarchical clustering optimization method based on root insertion strategy
By introducing dynamic similarity measurement and time weighting mechanisms into the stream data hierarchical clustering optimization method, combined with local updates and structural optimization operations, the existing methods have solved the problems of clustering quality and computing efficiency in large-scale real-time data stream processing, and achieved efficient and accurate clustering effect.
Patent Information
- Application Number
- CN202510189365.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing stream data hierarchical clustering optimization method based on root insertion strategy has problems such as degradation in clustering quality, low computing efficiency, poor stability and difficult high-dimensional data processing when processing large-scale real-time data streams.
Dynamic similarity measurement and time weighting mechanism are adopted, combined with local update, merging and pruning operations, the structure of the cluster tree is dynamically adjusted, the depth and breadth of the tree are optimized, and the clustering accuracy and calculation complexity are balanced.
It improves the computing efficiency and accuracy in the process of stream data clustering, reduces the computing burden of global reconstruction, adapts to the dynamic changes of data flow in real time, maintains a low computational complexity, and improves clustering accuracy in high-dimensional data scenarios.
Smart Images

Figure CN120123797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an optimization method for hierarchical clustering of streaming data, specifically an optimization method for hierarchical clustering of streaming data based on a root insertion strategy. Background Art
[0002] Elaborate in detail the deficiencies and drawbacks existing in the current optimization method for hierarchical clustering of streaming data based on the root insertion strategy without paragraph breaks; the number of words should be more than 800.
[0003] Although the optimization method for hierarchical clustering of streaming data based on the root insertion strategy has significant advantages in processing large-scale real-time data streams, there are still some deficiencies and drawbacks in practical applications. First of all, when the root insertion strategy dynamically adjusts the clustering tree structure, although it can reduce the computational overhead of global reconstruction, when the similarity between data points is relatively complex or the density difference is large, it may still lead to a decline in clustering quality. In this case, local updates and root insertions may not be able to promptly reflect the details in the data stream, resulting in a large deformation of the clustering tree in the local area and even affecting the overall clustering effect. Secondly, the root insertion strategy relies on the rapid calculation and real-time update of the similarity between data points. However, as the data stream continuously increases, the time complexity of calculating the similarity will also gradually increase. Although dynamic similarity metrics and weighting mechanisms are adopted to adjust the calculation of similarity, these mechanisms themselves may gradually lose their efficiency as the data volume increases. Especially in the scenario of large-scale data streams, how to accurately and efficiently calculate the similarity has become an urgent problem to be solved. In addition, the influence of the time decay factor on the data stream also needs to be carefully handled. Too strong a time decay may ignore the historical information of the data, while too weak a time decay may cause the clustering tree to overly rely on recent data, resulting in too strong a historical dependence on the data and affecting the mining of long-term trends.
[0004] Secondly, although local merging and pruning operations can optimize the tree structure and reduce redundant nodes and computational complexity, in some cases, excessive merging or pruning operations may lead to the loss of important clustering information, especially when the diversity and complexity of the data stream are high. This situation may result in the over-simplification of the tree, making the clustering results less refined and reducing the effectiveness and stability of clustering. Moreover, although the introduction of the local update mechanism can reduce the computational burden, its performance under extreme data distributions still needs to be further verified. Especially when sudden changes occur in the data stream, how to balance the relationship between local updates and global optimization to ensure the stability and accuracy of the clustering tree is a major challenge for current methods. Finally, although the depth and breadth adjustment mechanism of the clustering tree can optimize the structure of the clustering tree in real time, in some application scenarios, the dynamic adjustment of the depth and breadth of the tree may introduce unnecessary complexity, especially in scenarios with extremely high real-time requirements. How to find a suitable balance between complexity and accuracy and avoid over-adjustment during dynamic adjustment has also become a bottleneck in the application of the root insertion strategy method in large-scale streaming data. In addition, when dealing with high-dimensional data streams, existing methods may encounter the curse of dimensionality problem, resulting in a decline in clustering performance. Although a weighting mechanism is introduced to adjust the influence of each dimension, in the high-dimensional data space, the complexity and interpretability of distance calculation itself may make the construction process of the clustering tree unstable.
[0005] Therefore, in high-dimensional streaming data clustering, how to improve clustering accuracy through more effective dimensionality reduction or feature selection mechanisms remains an urgent problem to be solved. Generally speaking, although the hierarchical clustering optimization method for streaming data based on the root insertion strategy has shown good performance in multiple fields, it still faces challenges in terms of computational efficiency, clustering quality, stability, and high-dimensional data processing. Especially when dealing with large-scale and complex streaming data, existing methods still need further improvement and optimization to cope with a wider range of application scenarios. Summary of the Invention
[0006] The object of the present invention is to provide an optimized hierarchical clustering method for streaming data based on the root insertion strategy, so as to solve some of the drawbacks and deficiencies pointed out in the background technology.
[0007] The technical solutions adopted by the present invention to solve its above technical problems include the following steps:
[0008] S1. Dynamic input and incremental preprocessing of streaming data:
[0009] S1.1. Receive streaming data from the data source in real time;
[0010] S1.2. According to the set time window or data volume window;
[0011] S1.3. Standardize the data in the window; extract key features including distance, density, and time information;
[0012] S2. Hierarchical Clustering Tree Construction and Root Insertion Strategy:
[0013] S2.1. Construct a hierarchical clustering tree by presetting streaming data and clustering parameters including a clustering threshold and a similarity metric;
[0014] S2.2. Judge the similarity between the new data and the root node of the current clustering tree; perform local updates; and create a new root node;
[0015] S2.3. The new data points gradually expand the clustering tree structure and affect the expansion of the tree and the adjustment of the original subtrees;
[0016] S3. Dynamic Update and Local Optimization of the Clustering Tree:
[0017] S3.1. Calculate the similarity between the new data and the existing subtrees, adjust the subtree structure according to the similarity, and incorporate it into the existing subtrees or create new subtree branches;
[0018] S3.2. Adjust the shape of the tree through local merging or pruning operations; optimize the algorithm and adjust the structure;
[0019] S3.3. Maintain the balance between the accuracy and complexity of hierarchical clustering by dynamically adjusting the depth and breadth of the tree;
[0020] S4. Clustering Merging and Splitting Decision:
[0021] S4.1. Monitor the distances and similarities between adjacent subtrees in the clustering tree in real time;
[0022] S4.2. Trigger splitting operations according to the preset splitting rules;
[0023] S4.3. Evaluate the clustering tree structure and optimize the tree depth, width, and similarity metric.
[0024] Furthermore, the method for hierarchical clustering tree construction and root insertion strategy includes:
[0025] Introduce a dynamic similarity metric and a weighted mechanism based on the timestamp of streaming data to adjust the similarity calculation between data points; controlled by the w k and p k parameters; introduce a time weighting factor and adopt a time decay design to reduce the impact of data points with a large time span on the clustering tree;
[0026] The similarity metric formula is as follows:
[0027]
[0028] Where:
[0029] S ij is the similarity metric value between data points i and j; and respectively represent the feature values of data points i and j in the k-th dimension; w k is the weighting coefficient of the k-th dimensional feature; p k is the distance metric exponent of the k-th dimension; p is the scale factor of the overall distance metric; α represents the time weighting factor; t i and t j are the timestamps of data points i and j; is the time decay term, which increases as the time difference |t i t j | increases.
[0030] Furthermore, the method for constructing the hierarchical clustering tree and the root insertion strategy includes:
[0031] Adopt a dynamic root insertion strategy to insert new data points into the existing tree structure; evaluate the similarity between the data point and the root node of the existing cluster and merge with the existing node; by comparing the similarity between the data point and the root of the tree; reduce the computational burden of global reconstruction;
[0032] The calculation formula is as follows:
[0033]
[0034] Where:
[0035] D insert is the insertion metric between data point i and the root node of the cluster, representing the similarity between the data point and the root of the tree; and x rootk respectively represent the feature values of data point i and the root node of the tree in the k-th dimension; λ and β are constants that control the similarity decay; λ controls the amplification coefficient of the distance, and β controls the decay rate of the similarity by the time interval; t i and t root are the timestamps of data point i and the root node; is the time decay term.
[0036] Furthermore, the method for constructing the hierarchical clustering tree and the root insertion strategy includes:
[0037] A local update mechanism is constructed, and each node of the clustering tree is adjusted according to the change of the data point; the tree structure is kept in a balanced state; the idea of calculus is used, and the structure of the clustering tree is adjusted by calculating the relative position change between nodes;
[0038] The optimization formula for local update is:
[0039]
[0040] Wherein:
[0041] ΔT is the change amount of the clustering tree structure during the local update process; dist(x i , x k ) is the distance metric between data point i and node j; x i and x k represent the feature vectors of data point i and node k respectively; is the partial calculation of the local change rate of node k on all data points; Δx k represents a small change in the position of node k, and the local optimization process will adjust the tree structure based on the change.
[0042] Furthermore, the dynamic update and local optimization of the clustering tree include:
[0043] Through local merging and pruning operations; triggering the merging operation by dynamically evaluating the node similarity, merging multiple similar nodes; removing redundant nodes through pruning operations to reduce the width of the tree; combining time factors to measure the relevance of data points;
[0044] The formula is as follows:
[0045]
[0046] Wherein:
[0047] ΔC represents the change amount between nodes i and j; x i and x j represent the feature vectors of nodes i and j respectively; w k is the weighting coefficient of the k-th dimension; w k is used to adjust the weight of each feature; p k is the distance metric exponent of the k-th dimension; t i and t j represent the timestamps of nodes i and j respectively; represents the time decay factor.
[0048] Furthermore, the dynamic update and local optimization of the clustering tree include:
[0049] Construct a dynamic adjustment mechanism to adjust the structure based on real-time data; and increase or decrease the depth of the clustering tree by real-time analyzing the similarity distribution and density change of the data;
[0050] The formula is as follows:
[0051]
[0052] Wherein:
[0053] D adjustRepresents the adjustment amount of the depth or breadth of the tree; x i Represents the feature vector of data point i; x root Represents the feature vector of the root node of the clustering tree; p is the exponent for adjusting the depth or breadth; λ and α are coefficients for adjusting the time decay and similarity weight; λ adjusts the overall weight, and α controls the speed of time decay; t i and t root Represents the timestamps of data point i and the root node; The time decay factor is used to calculate the time relationship between the data point and the root node when adjusting the structure of the tree.
[0054] Furthermore, the dynamic update and local optimization of the clustering tree include:
[0055] Dynamically adjust the tree structure, combine local merging and pruning operations, balance accuracy and complexity in real time, adjust the computational amount according to the characteristics of the current data, maintain clustering accuracy, and reduce computational overhead at the same time; optimize the computational burden of global reconstruction caused by incremental update and local optimization mechanisms;
[0056]
[0057] Among them:
[0058] T balance Represents the balance amount between accuracy and complexity; dist(x i ,x cluster ) represents the distance metric between data point i and the clustering center x cluster ; γ and β are respectively the coefficients for adjusting the balance of accuracy and complexity; γ controls the accuracy, and β is to control the impact of time decay on accuracy; t i and t cluster Represents the timestamps of data point i and the clustering center, which are used to quantify the time difference between the data point and the clustering center.
[0059] Describe the beneficial effects of the optimized method for hierarchical clustering of streaming data based on the root insertion strategy of the present invention:
[0060] The optimized method for selecting and arranging garden landscape plants of the present invention has the following beneficial effects: By dynamically adjusting the structure of the clustering tree, the computational efficiency and accuracy in the clustering process of streaming data are effectively improved. Compared with the traditional hierarchical clustering method, the root insertion strategy of the present invention avoids the high computational burden of global reconstruction, and only through local updates, merging and pruning operations, the tree structure is refined. This method not only reduces unnecessary computational overhead, but also can adapt to the dynamic changes of the data stream in real time, ensuring that the clustering tree always maintains a low computational complexity while ensuring accuracy.
[0061] By introducing a dynamic similarity measure and a time-weighting mechanism, the system can automatically adjust the depth and breadth of the clustering tree according to the similarity between data points and time factors. This adaptive adjustment can flexibly optimize the tree structure during the continuous inflow of data points, avoiding the limitations of traditional static methods in dynamic data scenarios. In addition, the introduction of local merging and pruning operations further optimizes the morphology of the clustering tree, enabling redundant nodes to be removed in a timely manner, reducing the width of the tree, and ensuring the efficiency of clustering calculations.
[0062] Overall, the present invention can achieve a good balance between the real-time performance and accuracy of streaming data clustering, avoid overly complex global calculations, and can flexibly cope with the rapid changes in data characteristics. It has broad application prospects and is particularly suitable for large-scale real-time data stream analysis and dynamic clustering scenarios. Brief Description of the Drawings
[0063] Figure 1 It is a flowchart of the optimized method for hierarchical clustering of streaming data based on the root insertion strategy of the present invention.
[0064] Figure 2 It is a flowchart of the construction of the hierarchical clustering tree and the root insertion strategy of the present invention.
[0065] Figure 3 It is a flowchart of the dynamic update and local optimization of the clustering tree of the present invention. Detailed Embodiment
[0066] Combined with the Figure 1 process, the optimized method for hierarchical clustering of streaming data based on the root insertion strategy includes the following steps:
[0067] S1. Dynamic input and incremental preprocessing of streaming data: In a streaming data environment, data is input into the system in real time and continuously. Therefore, the first step is to efficiently receive the streaming data source. At this time, the system will set a time window or a data volume window, and decide when to receive and process data according to the size or time limit of this window. When new data arrives, these data need to be standardized first. Standardization is to make the features of the data have a unified dimension, eliminating the influence caused by different units or scales of each data source, so as to ensure the comparability of data in subsequent processing. Then, the system extracts key features from these data, and these features include but are not limited to distance, density, and time information. Among them, distance refers to the relative position between data points, density reflects the distribution of data, and time information helps the system understand the dynamic changes of the data stream. In this way, the system can make full use of the characteristics of each data point to prepare for the subsequent clustering process while ensuring the efficiency of data processing.
[0068] S2. Hierarchical Clustering Tree Construction and Root Insertion Strategy: Once the streaming data enters the system and undergoes normalization and feature extraction, the next task is to construct a hierarchical clustering tree. At this stage, by setting clustering parameters including the clustering threshold, similarity measure, etc., the system can construct a preliminary hierarchical clustering tree according to the preset rules. The clustering threshold is used to define the critical value of similarity between data points. Only when the similarity between a new data point and the data points in the existing cluster reaches a certain standard will it be classified into the same cluster. The similarity measure is used to calculate the similarity between data points. Common measurement methods include Euclidean distance, Manhattan distance, etc. These parameters determine the structure of the clustering tree and its growth pattern. Next, the system determines the similarity between the new data and the root node of the current clustering tree and decides whether to directly add the new data to the root node by evaluating this similarity. If the similarity between the new data and the root node is high enough, the system will perform a local update, that is, insert the new data into the existing structure instead of reconstructing the entire tree. If the similarity between the new data and the root node is not high, a new root node will be created and the root insertion strategy will be initiated, that is, the tree structure will be extended by introducing the new root node into the clustering tree. In this way, the clustering tree can be dynamically extended, and the addition of new data will affect the adjustment of the original tree structure.
[0069] S3. Dynamic Update and Local Optimization of the Clustering Tree: During the construction and update of the hierarchical clustering tree, new data points continuously flow in. Whenever new data is added, the system needs to calculate the similarity between the new data and the existing subtrees and determine whether to classify the new data into the existing subtrees or create a new subtree branch based on this similarity. This decision is achieved by dynamically adjusting the clustering tree structure. If the similarity between the new data point and a certain subtree is high, it will be classified into that subtree; otherwise, if the similarity is low, the system will create a new subtree branch for this data point. To further optimize the tree structure, the system will also perform local merging or pruning operations to adjust the shape of the tree. In local merging, when the similarity between nodes within a subtree is too high, the system will merge multiple similar nodes to reduce the depth and complexity of the tree. The pruning operation removes certain nodes when their contribution to the overall clustering is small to reduce the width and complexity of the tree. In addition, the system will also optimize the algorithm according to the dynamic changes of the data. By adjusting the depth and breadth of the tree, the clustering tree can adaptively balance accuracy and computational complexity, avoiding the tree structure from being too long or too complex.
[0070] S4. Clustering Merging and Splitting Decision: As the clustering tree continues to expand and adjust, the system needs to monitor the distances and similarities between adjacent subtrees in the clustering tree in real time. These monitored data are crucial for subsequent clustering optimization. When the similarity or distance between certain subtrees reaches the preset splitting rules, the system will trigger a splitting operation to split these subtrees into multiple smaller subtrees to adapt to the new data distribution. Through the splitting operation, the clustering tree can more flexibly respond to the dispersion or aggregation trend of data points, maintain clustering accuracy, and avoid the phenomenon of over-aggregation. The system will also regularly evaluate the structure of the clustering tree, and optimize the structure by calculating the depth, width of the tree, and the similarities between each node. For example, if the depth of certain subtrees is too large and the computational complexity is high, the system will optimize the query efficiency and clustering accuracy by adjusting the structure of the tree. This series of dynamic adjustment and optimization operations enable the clustering tree to always maintain high efficiency and accuracy when facing continuously changing streaming data.
[0071] Example 1:
[0072] Combined with the process shown in the attached Figure 2 For a traffic monitoring system based on streaming data under development, this system is used to analyze urban traffic conditions in real time and dynamically identify and predict traffic congestion areas through the streaming data hierarchical clustering optimization method. The system obtains data such as vehicle positions, speeds, and road conditions from multiple traffic sensors. The data is updated every minute, and sensors from different geographical locations continuously transmit data streams in real time. To improve the real-time performance and computational efficiency of the system, a streaming data hierarchical clustering optimization method based on the root insertion strategy is adopted, combined with a dynamic similarity metric and a timestamp-based weighting mechanism. This method will help to timely identify and adjust the clustering structure of traffic flows for accurate traffic prediction and real-time decision-making.
[0073] Suppose the system receives data from 5 different traffic sensors every minute, and each sensor provides the following information:
[0074] Vehicle position (longitude, latitude)
[0075] Vehicle speed (km / h)
[0076] Traffic status (e.g., smooth, congested, slow moving, etc.)
[0077] The data comes from different regions, so the data characteristics provided by each sensor may vary. The timestamp of each data point is matched in real time with the received data stream, and the timestamp indicates the collection time of the data.
[0078] Suppose at a certain point in time, the system receives the following vehicle information from Sensor 1:
[0079] Vehicle position (longitude: 120.5, latitude: 30.2)
[0080] Vehicle speed (60 km / h)
[0081] Traffic condition (smooth)
[0082] Timestamp (t 1 = 10:00)
[0083] Then, the system receives data streams from other sensors with timestamps t 2 = 10:01, t 3 = 10:02, and so on. Whenever new data arrives, the system processes the data according to the dynamic input and incremental preprocessing method of the stream data and constructs and expands the hierarchical clustering tree using the root insertion strategy.
[0084] After receiving the data, the data is first normalized to ensure that the data from each sensor can be compared on the same scale. For example, for each vehicle speed, position, and traffic condition, the system converts them into a standard normal distribution to eliminate differences in units and dimensions, enabling effective clustering of data from different sensors.
[0085] Next, the system extracts key features from each data point, including:
[0086] Vehicle speed (x1)
[0087] Longitude (x2)
[0088] Latitude (x3)
[0089] Traffic condition (x4)
[0090] Dynamic similarity metric and time-weighting mechanism
[0091] To measure the similarity between new data and existing data, a dynamic similarity metric formula is introduced:
[0092]
[0093] Where:
[0094] S ij represents the similarity metric between data points i and j.
[0095] and are the feature values of data points i and j in the k-th dimension (such as vehicle speed, longitude, latitude, traffic condition, etc.), respectively.
[0096] w k is the weighting coefficient of the k-th dimension feature, indicating the importance of different features in the similarity calculation.
[0097] pk is the distance metric exponent of the k-th dimension, controlling the metric methods in different dimensions.
[0098] p is the scale factor of the overall distance metric, usually taking the value of 2 for the Euclidean distance metric.
[0099] α is the time weighting factor, controlling the speed of time decay.
[0100] t i and t j are the timestamps of data points i and j, measuring the time difference between data points.
[0101] is the time decay term. As the time difference increases, the decay factor gradually decreases, making the influence of older data points on the current clustering tree gradually weaken.
[0102] Suppose the following two pieces of data have been received:
[0103] Data point 1: Vehicle position (120.5, 30.2), vehicle speed (60 km / h), traffic condition (smooth), timestamp t1 = 10:00
[0104] Data point 2: Vehicle position (120.51, 30.21), vehicle speed (62 km / h), traffic condition (smooth), timestamp t2 = 10:05
[0105] First, calculate the similarity of these two data points in the feature dimension. Suppose the following parameters are set:
[0106] w 1 = 0.5 (weighting coefficient of vehicle speed)
[0107] w 2 = 0.3 (weighting coefficient of position)
[0108] w 3 = 0.2 (weighting coefficient of traffic condition)
[0109] p 1 = 2 (distance metric exponent of vehicle speed)
[0110] p 2 = 1 (distance metric exponent of position)
[0111] p 3 = 1 (distance metric exponent of traffic condition)
[0112] p = 2 (scale factor of the overall distance metric)
[0113] α = 0.1 (time weighting factor)
[0114] First, calculate the similarity of vehicle speed, location, and traffic status:
[0115] S speed =|6062| 2 =4
[0116] S location =|120.5120.51| 2 +|30.230.21| 2 =0.0001+0.0001=0.0002
[0117] S state =0 (because the traffic status is the same)
[0118] Next, weight the similarity of each dimension by the weighting factor:
[0119] S weighted =0.5·4 + 0.3·0.0002 + 0.2·0=2 + 0.00006=2.00006
[0120] Then, apply the overall scale factor p = 2 to calculate the total distance metric:
[0121] S total =(2.00006) 1 / 2 ≈1.414
[0122] Finally, calculate the time decay factor:
[0123]
[0124] Therefore, the total similarity is:
[0125] S ij =1.414·0.6065≈0.857
[0126] This similarity value indicates that data point 1 and data point 2 are similar, and due to the time difference, the similarity has decayed.
[0127] Through the above similarity calculation, the system can determine whether a new data point should be inserted as the root node or expand the existing clustering tree. If the new data has a high similarity with the root node of the current tree, the system will perform a local update and insert the new data point into the existing subtree; if the similarity is low, a new root node will be created, thus initiating the root insertion strategy to expand the clustering tree structure.
[0128] In the traffic monitoring system of the foregoing embodiment, the core objective of the system is to analyze and predict traffic conditions in real time. To this end, the system will use a flow data hierarchical clustering optimization method based on a root insertion strategy to more precisely monitor and predict traffic flow by efficiently expanding the existing clustering tree structure. In this example, the system will dynamically receive vehicle data from different traffic sensors and dynamically expand and optimize the clustering tree structure through the root insertion strategy, reducing the computational burden of global reconstruction.
[0129] Suppose data streams are obtained from five traffic sensors in the city. The data provided by each sensor includes:
[0130] Vehicle position (longitude, latitude)
[0131] Vehicle speed (km / h)
[0132] Traffic status (e.g., smooth, slow, congested, etc.)
[0133] Timestamp (e.g., t 1 = 10:00, t 2 = 10:01, t 3 = 10:02, etc.)
[0134] At a certain moment, the system receives the following information from Sensor 1:
[0135] Vehicle position (longitude: 120.5, latitude: 30.2)
[0136] Vehicle speed (60 km / h)
[0137] Traffic status (smooth)
[0138] Timestamp (t 1 = 10:00)
[0139] Suppose the root node of the current clustering tree is the following data:
[0140] Vehicle position (longitude: 120.4, latitude: 30.15)
[0141] Vehicle speed (58 km / h)
[0142] Traffic status (smooth)
[0143] Timestamp (t 0 = 09:50)
[0144] As new data arrives, the system uses a dynamic root insertion strategy to determine whether to insert the new data point into the existing clustering tree or expand the clustering tree with it as a new root node.
[0145] According to the set root insertion strategy, the system will calculate the insertion metric between the new data point and the root node of the current clustering tree to measure the similarity between the new data point and the root node of the tree. The insertion metric is calculated by the following formula:
[0146]
[0147] Where:
[0148] D insert represents the insertion metric between data point i and the root node of the tree, measuring the similarity between the new data and the root node of the tree.
[0149] and represent the feature values of data point i and the root node of the tree in the k-th dimension (e.g., vehicle speed, position, etc.) respectively.
[0150] λ and β are constants that control the attenuation of similarity.
[0151] λ controls the amplification factor of the distance, and its general value range is [0, 5].
[0152] β controls the influence of the time interval on the similarity attenuation rate, and its general value range is [0, 1].
[0153] t i and t root are the timestamps of data point i and the root node of the tree.
[0154] is the time attenuation term. As the time difference increases, the similarity will decrease.
[0155] Assume that the following parameters are applied to the calculation:
[0156] λ = 3
[0157] β = 0.2
[0158] Data point i (from sensor 1):
[0159] Vehicle speed (60 km / h)
[0160] Position (longitude: 120.5, latitude: 30.2)
[0161] Traffic condition (smooth)
[0162] Timestamp (t 1 = 10:00)
[0163] Data of the root node of the clustering tree:
[0164] Vehicle speed (58 km / h)
[0165] Position (longitude: 120.4, latitude: 30.15)
[0166] Traffic condition (smooth)
[0167] Timestamp (t 0 = 09:50)
[0168] First, calculate the differences among vehicle speed, position, and traffic condition:
[0169]
[0170] The traffic conditions are the same, so there is no difference.
[0171] Next, calculate the sum of squared Euclidean distances:
[0172]
[0173] Then, calculate the time decay term:
[0174]
[0175] Finally, calculate the insertion metric:
[0176]
[0177] According to the calculation result, the insertion metric D insert = 2.85 indicates a high similarity between the new data point and the current root node of the tree. Therefore, the system decides to insert the new data point under the root node of the existing clustering tree instead of creating a new root node. The insertion operation is completed through local update, and the system only updates the root node of the tree, avoiding global reconstruction and thus reducing the computational complexity.
[0178] Through the dynamic root insertion strategy, the system effectively reduces the computational overhead caused by global reconstruction. Each time new data flows in, the system only calculates the similarity between the data point and the root node of the current clustering tree, thus achieving incremental update. This method adapts to the rapid changes of streaming data and ensures the balance between the accuracy of the clustering tree and the computational complexity.
[0179] In the context of the traffic monitoring system of this embodiment, considering that traffic flow data is constantly changing in real time, traditional clustering methods often cannot reflect these changes in a timely manner. Therefore, a hierarchical clustering optimization method for streaming data based on the root insertion strategy is adopted, combined with a local update mechanism to ensure that the structure of the clustering tree is always balanced while optimizing the computational efficiency. In this process, each node makes local adjustments according to the changes of the data points after receiving new data streams, rather than reconstructing the entire tree, thus greatly reducing the computational overhead.
[0180] In a traffic monitoring system, consider a system for monitoring highway traffic flow. The system continuously receives data streams from different locations through on-vehicle sensors. Each piece of data includes information such as the current position, speed, and traffic status of the vehicle. The system analyzes this data in real-time, clustering similar vehicle states to infer future traffic trends.
[0181] Suppose the current traffic monitoring system has already established a preliminary clustering tree based on historical data. The root node represents the normal traffic state of a certain highway section, and other branches represent different traffic flow situations (e.g., smooth, slow-moving, congested, etc.). As new vehicle data flows in, the system needs to continuously adjust the tree structure to maintain clustering accuracy and avoid global reconstruction. At this time, the local update mechanism comes in handy, optimizing the structure of the clustering tree by calculating the relative position changes between tree nodes.
[0182] In this example, the latest data sent by vehicle sensor 1 (data point x 1 ) is as follows:
[0183] Location: longitude 120.5, latitude 30.2
[0184] Speed: 55 km / h
[0185] Traffic status: slow-moving
[0186] Timestamp: t 1 = 10:10
[0187] At this time, the traffic state represented by the root node of the clustering tree is smooth, with a location of longitude 120.4, latitude 30.15, a speed of 58 km / h, and a timestamp t 0 = 10:00. It is necessary to evaluate the distance between data point x 1 and the root node of the tree, and adjust the tree structure according to its change.
[0188] By calculating the relative distance of the vehicle, the following formula can be used to calculate the change amount of the clustering tree structure during local update:
[0189]
[0190] Where:
[0191] ΔT is the change amount of the clustering tree structure during the local update process, measuring the degree of adjustment of the tree node structure.
[0192] dist(x i ,x k ) is the distance metric between data point i and node k, and the Euclidean distance or other distance metric methods suitable for traffic data can be selected.
[0193] x i and xk They respectively represent the feature vectors of data point i and node k, and the feature vectors may include information such as position, velocity, time, etc.
[0194] To calculate the local change rate of node k over all data points, that is, the relative position change of this node with respect to all data points.
[0195] Δx k It represents the small change in the position of node k, reflecting the adjustment of the node position by the local optimization process.
[0196] First, calculate the distance between data point x 1 (data of sensor 1) and the root node of the tree. Use the Euclidean distance formula:
[0197]
[0198] For the root node of the tree, calculate the local change rate of all data points relative to the root node. Assume that the root node of the current clustering tree represents a smooth traffic state, and data point x 1 (data of sensor 1) represents a slow-moving state, then the local change rate can be calculated by the following formula:
[0199]
[0200] According to calculus theory, the change rate can be estimated by the relative change between the data point and the root node of the tree, and is usually calculated by the backpropagation algorithm. Set a suitable change rate to 0.2.
[0201] Set the small change amount Δx of the node k =0.05, indicating that the position of the root node has changed slightly.
[0202] Substitute the above data into the formula to obtain the change amount ΔT of the tree structure in the local update process:
[0203] ΔT=0.2·0.05=0.01
[0204] This means that through local update, the structure of the clustering tree will change slightly, and the adjusted tree structure is more in line with the actual situation of the current streaming data.
[0205] After local update, the position of the root node of the tree will be fine-tuned to reflect the new traffic flow state. Since the structure of the clustering tree has been fine-tuned according to the changes of the new data points, the system avoids global reconstruction, thus improving the calculation efficiency.
[0206] After adopting the local update mechanism, the system can reflect the slightest changes in traffic conditions in real time without having to recalculate the entire tree structure. By continuously adjusting the relative positions between nodes, the system can maintain the balance of the clustering tree, ensuring the accuracy and real-time nature of the clustering results while reducing the computational complexity. This mechanism is particularly suitable for processing streaming data, enabling the system to operate efficiently in a dynamically changing environment.
[0207] Embodiment 2:
[0208] Combined with the attached Figure 3 , in an intelligent agricultural monitoring system, data from multiple field sensors is collected in real time to monitor the growth status of crops and environmental conditions. Consider a system for monitoring the greenhouse environment, where sensors regularly collect information such as temperature, humidity, and soil moisture. The data flowing into the system is constantly changing, and hierarchical clustering processing of this real-time data is required to help the farm administrator understand the status of different areas in the greenhouse in real time and thus make appropriate decisions.
[0209] In this scenario, the hierarchical clustering optimization method for streaming data based on the root insertion strategy is particularly important. The system needs to dynamically adjust the clustering tree structure according to characteristics such as time and environment in a large amount of continuously changing data streams. To achieve efficient data management, local merging and pruning operations are used to optimize the tree's form, reduce redundant nodes, and at the same time retain the accuracy of the data. Specifically, the dynamic evaluation of data point similarity can trigger the merging operation, while the pruning operation can effectively reduce the width of the tree, thereby improving the computational efficiency of the algorithm.
[0210] Suppose the sensor data in the current greenhouse flows into the system. The first piece of data comes from Sensor 1, with a timestamp of t 1 = 10:00, and the characteristic data it collects is:
[0211] Temperature: 25°C
[0212] Humidity: 65%
[0213] Soil moisture: 30%
[0214] The latest data timestamp of Sensor 2 is t 2 = 10:05, and the characteristic data it collects is:
[0215] Temperature: 24.5°C
[0216] Humidity: 66%
[0217] Soil moisture: 29%
[0218] The latest data timestamp of Sensor 3 is t 3 = 10:10, and the characteristic data it collects is:
[0219] Temperature: 26°C
[0220] Humidity: 64%
[0221] Soil humidity: 32%
[0222] The system gradually constructs a clustering tree through the root insertion strategy of the clustering tree. First, the data of sensor 1 is used as the root node to start building the tree structure. At this time, the clustering threshold has been set. Assume that the weight coefficients w 1 = 0.7 and w 2 = 0.3 for the two features of temperature and humidity. Whenever new data flows in, the system determines whether local merging or pruning is needed based on the similarity between the new data and the existing tree nodes. The following will show how to optimize the tree structure through local merging and pruning operations.
[0223] First, the system calculates the similarity between data points using the given similarity metric formula:
[0224]
[0225] Where:
[0226] ΔC is the change amount between nodes i and j, which measures the change in similarity between the two nodes.
[0227] x i and x j are the feature vectors of nodes i and j respectively, that is, the values of each node in different feature dimensions (such as temperature, humidity, soil humidity, etc.).
[0228] w k is the weighted coefficient of the k-th dimension feature, which controls the weight of each feature (for example, the importance of temperature and humidity).
[0229] p k is the distance metric exponent of the k-th dimension, which is used to adjust the sensitivity of the distance. Assume that when using common distance metric methods such as Euclidean distance or Manhattan distance, p k = 2.
[0230] α is the time decay factor, which is used to control the influence of the time span on the similarity. Assume that α = 0.05.
[0231] t i and t j are the timestamps of data points i and j, which are used to calculate the time decay factor represents the influence of the time difference on the similarity.
[0232] First, calculate the similarity between sensor 1 and sensor 2:
[0233] For the data points of sensor 1 and sensor 2, the eigenvectors of temperature, humidity, and soil humidity are calculated first:
[0234] x 1 =(25, 65, 30)
[0235] x 2 =(24.5, 66, 29)
[0236] Assume that the Euclidean distance is used as the distance metric index p k = 2, then calculate the similarity between them:
[0237] ΔC 12 =(0.7·|25 - 24.5| 2 + 0.3·|65 - 66| 2 )·e -α·|10:0010:05|
[0238] According to the above formula, the similarity obtained is:
[0239] ΔC 12 =(0.7·0.25 + 0.3·1)·e -0.05·5 =(0.175 + 0.3)·e -0.25 ≈0.475·0.7788 = 0.369
[0240] Secondly, judge whether merging is needed:
[0241] By calculation, ΔC 12 ≈0.369. According to the set similarity threshold, if ΔC 12 exceeds this threshold (for example, 0.3), it is considered that the states of sensor 1 and sensor 2 are similar enough to be merged.
[0242] At this time, the two nodes will be merged into a new node, and the eigenvector of the merged node is the weighted average of, for example:
[0243] x merged =(0.7·25 + 0.3·24.5, 0.7·65 + 0.3·66, 0.7·30 + 0.3·29) = (24.85, 65.3, 29.7)
[0244] Finally, pruning operation:
[0245] As new data points flow in, some redundant nodes may need to be pruned. If the similarity of a certain node is low, or it no longer makes a significant contribution to the accuracy of the clustering tree, it is removed through the pruning operation, thereby reducing the width of the tree and improving the calculation efficiency.
[0246] Assume the data point x of sensor 3 3=(26, 64, 32) and the merged node x merged =(24.85, 65.3, 29.7) has a relatively low similarity. After calculation, the similarity between them is obtained as follows:
[0247] ΔC 13 =(0.7 · |26 - 24.85| 2 + 0.3 · |64 - 65.3| 2 ) · e -α·|10:1010:00|
[0248] The calculation result is:
[0249] ΔC 13 ≈0.465
[0250] Due to this relatively low similarity, the system will choose to prune, remove the node of sensor 3, so as to reduce unnecessary expansion of the tree structure.
[0251] In the application scenario of the intelligent agricultural monitoring system in this embodiment, the sensor data of the greenhouse environment is constantly changing. The farm management personnel need to understand the status of the greenhouse in real time and then make corresponding adjustment decisions. In order to optimize the structure of the clustering tree and make it better adapt to the changes of the streaming data, the system adopts a streaming data hierarchical clustering optimization method based on the root insertion strategy. In this method, the dynamic adjustment mechanism plays a crucial role. The system adjusts the depth of the clustering tree according to the similarity distribution and density change of the real-time data to ensure the balance between clustering accuracy and computing efficiency.
[0252] Suppose the sensor data in the greenhouse flows into the system, and the sensor data is constantly changing. The farm administrator needs to adjust the control system of the greenhouse environment in time according to these data. The data measured by the sensors includes parameters such as temperature, humidity, CO2 concentration, light intensity, etc. These data continuously flow into the system and quickly form a large amount of data stream in a short time. In order to manage these real-time data, the system needs to hierarchically organize and manage the data through a clustering tree.
[0253] Set the sensor data flow rate in the greenhouse to 20 pieces of data per minute, and the data includes multiple dimensions such as temperature, humidity, and soil humidity. The system gradually inserts the data into the existing clustering tree through the root insertion strategy and dynamically adjusts the depth and breadth of the tree to optimize the storage and query efficiency of the data.
[0254] In order to make the clustering tree structure adapt to the continuous changes of the data stream, the system introduces a dynamic adjustment mechanism. The core of this mechanism is to adjust the depth or breadth of the tree according to the similarity distribution and density change of the real-time data. This means that if the similarity between the new data point and the root node of the current tree is high, the system will adjust the depth of the tree; if the distribution of the new data point is sparse, it may increase the breadth of the tree to maintain the clustering accuracy.
[0255] During the dynamic adjustment process, the following formula is used to calculate the adjustment amount of the depth or breadth of the tree:
[0256]
[0257] Where:
[0258] D adjust represents the adjustment amount of the depth or breadth of the tree, that is, the amount of change required to adjust the structure of the tree.
[0259] x i is the feature vector of data point i, containing information in multiple dimensions, such as temperature, humidity, etc.
[0260] x root is the feature vector of the root node of the clustering tree, representing the core data point of the clustering tree.
[0261] p is the exponent for adjusting the depth or breadth, which determines the sensitivity of the distance metric.
[0262] λ is the weight coefficient for adjusting the overall similarity.
[0263] α is the time decay coefficient, which controls the decay rate of the time difference between the data point and the root node.
[0264] t i and t root are the timestamps of data point i and the root node, used to calculate the time decay factor.
[0265] is the time decay factor, which makes data points with a larger time interval have less influence on the structure of the clustering tree.
[0266] Suppose at a specific time point, three groups of sensor data in the greenhouse have flowed into the system and been clustered into the current clustering tree. The current root node data of the system is the measurement value of sensor 1, with timestamp t root = 10:00, and the feature vector is x root = (25, 65, 30), representing temperature, humidity, and soil humidity.
[0267] Next, the data point of sensor 2 flows into the system at 10:05, with a feature value of x 2 = (24.5, 66, 29), and timestamp t 2 = 10:05. The data point of sensor 3 flows into the system at 10:10, with a feature value of x 3 = (26, 64, 32), and timestamp t 3 = 10:10.
[0268] The following parameters are used to calculate the adjustment amount of the tree:
[0269] p = 2 indicates the use of Euclidean distance metric.
[0270] λ = 0.5 represents the weighting coefficient of the distance.
[0271] α = 0.1 controls the speed of time decay.
[0272] Calculate the similarity between sensor 2 and the root node:
[0273]
[0274] Calculate the similarity between sensor 3 and the root node:
[0275]
[0276] According to the calculation results, the similarity between sensor 2 and the root node is 1.725, while the similarity of sensor 3 is 5.08. This indicates that the distance between sensor 3 and the root node is relatively far. Therefore, the system may need to separate the data by increasing the depth of the tree to avoid over-merging. The system dynamically adjusts the structure of the tree according to the value of the adjustment amount D adjust to ensure that the tree structure can adapt to the continuously incoming new data.
[0277] When continuously optimizing the sensor data management system for the greenhouse environment in the above embodiment, the farm manager hopes to effectively control the computational complexity while ensuring the clustering accuracy of the data. As the data stream continuously inputs, the structure of the clustering tree constantly changes. To avoid the high computational cost of global reconstruction, the system adopts a hierarchical clustering optimization method for stream data based on the root insertion strategy, combined with local merging and pruning operations, to balance the clustering accuracy and computational complexity.
[0278] The sensors of the greenhouse environment monitoring system collect parameter data such as temperature, humidity, and light intensity every minute. Inside the greenhouse, the temperature data flow is stable, and about 10 data points are collected per minute. As time goes by, the system continuously receives new data and performs clustering to ensure that the management personnel can obtain the status of the greenhouse in a timely manner.
[0279] In this case, the goal of the manager is to optimize the clustering tree structure so that it can accurately reflect the changes inside the greenhouse every time data flows in, while avoiding the overhead caused by global reconstruction calculations. To achieve this goal, the system introduces a mechanism for dynamically adjusting the tree structure and combines local merging and pruning operations, enabling the clustering tree to dynamically balance the computational complexity and accuracy during incremental updates.
[0280] Under this mechanism, the system will dynamically adjust the depth and breadth of the clustering tree based on the distance measurement between the data point and the current cluster center. This process uses the precision and complexity balance formula to calculate and control the balance between the precision and complexity of the clustering tree in real time. The specific balance calculation formula is as follows:
[0281]
[0282] in:
[0283] T balance Represents the balance between accuracy and complexity, that is, the balance factor calculated in the incremental update.
[0284] dist(x i ,x cluster ) represents the data point x i With the current cluster center x cluster The distance measure between .
[0285] γ and β are coefficients for adjusting the balance between accuracy and complexity, respectively, where:
[0286] γ controls the impact of accuracy on the amount of calculation, and usually takes values between 0.1 and 1.0. Larger values tend to improve accuracy.
[0287] β controls the effect of time decay on accuracy. The value range is usually 0.01, 0.5. A larger value will cause the time factor to have a greater impact on the calculation.
[0288] t i and t cluster Represents the data point x i and the timestamp of the current cluster center, which is used to quantify the time difference between the data point and the cluster center.
[0289] Assume that the cluster center of the current system is based on the temperature and humidity data in the previous few minutes, and the characteristic vector of the cluster center is x cluster =(26,70), timestamp t cluster =10:00, indicating that the temperature is 26°C and the humidity is 70%. At this time, the root node of the clustering tree is this data point.
[0290] Next, the system received three new sets of sensor data:
[0291] Data point 1 (temperature: 25°C, humidity: 68%) flows in at 10:02, timestamp t 1 =10:02, eigenvector x 1 =(25,68)
[0292] Data point 2 (temperature: 27°C, humidity: 72%) flows in at 10:05, timestamp t2 = 10:05, eigenvector x 2 = (27, 72)
[0293] Data point 3 (Temperature: 28°C, Humidity: 73%) flows in at 10:08, timestamp t 3 = 10:08, eigenvector x 3 = (28, 73)
[0294] Now, it is desired to calculate the distance metric between these three data points and the current cluster center through a formula and obtain the balance between accuracy and complexity. Assume the following parameters are selected:
[0295] γ = 0.8, meaning it is biased towards accuracy.
[0296] β = 0.1, with a slow time decay rate, and newer data will affect the structure of the cluster tree.
[0297] Balance of data point 1 with the cluster center:
[0298]
[0299] Balance of data point 2 with the cluster center:
[0300]
[0301] Balance of data point 3 with the cluster center:
[0302]
[0303] According to the calculation results, the balance of data point 1 with the cluster center is 1.35, data point 2 is 1.51, and data point 3 is 2.65. The system determines the relative relationship between the data points and the current cluster center through these calculations, and then decides whether to adjust the cluster tree structure. For example, the higher balance of data point 3 indicates a lower similarity with the current cluster center, and it may be necessary to adjust the depth of the cluster tree and increase the levels of the tree to accommodate this more distant data point. While data points 1 and 2 have a higher similarity, the computational efficiency can be optimized through local merging or pruning operations to reduce the computational overhead.
Claims
1. A stream data hierarchical clustering optimization method based on root insertion strategy, characterized by The following steps are involved: S1. Dynamic input and incremental preprocessing of streaming data: S1.1, receive streaming data from the data source in real time; S1.2, according to the set time window or data volume window; S1.3, standardize the data in the window; extract key features including distance, density, and time information; S2. Hierarchical clustering tree construction and root insertion strategy: S2.1, constructing a hierarchical clustering tree by presetting the streaming data and clustering parameters including clustering threshold and similarity metric; S2.2, determine the similarity between the new data and the root node of the current clustering tree; perform local update; And create a new root node; S2.3, new data points gradually expand the clustering tree structure and affect the expansion of the tree and the adjustment of the original subtrees; S3. Dynamic update and local optimization of clustering tree: S3.
1. Calculate the similarity between the new data and the existing subtree, adjust the subtree structure according to the similarity, and classify it into the existing subtree or create a new subtree branch; S3.
2. Adjust the tree shape through local merging or pruning operations; optimize the algorithm and adjust the structure; S3.3, by dynamically adjusting the depth and breadth of the tree, the accuracy and complexity of hierarchical clustering are balanced; S4. Cluster merging and splitting decisions: S4.1, real-time monitoring of the distance and similarity between adjacent subtrees in the clustering tree; S4.
2. triggering a split operation according to a preset split rule; S4.
3. Evaluate the clustering tree structure and optimize the tree depth, width and similarity metrics.
2. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 1 is characterized in that The method of constructing the hierarchical clustering tree and the root insertion strategy comprises: Introduce dynamic similarity measurement and a weighting mechanism based on streaming data timestamps to adjust the similarity calculation between data points; through the formula w k and p k Parameter control: introduce time weighting factor and adopt time decay design to reduce the impact of data points with large time span on clustering tree; The similarity measurement formula is as follows: in: S ij is the similarity measure between data points i and j; and Respectively represent the eigenvalues of data points i and j in the kth dimension; w k The weight coefficient of the k-th dimension feature; p k is the distance metric index of the kth dimension; p is the scale factor of the overall distance metric; α represents the time weighting factor; t i and t j The timestamps of data points i and j; is the time decay term, as the time difference |t i t j |Increase.
3. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 2 is characterized in that The method of constructing the hierarchical clustering tree and the root insertion strategy comprises: A dynamic root insertion strategy is used to insert new data points into the existing tree structure; the similarity between the data point and the existing clustering tree root node is evaluated and merged with the existing node; by comparing the similarity between the data point and the tree root; the computational burden of global reconstruction is reduced; The calculation formula is as follows: in: D insert It is the insertion measure between data point i and the root node of the cluster tree, indicating the similarity between the data point and the root of the tree; and They represent the eigenvalues of data point i and the root node in the kth dimension respectively; λ and β are constants that control the attenuation of similarity; λ controls the magnification factor of distance, and β controls the attenuation speed of similarity due to time interval; t i and t root is the timestamp of data point i and the root node; is the time attenuation term.
4. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 3 is characterized in that The method of constructing the hierarchical clustering tree and the root insertion strategy comprises: A local update mechanism is constructed, and each node of the clustering tree is adjusted according to the changes in data points; the tree structure is kept in a balanced state; the idea of calculus is used to adjust the structure of the clustering tree by calculating the relative position changes between nodes.
5. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 1 is characterized in that The dynamic update and local optimization of the clustering tree include: Through local merging and pruning operations; triggering merging operations by dynamically evaluating node similarity, merging multiple similar nodes into one; removing redundant nodes through pruning operations to reduce the width of the tree; combining time factors to measure the relevance of data points; The formula is as follows: in: ΔC represents the change between nodes i and j; x i and x j Represent the feature vectors of nodes i and j respectively; w k The weighting coefficient of the kth dimension; w k Used to adjust the weight of each feature; p k The distance metric index of the kth dimension; t i and t j Represent the timestamps of nodes i and j respectively; Represents the time decay factor.
6. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 5 is characterized in that The dynamic update and local optimization of the clustering tree include: Build a dynamic adjustment mechanism to adjust the structure based on real-time data; and increase or decrease the depth of the clustering tree by analyzing the similarity distribution and density changes of the data in real time; The formula is as follows: in: D adjust Indicates the depth or breadth adjustment of the tree; x i represents the feature vector of data point i; x root Represents the root node feature vector of the clustering tree; p adjusts the index of depth or breadth; λ and α adjust the coefficients of time decay and similarity weight; λ adjusts the overall weight, and α controls the speed of time decay; t i and t root Represents the timestamp of data point i and the root node; The time decay factor is used to calculate the time relationship between the data point and the root node when adjusting the tree structure.
7. The stream data hierarchical clustering optimization method based on root insertion strategy according to claim 6 is characterized in that The dynamic update and local optimization of the clustering tree include: Dynamically adjust the tree structure, combine local merging and pruning operations, balance accuracy and complexity in real time, adjust the amount of calculation according to the characteristics of the current data, maintain clustering accuracy while reducing computational overhead; optimize the computational burden of global reconstruction caused by incremental updates and local optimization mechanisms.