Density data stream clustering method based on adaptive online learning

By introducing microcluster structure and adaptive adjustment strategies into the data flow clustering method, dynamically updating the clustering parameters is solved, and the problem of feedback lag and parameter settings of clustering results in the existing technology depends on prior knowledge, realizing high-precision data flow clustering and fast online output.

CN115496133BActive Publication Date: 2025-05-23XIDIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211094825.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-05-23
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

When the existing density-based data flow clustering algorithm processes high-speed changing data flows, the clustering results are feedback lagging, and the parameter settings rely on prior knowledge and cannot adapt to the rapid evolution of the data flow, resulting in low clustering accuracy.

Method used

A data flow clustering method based on density is proposed. By introducing microcluster structure and adaptive adjustment strategy, clustering parameters are dynamically updated, and adaptive parameter adjustment is used to adjust the spatial information of the data flow to process the continuous evolution of the data flow.

Benefits of technology

It improves the accuracy of clustering results, reduces the input of expert prior knowledge, reduces the negative impact caused by user parameter settings, realizes the online rapid output of clustering results, and can check the current clustering status at any time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496133B_ABST
    Figure CN115496133B_ABST
Patent Text Reader

Abstract

The present invention discloses a data stream clustering method based on adaptive online learning, which mainly solves the problem of low accuracy of clustering results caused by fixed parameters and model dependence on parameters in the prior art. Its implementation scheme is: create micro-clusters according to data information in the data stream; after the micro-cluster receives a new data point, use the radius adaptive growth strategy to make the active micro-cluster absorb more data to learn the micro-cluster structure, and process the continuous evolution of the data stream through the energy update strategy of different types of micro-clusters, so that the energy decay of the micro-cluster simulates the evolution process, and the extinction of the micro-cluster causes the clustering change; according to the changed distance between micro-clusters, the aggregation of similar data is realized, and the data clustering results are output. The present invention dynamically optimizes the parameters of the micro-cluster model through an adaptive adjustment strategy, updates the micro-cluster model through online learning data, and improves the accuracy of data stream clustering in a dynamic data environment. It can be used for model learning of Internet data, network intrusion detection, network click stream and weather monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent information processing, and in particular relates to a density data stream clustering method, which can be used for model learning of Internet data, bank data processing, network intrusion detection, transaction flow, network click flow and weather monitoring. Background Art

[0002] Data stream refers to a huge sequence of data transmitted at high speed, which can only be read in a predetermined order. Traditional data is static and stable. It can be accessed and processed multiple times at any time. Data stream is dynamic, continuous, and changes over time. "Real-time", "continuous", and "ordered" are common words to describe data streams, and "large data volume", "potentially unlimited", and "uncertain arrival rate" are also its obvious characteristics. In most applications of Internet data model learning, bank data processing, network intrusion detection, transaction flow, network click stream, and weather monitoring, the real class label is not available for data stream instances, and because there is no prior knowledge about the number of categories, the staff needs to use unsupervised data clustering methods for the data in the data stream. Since data streams are usually transmitted at very high speeds, the calculation and storage of data stream data will become very difficult. Usually, there is only an opportunity to process the data once when it first arrives, and it is difficult to access the data at other times. Therefore, online data stream clustering is very valuable for research. This method can quickly feedback the clustering results. In addition, in a dynamic data environment, the data generated by the data stream is unstable, and there is a phenomenon that the data distribution changes over time, that is, concept drift. When dealing with data stream clustering in a dynamic data environment, how to save historical information, how to use historical information, and how to maintain historical information are important issues that affect the accuracy of data stream clustering.

[0003] Existing data stream clustering methods are mainly divided into hierarchical data stream clustering methods, partition-based data stream clustering methods, density-based data stream clustering methods, grid-based data stream clustering methods, and model-based data stream clustering methods, among which:

[0004] The hierarchical data stream clustering method is based on the binary tree data structure, which groups the given data into a cluster tree. The hierarchical clustering is divided into two types: agglomerative and divisive. The agglomerative algorithm adopts a bottom-up approach, that is, it assumes that each instance itself is a cluster, and creates clusters by gradually merging instances; the divisive algorithm adopts a top-down approach, that is, it assumes that a starting cluster contains all the data, and then splits the starting cluster into smaller clusters. The classic hierarchical algorithm is the BRITH algorithm, and the CF data representation structure it uses is adopted by many subsequent two-stage algorithms. Its limitation is that the process of forming micro-clusters is irreversible.

[0005] The data stream clustering method based on partitioning divides the data instance into several predefined partitions according to the similarity (or distance) between the data instance and the cluster centroid, where each partition represents a cluster. The partitioning algorithms include Clustream algorithm and strAP algorithm. Although this method is easy to implement, it can only find spherical clustering results, and the clustering results are easily affected by noise and the number of partitions.

[0006] The grid-based data stream clustering method uses a grid structure to divide multiple grid cells. Each instance is mapped to a grid cell, and the algorithm clusters the grid cells according to the density of the grid cells. In the grid-based algorithm, the running time does not depend on the number of input data but on the number of grid cells.

[0007] Grid-based algorithms include D-stream and MR-stream algorithms. Grid-based algorithms are fast algorithms that are also highly robust to noise and can find clusters of any shape. However, since the complexity of the algorithm depends on the dimension of the data, grid-based algorithms are more suitable for low-dimensional data. In addition, grid-based algorithms require the size of the grid to be predefined.

[0008] Model-based data stream clustering methods are generally based on the idea that "data sets conform to a certain distribution" and fit the data sets to be optimized with various data models. The EM algorithm can be seen as an extension of k-means. EM assigns objects to clusters based on weights that represent membership probabilities. Model-based algorithms have great limitations, and it is difficult to find data models with universal applicability.

[0009] Density-based data stream clustering method, which is divided into two types: two-stage mode and online mode. The two-stage mode uses micro-cluster structure to save the summary information of input data. Micro-cluster is a group of data instances that are very close to each other. The position and outline of micro-cluster are calculated according to the feature vector, and then these micro-clusters are merged into the final cluster according to the concepts of density reachability and density connectivity; the online mode uses micro-cluster structure to divide the data space into a core area and a subspace of non-core area. The micro-cluster structure can also characterize the spatial position, radius and life cycle. The online mode algorithm maintains a graph structure to represent the intersection relationship between the current micro-cluster and other micro-clusters. The intersecting core area constitutes a cluster. The graph structure will greatly reduce the micro-cluster separation operation and improve the processing speed. The density-based data stream clustering algorithm is the most popular method in data stream clustering. This method can process clusters of any shape, is also very robust to noise, and has high accuracy. The existing density-based two-stage algorithms are Denstream algorithm and SOStream algorithm, and the density-based online processing algorithms are CODAS algorithm and BOCEDS algorithm.

[0010] Since data streams all require fast and real-time processing of data objects, and most of the existing Denstream algorithm and SOStream algorithm adopt a two-stage clustering framework, the feedback of the data clustering result has a lag. The CODAS algorithm cannot handle the data evolution problem. The attenuation factor proposed by the BOCEDS algorithm for the data evolution problem adjusts the micro-cluster structure. Its fixed attenuation factor set by the user cannot adapt to the change of the underlying structure caused by the high-speed evolution of the data stream, resulting in a certain degree of influence on the final clustering result. In addition, many parameter settings of the existing density-based algorithms use prior knowledge and are fixed, and cannot make good use of the existing information to adjust the parameters to handle the continuous evolution of the data stream, resulting in low accuracy of the clustering result. Summary of the Invention

[0011] The object of the present invention is to overcome the deficiencies of the existing technologies, and propose a density-based adaptive online learning data stream clustering method to utilize the existing information for adaptive parameter adjustment, handle the continuous evolution of the data stream, and improve the accuracy of the clustering result.

[0012] To achieve the above object, the technical solution of the present invention includes the following:

[0013] (1) Receive the data stream in a dynamic data environment, and divide the data stream into n data blocks at an interval of 1000 data points according to the reception order, where n≥3;

[0014] (2) Create a micro-cluster structure separately according to the information of the first data point in the data block, and add the micro-cluster to the initially empty micro-cluster list;

[0015] (3) Calculate the Euclidean distance d between other data points X i in the data block and the micro-cluster center C in the micro-cluster list one by one, and map the data point to the micro-cluster with the smallest Euclidean distance, and judge whether the current data point is added to the micro-cluster:

[0016] If the Euclidean distance is less than the minimum radius R of the mapped micro-cluster, that is, when d<R, and the micro-cluster is a weak micro-cluster in the buffer, at this time the micro-cluster is activated into a core micro-cluster, then change the energy of the micro-cluster to 1, and then add the current data point to the mapped micro-cluster, and execute step (4);

[0017] Otherwise, create a new micro-cluster structure separately by the current data point and add it to the existing micro-cluster list, and execute step (5);

[0018] (4) After a micro-cluster in the micro-cluster list receives a new data point, perform an update operation:

[0019] (4a) When the data point Xi After adding, the radius R of the microcluster t Perform adaptive update to obtain the latest value R of the updated micro-cluster radius t+1 :

[0020]

[0021] Among them, N t+1 =N t +1 is the local density threshold of microclusters, N t ' is the spatial information count value of the micro-cluster, Decay is the decay factor of the micro-cluster, The ratio is the adaptive adjustment factor, R max is the maximum radius of the microcluster;

[0022] (4b) When the data point X i Located in the shell region, that is, the data point is at the center of the microcluster [0.5*R t+1 ,R t+1 ] range, the micro-cluster center C t Perform adaptive update to obtain the updated center latest value C t+1 :

[0023]

[0024] Among them, N t ' +1 =N t '+1 is the spatial information count value of the micro-cluster;

[0025] (4c) The energy E of the microcluster t Perform adaptive update to obtain the latest energy value E after update t+1 :

[0026] Execute step (5);

[0027] (5) The micro-cluster energy E' in the micro-cluster list t Attenuate and get the latest energy value E' after attenuation t+1 :

[0028] Execute step (6);

[0029] (6) According to the relationship between the attenuated micro-cluster energy value and 0, determine whether the micro-cluster has changed at the current moment:

[0030] If the energy value of the micro-cluster after attenuation is less than 0, the micro-cluster changes and changes accordingly according to the type of the micro-cluster: if the micro-cluster is a local density threshold N t Greater than the density threshold N thThe core micro-cluster is transformed into a weak micro-cluster in the buffer zone, and the structure of the weak micro-cluster is changed accordingly; otherwise, the micro-cluster is a weak micro-cluster in the buffer zone and is directly deleted;

[0031] If the energy value of the micro-cluster after attenuation is greater than or equal to 0, the micro-cluster does not change, and step (7) is executed;

[0032] (7) Calculate the intersection distance d′ between the current micro-cluster and all micro-clusters in the micro-cluster list, and add the micro-clusters whose Euclidean distance d of the micro-cluster center is less than the intersection distance d′ to their respective edge lists EL, that is, classify the intersecting micro-clusters into the same category, thus updating the macro cluster;

[0033] (8) After updating the macro cluster, the results belonging to the same category are output online to complete the clustering of the data stream.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] First, the present invention utilizes the potential spatial information of the data for the first time, introduces an adaptive adjustment strategy for the radius, energy, and center parameters in the clustering model, and dynamically optimizes the clustering parameters online. Through the adaptive update process of the radius, it reduces the expert prior knowledge input and reduces the negative impact of over-reliance on user parameters on the clustering parameters, thereby improving the accuracy of the clustering results.

[0036] Second, the present invention realizes the regeneration and extinction of micro-clusters through adaptive energy update and energy decay process, which can better adapt to the high-speed and constantly changing evolution process of data streams while reducing memory.

[0037] Third, by updating the macro clusters, the present invention can not only quickly output the clustering results, but also check the clustering results at any time, so as to better achieve timely interaction between the user and the clustering process. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flow chart for realizing the present invention;

[0039] Figure 2 The simulation results of clustering the KDD-CUP data set using the F1 evaluation index using the present invention and the existing CODAS algorithm and BOCEDS algorithm respectively;

[0040] Figure 3 The diagram is a simulation result diagram of clustering the KDD-CUP data set using the recall evaluation index using the present invention and the existing CODAS algorithm and BOCEDS algorithm respectively. DETAILED DESCRIPTION

[0041] The specific embodiments and effects of the present invention are further described in detail below with reference to the accompanying drawings.

[0042] Reference Figure 1 , the implementation steps of this example are as follows:

[0043] Step 1: Data stream segmentation.

[0044] The data stream is not read in all at once. The read data will change over time. In order to better reflect the clustering accuracy of the data in the set stage, the data stream will be divided into blocks according to the number of data read in a batch when it is used. Too much data in the data block will cause concept drift in the data block, and too little data in the data block will increase the weight of the noise data. These will affect the clustering algorithm's learning of the data, and then affect the accuracy of the final data stream clustering. Therefore, the data points in the data stream need to be divided into blocks in the specified order.

[0045] The specific implementation of this step is: the data stream received in the dynamic data environment is divided into n data blocks according to the receiving order with 1000 data points as an interval, where n≥3.

[0046] In this example, the data stream of the guassion dataset is divided into 28 data blocks, the Spiral dataset is divided into 9 data blocks, and the KDD-CUP dataset is divided into 480 data blocks. However, the division is not limited to the guassion dataset.

[0047] Step 2: Find the target micro-clusters of data points.

[0048] When new data X in the data stream i When arriving, according to the micro-cluster center C and the data point X i The Euclidean distance d between them needs to be mapped to the target micro-cluster.

[0049] The target micro-clusters include three types of micro-clusters, namely core micro-clusters, weak micro-clusters, and potential micro-clusters.

[0050] A core micro-cluster refers to one whose local density is greater than the minimum density threshold, and it can participate in the final clustering result output;

[0051] A weak microcluster refers to a microcluster that is degenerated from a core microcluster in the buffer. It can be reactivated as a core microcluster or completely deleted as data evolution occurs.

[0052] Potential microclusters are those microclusters whose local density is less than the minimum density threshold. When a potential microcluster receives a data point, the local density of the potential microcluster is checked and it grows into a core microcluster when the local density is greater than the minimum density threshold.

[0053] The specific implementation of this step is as follows:

[0054] 2.1) Create a micro-cluster structure separately according to the information of the first data point in the data block, and add this micro-cluster to the initially empty micro-cluster list. The micro-cluster structure includes the center C, radius R, energy E, local density threshold N, spatial information count N', and edge list EL. Among them, C = X i , X i is the current data point, E = 1, R = R min , R min is the minimum radius input by the user, N = 1, N' = 1, and EL is initialized as an empty set The micro-cluster list can store core micro-clusters and potential micro-clusters;

[0055] 2.2) For each of the other data points X in the data block j calculate its Euclidean distance d i from the center C of the micro-clusters in the micro-cluster list ij :

[0056]

[0057] where m is the dimension of the data, and k is the range of the k-th dimension of the data point from 1 to m;

[0058] 2.3) Find the minimum distance and the corresponding micro-cluster in the Euclidean distances. At this time, this micro-cluster is the target micro-cluster found, and map the current data point to the target micro-cluster, and judge whether the current data point is added to the micro-cluster:

[0059] If the Euclidean distance is less than the minimum radius R of the target micro-cluster, that is, when d < R, and this micro-cluster is a weak micro-cluster in the buffer area, at this time the micro-cluster is activated into a core micro-cluster, then change the energy of this micro-cluster to 1, and then add the current data point to the target micro-cluster;

[0060] Otherwise, create a new potential micro-cluster separately from the current data point. The creation of the potential micro-cluster is the same as the method of creating a micro-cluster separately, and add the potential micro-cluster to the existing micro-cluster list for subsequent data points to select.

[0061] Step 3, update the current micro-cluster.

[0062] When any micro-cluster receives a new data point, the micro-cluster will perform an update operation on the micro-cluster structure. Since the radius will not exceed the maximum radius set by the user when updating and growing, and only when the data point is within the specified area range will the micro-cluster center be updated. Therefore, the purpose of the update is to limit the micro-cluster from drifting endlessly following the data stream. The specific implementation is as follows:

[0063] 3.1) After the data point X i is added, for the radius R of the micro-cluster tPerform adaptive update to obtain the latest value R of the updated micro-cluster radius t+1 :

[0064]

[0065] Among them, N t+1 =N t +1 is the local density threshold of microclusters, N' t is the spatial information count value of the micro-cluster, The ratio is the adaptive adjustment factor, and Decay is the attenuation coefficient of the micro-cluster. By attenuating the micro-cluster energy with this coefficient, outdated micro-clusters can be removed in time, reducing memory usage while better adapting to the evolution of data streams. Since the radius will not exceed the maximum radius R set by the user when it is updated and increased max ,This avoids the unlimited growth of micro-clusters and is more in line with the actual clustering requirements;

[0066] 3.2) When data point X i Located in the shell region, that is, the data point is at the center of the microcluster [0.5*R t+1 ,R t+1 ] range, the micro-cluster center C t Perform adaptive update to obtain the updated center latest value C t+1 :

[0067]

[0068] Among them, N' t+1 =N' t +1 is the spatial information count value of the micro-cluster;

[0069] 3.3) Energy E of microclusters t Perform adaptive update to obtain the latest energy value E after update t+1 :

[0070]

[0071] By updating the micro-cluster energy, the micro-clusters that conform to the underlying structure of the current data stream can gain a larger life value, and by the same token, they will have a greater opportunity to learn other data points.

[0072] Step 4: Remove the dead microclusters.

[0073] 4.1) For the micro-cluster energy E' in the micro-cluster list t Attenuate and get the latest energy value E' after attenuation t+1 :

[0074]

[0075] 4.2) According to the relationship between the attenuated micro-cluster energy value and 0, determine whether the micro-cluster has changed at the current moment:

[0076] If the energy is greater than 0, that is, no micro-cluster dies, then continue to perform data stream clustering;

[0077] Otherwise, the following three extinction situations are judged and corresponding operations are performed:

[0078] If the energy value of the core micro-cluster is less than 0, that is, the core micro-cluster disappears, and the core micro-cluster becomes a weak micro-cluster. The weak micro-cluster is placed in the buffer, the micro-cluster energy becomes 0.5, and the edge list of the micro-cluster is cleared. The weak micro-cluster may not be applicable to the current data stream evolution, but as the data stream evolves, it may be recaptured at some point, thus forming the protection of important historical information. The use of historical information can improve accuracy and discover the evolution process and evolution trend of the data stream;

[0079] If the energy value of a potential micro-cluster is less than 0, that is, when the potential micro-cluster dies, the potential micro-cluster is directly deleted to reduce memory consumption.

[0080] If the energy value of the weak micro-cluster in the buffer is less than 0, that is, the weak micro-cluster disappears, indicating that the historical information it contains is eliminated by data evolution, or there are new micro-clusters that can better replace the weak micro-cluster to handle data stream changes. At this time, in order to reduce memory consumption and adapt to the characteristics of high-speed changes in data streams, it is completely deleted.

[0081] Step 5, update the macro cluster.

[0082] Maintain a cluster graph structure to generate macro clusters online and achieve real-time output of clustering results. The graph structure will be updated and maintained in the following situations:

[0083] When a microcluster becomes a core microcluster, the local density N t Greater than the density threshold N th ,It shows that the micro-cluster conforms to the current data stream evolution trend and can represent the data information well;

[0084] When a weak micro-cluster is captured by the current data point and activated into a core micro-cluster, it means that there may be some underlying connection between the current data and the previous historical information. This situation can well reflect the changes and is of great significance for in-depth research on subsequent changes.

[0085] When the center of the core micro-cluster changes, it involves the connection of new graph node edges and the disconnection of existing edges;

[0086] When the core micro-cluster degenerates into a weak micro-cluster and moves into the buffer zone.

[0087] The specific implementation is as follows:

[0088] 5.1) Calculate the intersection distance d′ between the current micro-cluster and all micro-clusters in the micro-cluster list:

[0089]

[0090] Where R is the radius of the current micro-cluster, and R′ is the radius of the micro-cluster in the micro-cluster list;

[0091] 5.2) Add the microclusters whose Euclidean distance d of the microcluster center is less than the intersection distance d′ to their respective edge lists EL. The edge list of a microcluster maintains the information of other microclusters that intersect with the current microcluster, and classifies the intersecting microclusters into the same category to update the macro cluster;

[0092] 5.3) Based on the updated macro clusters, the obtained clustering results are output online in real time.

[0093] The clustering results of the present invention are interrelated with the cluster graph structure. When the graph structure changes, the clustering results will also change accordingly. The present invention can obtain the clustering results at any time online in real time and capture the changes in time. Different from the previous algorithms that require a specified time interval to obtain the clustering result output, at the same time, in some fields, users are more interested in the occurrence of changes and the reasons for the changes. The present invention can provide assistance for subsequent potential information mining.

[0094] The following is a description of the technical effects of the present invention in combination with simulation experiments:

[0095] 1. Simulation conditions

[0096] The experimental data uses the KDD-CUP dataset, the Gaussian dataset, and the Spiral dataset. The simulation platform is: an Intel Core i5-4590 CPU with a main frequency of 3.30GHz, 12.0GB of memory, Windows 10 operating system, and Matlab2021a development platform.

[0097] 2. Simulation content

[0098] Simulation 1: Clustering simulation is performed on gaussian data using the present invention, the existing CODAS algorithm and the BOCEDS algorithm respectively. The F1 score, NMI, RI, recall, Purity purity and Ac accuracy indicators are used to evaluate their respective clustering performances. The results are shown in Table 1.

[0099] Table 1 Evaluation of guassian data simulation results

[0100] It can be seen from Table 1 that among the six evaluation indicators used in the experiment, the present invention is higher than the prior art, which proves that compared with the prior art, the present invention can improve the data stream clustering results in a dynamic data environment.

[0101] Simulation 2: Clustering simulation of Spiral data is performed using the present invention and the existing CODAS algorithm and BOCEDS algorithm respectively, and the F1 score, NMI, RI, recall, Purity purity and Ac accuracy indicators are used to evaluate their respective clustering performances. The results are shown in Table 2.

[0102] Table 2 Evaluation of Spiral data simulation results

[0103]

[0104]

[0105] It can be seen from Table 2 that in the presence of noise, 5 of the 6 evaluation indicators used in the experiment are higher than the prior art and have been greatly improved, and one of them is equivalent to the prior art, which proves that the present invention can still improve the clustering results in the presence of noise and has noise resistance.

[0106] Simulation 3, using the present invention and the existing CODAS algorithm and BOCEDS algorithm to perform clustering simulation on the KDD-CUP data set, and using the F1 score to evaluate the clustering performance. The results are as follows: Figure 1 , use recall to evaluate the clustering results, such as Figure 2 , the i-boceds corresponding curve in the figure is the simulation result of the present invention.

[0107] from Figure 1 and Figure 2 It can be seen that on the real network intrusion data set, the clustering result of the present invention is better than that of the prior art BOCEDS, and although the index of CODAS is higher than that of the present invention to a certain extent, since CODAS does not involve the evolution of clusters, it has higher time complexity and larger memory consumption, and is not suitable for the existing data stream clustering development trend. By comparing with the advanced BOCEDS algorithm, it can be concluded that the present invention can handle real network intrusion data well and achieve good clustering accuracy in real data. At the same time, the fast and efficient clustering ability of the present invention can be applied to other application fields.

[0108] The above simulation results show that the present invention can better meet the high-speed and constantly changing evolution of data streams, quickly output clustering results online, and can check the clustering status at any time, so as to achieve timely interaction. It also reduces the negative impact of over-reliance on user parameters on clustering parameters, and while ensuring that data distribution of any shape can be processed, it can achieve online learning and updating of clustering models through the use of adaptive and spatial information, so as to better improve the clustering accuracy of data streams.

Claims

1. A data stream clustering method based on adaptive online learning, characterized in that, it includes the following steps: (1) Receive the data stream in a dynamic data environment, and divide the data in the data stream into n data blocks at intervals of 1000 data points according to the receiving order, where n ≥ 3; (2) Create a micro-cluster structure separately according to the information of the first data point in the data block, and add the micro-cluster to the initially empty micro-cluster list; (3) Calculate other data points X in the data block i The Euclidean distance d between each data point and the micro-cluster center C in the micro-cluster list is calculated one by one, and the data point is mapped to the micro-cluster with the smallest Euclidean distance to determine whether the current data point is added to the micro-cluster: If the Euclidean distance is less than the minimum radius R of the mapped micro-cluster, that is, when d < R, and the micro-cluster is a weak micro-cluster in the buffer, at this time the micro-cluster is activated into a core micro-cluster, then change the energy of the micro-cluster to 1, and then add the current data point to the mapped micro-cluster, and execute step (4); Otherwise, create a new micro-cluster structure separately from the current data point and add it to the existing micro-cluster list, and execute step (5); (4) After the micro-cluster in the micro-cluster list receives a new data point, perform an update operation: (4a) When the data point X i After adding, the radius R of the microcluster t Perform adaptive update to obtain the latest value R of the updated micro-cluster radius t+1 : Among them, N t+1 =N t +1 is the local density threshold of microclusters, N′ t is the spatial information count value of the micro-cluster, Decay is the decay factor of the micro-cluster, The ratio is the adaptive adjustment factor, R max is the maximum radius of the microcluster; (4b) When the data point X i Located in the shell region, that is, the data point is at the center of the microcluster [0.5*R t+1 ,R t+1 ] range, the micro-cluster center C t Perform adaptive update to obtain the updated center latest value C t+1 : Among them, N′ t+1 =N′ t +1 is the spatial information count value of the micro-cluster; (4c) The energy E of the microcluster t Perform adaptive update to obtain the latest energy value E after update t+1 : Execute step (5); (5) The microcluster energy E′ in the microcluster list t Attenuate and get the latest energy value E′ after attenuation t+1 : Execute step (6); (6) According to the size relationship between the attenuated micro-cluster energy value and 0, judge whether the micro-cluster changes at the current moment: If the energy value of the micro-cluster after attenuation is less than 0, the micro-cluster changes and changes accordingly according to the type of the micro-cluster: if the micro-cluster is a local density threshold N t Greater than the density threshold N th The core micro-cluster is transformed into a weak micro-cluster in the buffer zone, and the structure of the weak micro-cluster is changed accordingly; otherwise, the micro-cluster is a weak micro-cluster in the buffer zone and is directly deleted; If the attenuated energy value of the micro-cluster is greater than or equal to 0, then the micro-cluster does not change, and execute step (7); (7) Calculate the intersection distance d′ between the current micro-cluster and all micro-clusters in the micro-cluster list, and add the micro-clusters whose Euclidean distance d of the micro-cluster center is less than the intersection distance d′ to their respective edge lists EL, that is, divide the intersecting micro-clusters into the same class to realize the update of the macro cluster; (8) Output the result of the same class after updating the macro cluster online to complete the clustering of the data stream.

2. The method according to claim 1, characterized in that, The micro-cluster structure in step (2) includes a center C, a radius R, an energy E, a local density threshold N, a spatial information count N′ and an edge list EL, wherein C=X i , X i is the current data point, E=1, R=R min , R min is the minimum radius entered by the user, N = 1, N' = 1, EL is initialized to an empty set 3. The method according to claim 1, characterized in that, In step (6), the structure of the weak micro-cluster is changed accordingly, that is, the energy of the weak micro-cluster is first changed to 0.5, then the micro-cluster in the weak micro-cluster edge list is found, then the information record of the micro-cluster and the weak micro-cluster is cleared, and finally the edge list of the weak micro-cluster is set to an empty set.

4. The method according to claim 1, characterized in that, In step (7), the intersection distance d′ between the current micro-cluster and all micro-clusters in the micro-cluster list is calculated according to the following formula: where R is the radius of the current micro-cluster and R′ is the radius of the micro-cluster in the micro-cluster list.

Citation Information

Patent Citations

  • Mixed attribute data flow clustering method for automatically determining clustering center based on density

    CN105139035A

  • Density peak clustering method based on adaptive micro-cluster fusion

    CN111914930A