Granular ball generation algorithm for data flow

Through the sphere generation algorithm for data flow, combined with LORE-GBG and NNGBM algorithms, the problem that the classic sphere generation algorithm cannot effectively process data flow is solved, and the effect of efficiently processing data flow and maintaining classification accuracy is achieved.

CN120217189APending Publication Date: 2025-06-27成佳睿
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510098605.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Classic pellet generation algorithms cannot effectively process data streams, resulting in low data processing accuracy.

Method used

A sphere generation algorithm for data flow is proposed. The sphere model is generated through the LORE-GBG algorithm, and the current sphere model and the sphere model obtained by the previous data flow are merged using the NNGBM algorithm to form a DSGBG algorithm to efficiently process the data flow.

Benefits of technology

Efficient processing of data flow is achieved, and the classification accuracy and number of pellet models obtained are comparable to those generated directly on the overall data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217189A_ABST
    Figure CN120217189A_ABST
Patent Text Reader

Abstract

The invention provides a data stream-oriented granular ball generation algorithm, which comprises the following steps of: 1, for a current data stream, generating a granular ball model by using an LORE-GBG algorithm; and 2, combining the current granular ball model with the granular ball model obtained by the previous data stream by using an NNGBM algorithm. According to the method, a DSGBG algorithm can effectively process data streams by means of the merging capacity of the NNGBM algorithm on the two granular ball models, the DSGBG algorithm can efficiently process the data streams, and the classification precision and the granular ball number of the obtained granular ball models can be comparable with those of granular ball models directly generated on a whole data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of algorithm technologies, and particularly to a granule ball generation algorithm for data streams. Background Art

[0002] With the development of the Internet of Things, the processing of various data streams based on real-time sensing data has gradually become the key to the construction of current Internet of Things applications. However, the classical granule ball generation algorithm cannot effectively process data streams.

[0003] After entering the 21st century, with the rapid development of computers, communication technologies, and social media, the growth rate of data has become faster and faster, and the era of big data has arrived. A large amount of data contains a lot of important information waiting to be explored. However, due to the huge volume of data, it is very difficult to perform traditional manual analysis and mining on this data. The big data mining technology can just make up for the deficiencies of manual classification, efficiently and accurately mine valuable information on the huge data body, provide a reference basis for people's decision-making, and greatly facilitate people's lives.

[0004] Data streams are different from traditional batch processing data. They are processed immediately when the data arrives, rather than accumulating a certain amount of data over a period of time and then processing it. Therefore, data stream processing usually requires the use of real-time algorithms and technologies in order to be able to process and analyze data in a timely manner.

[0005] Data streams have the following three characteristics: data arrives quickly, that is, a large amount of input data may need to be processed in a short period of time; the value range of data attributes is relatively wide, which means that the system cannot completely save this information, and usually can only access the data once when the data arrives; data arrives continuously, that is, the amount of data may be infinite. Data stream mining is a process of incremental learning and adapting to the dynamically changing data stream environment. First, it is necessary to perform data preprocessing on the data stream, and then continuously perform incremental learning. The specific method is: first establish an initial model based on part of the data; then update or modify the existing model according to the latest arrived data, and at the same time create a certain elimination rule to eliminate obsolete data; finally, useful information can be dynamically extracted.

[0006] This article combines the natural neighbor theory and proposes the granule ball natural neighbor, the granule ball merging algorithm based on the granule ball natural neighbor (NNGBM), and the DSGBG algorithm. The NNGBM algorithm can effectively compress the number of granule balls in the granule ball model while maintaining the classification accuracy of the original granule ball model, thereby improving the quality of the granule ball model. The DSGBG algorithm can efficiently process data streams, and the classification accuracy and the number of granule balls of the obtained granule ball model can be comparable to those of the granule ball model generated directly on the overall data set. Summary of the Invention

[0007] This application provides a granule generation algorithm for data streams to solve the problem of low data processing accuracy in the prior art, realizing efficient processing of data streams. The classification accuracy and the number of granules of the obtained granule model can be comparable to those of the granule model generated directly on the overall data set.

[0008] This application provides a granule generation algorithm for data streams, including the following steps:

[0009] First step, for the current data stream, use the LORE-GBG algorithm to generate a granule model;

[0010] Second step, use the NNGBM algorithm to merge the current granule model and the granule model obtained from the previous data stream. Relying on the merging ability of the NNGBM algorithm for the two granule models, the DSGBG algorithm can effectively process the data stream.

[0011] Preferably, the specific process of the DSGBG algorithm is as follows:

[0012] Input: data stream S, granule model GBs generated from the previous data stream, purity threshold p;

[0013] Output: granule model,

[0014] ①GBs new =LORE-GBG(S);

[0015] ②GBs = NNGBM(GBs,GBs new ,p)

[0016] Output GBs.

[0017] Preferably, the specific process of the NNGBM algorithm is as follows:

[0018] Input: granule model GBs, purity threshold p;

[0019] Output: granule model,

[0020]

[0021]

[0022] Preferably, the related definitions and formulas of the NNGBM algorithm are as follows:

[0023] If a given granule GB is provided, the following information is known: the radius r of the granule, the center c of the granule, the sample distribution {N1, N2, ···, N k}, the total number of samples N of the granule, where k represents that the data set has k categories, and N k represents the number of samples of the kth category in the given granule. The related definitions of natural neighbors of granules are as follows:

[0024] Given granule ball GB i , granule ball GB j , then it is easy to obtain GB i , GB j The purity of the merged granule ball is shown in the following formula:

[0025]

[0026] Given granule ball GB i , granule ball GB j , according to the formula, it is easy to obtain the center of the merged granule ball as shown in the following formula:

[0027]

[0028] According to the above two formulas, the purity and the center point of the two merged granule balls can be calculated quickly without all relevant data participating in the calculation, which enables the granule ball merging algorithm based on natural neighbors to efficiently merge granule balls.

[0029] Preferably, the NNGBM algorithm is based on natural neighbors, and the fast granule ball nearest neighbor solving algorithm based on natural neighbors has the following process:

[0030] Input: granule ball model GBs = {centers, radius};

[0031] Output: the nearest neighbor of the granule ball.

[0032]

[0033]

[0034] Output NN.

[0035] Preferably, the natural neighbor search algorithm includes the following steps:

[0036] Input: data set D;

[0037] Output: natural neighbor eigenvalue λ, λ neighbor NN.

[0038]

[0039] Output λ, NN.

[0040] Beneficial effects: The NNGBM algorithm of the present invention can effectively compress the number of granule balls in the granule ball model while maintaining the classification accuracy of the original granule ball model, thereby improving the quality of the granule ball model. The DSGBG algorithm can efficiently process data streams, and the classification accuracy and the number of granule balls of the obtained granule ball model can be comparable to those of the granule ball model directly generated on the overall data set.

[0041] The above description is only an overview of the technical solution of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically illustrates the specific implementation manners of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a data stream mining framework diagram of the prior art;

[0044] Figure 2 It is a diagram of the results of granule ball merging at each stage and finally after simulating the data stream using the fourclass dataset in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the drawings are intended to cover non-exclusive inclusion.

[0047] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase "embodiments" appearing in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0048] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings.

[0049] This application provides a granule generation algorithm for data streams, including the following steps:

[0050] In the first step, for the current data stream, the LORE-GBG algorithm is used to generate a granule model;

[0051] In the second step, the NNGBM algorithm is used to merge the current granule model and the granule model obtained from the previous data stream. Relying on the merging ability of the NNGBM algorithm for the two granule models, the DSGBG algorithm can effectively process data streams.

[0052] Among them, the specific process of the DSGBG algorithm is as follows:

[0053] Input: data stream S, granule model GBs generated from the previous data stream, purity threshold p;

[0054] Output: granule model,

[0055] ① GBs new = LORE-GBG(S);

[0056] ② GBs = NNGBM(GBs, GBs new , p)

[0057] Output GBs.

[0058] In the present invention, the specific process of the NNGBM algorithm is as follows:

[0059] Input: granule model GBs, purity threshold p;

[0060] Output: granule model,

[0061]

[0062] Output GBs.

[0063] Meanwhile, as a tool that can fuse multiple granule ball models, the NNGBM algorithm enables the granule ball calculation method to further improve the quality of the granule ball model through multiple learning processes on the dataset. At the same time, the granule ball model obtained by the granule ball calculation method has the ability to be extended. Based on this, various granule ball generation algorithms that are either fast or have better result quality can be proposed for various scenarios. For example, in the next section of this paper, a corresponding granule ball generation algorithm will be proposed for the data stream scenario based on the NNGBM algorithm. To accelerate the granule ball generation process, a sampling-based granule ball generation algorithm can also be proposed. The specific process of this algorithm is divided into two steps: First, randomly sample the dataset multiple times and generate granule ball models on the sampled datasets; then merge the obtained multiple granule ball models based on the NNGBM algorithm. This method is based on the divide-and-conquer idea and can accelerate the granule ball generation process. Finally, useful information can be dynamically extracted.

[0064] The relevant definitions and formulas of the NNGBM algorithm are as follows:

[0065] If a given granule ball GB is provided, the following information is known: the radius r of the granule ball, the center c of the granule ball, the sample distribution {N1, N2, ···, N k}, the total number of samples N of the granule ball, where k represents that the dataset has k categories, and N k represents the number of samples of the k-th category in the granule ball. The relevant definitions of the natural neighbors of the granule ball are as follows:

[0066] Given granule ball GB i and granule ball GB j , it is easy to obtain that the purity of the merged granule ball of GB i and GB j is shown by the following formula:

[0067]

[0068] Given granule ball GB i and granule ball GB j , according to the formula, it is easy to obtain that the center of the merged granule ball is shown by the following formula:

[0069]

[0070] According to the above two formulas, the purity and the center point of the merged two granule balls can be quickly calculated without all relevant data participating in the calculation, which enables the granule ball merging algorithm based on natural neighbors to efficiently merge granule balls.

[0071] The NNGBM algorithm also uses purity as the criterion for merging to avoid reducing the classification accuracy of the granule sphere model after merging. This algorithm mainly consists of three steps: First, all natural neighbor granule sphere pairs are obtained according to formula (2); Second, if the purity of the new granule sphere formed by this granule sphere pair is still greater than the purity threshold, then merge this granule sphere pair; Third, recalculate the natural neighbor pairs of granule spheres and try to merge until no more granule spheres can be merged.

[0072] The NNGBM algorithm merges granule spheres based on the natural neighbors of granule spheres rather than merging the nearest pair of granule spheres each time, which has two advantages: During one iteration of the granule sphere merging algorithm, multiple pairs of granule spheres can be merged instead of just the nearest pair, so the efficiency of the algorithm can be improved; If the nearest pair of granule spheres is merged each time, when the purity cannot meet the requirements of the purity threshold after the nearest pair of granule spheres is merged, the granule sphere merging algorithm will be forced to stop, but this does not mean that there are no granule spheres that should be merged, so it may lead to a low compression rate. Merging granule spheres based on the natural neighbors of granule spheres can just solve this problem.

[0073] If two objects are neighbors to each other, then these two objects are natural neighbors of each other, where the search range of neighbors is determined by the natural neighbor search algorithm.

[0074] The natural neighbor search algorithm works by continuously expanding the neighbor search range and calculating the number of times each object is regarded as a neighbor by other objects until all objects are regarded as neighbors by other objects at least once, or the number of objects not regarded as neighbors by other objects does not change. This search process can be completed automatically without any parameters.

[0075] The process of the natural neighbor search algorithm is as follows:

[0076] Input: Dataset D;

[0077] Output: Natural neighbor feature value λ, λ neighbors NN,

[0078]

[0079]

[0080] Given a granule sphere set GBs, granule sphere GB i , GB j ∈GBs, whose radius and center are r i , r j and c i , c j . According to the distance dist(GB i , GB j ) between every two granule spheres in GBs, if granule sphere GB iGranular Ball GB j If granular balls GB i are the nearest neighbors of each other, then granular ball GB j is called the natural neighbor of granular ball GB j , granular ball GB i is called the natural neighbor of granular ball GB i , that is, NaN(GB j ) = GB j and NaN(GB i ) = GB i . At the same time, (GB j , GB i ) is called the natural neighbor granular ball pair. Where N(GB) represents the nearest neighbor of granular ball GB

[0081]

[0082] ) = {GB j | GB j == N(GB i ) ∧ GB i == N(GB j )}

[0083] In the present invention, the NNGBM algorithm is based on natural neighbors. The following is the process of the fast granular ball nearest neighbor solving algorithm based on natural neighbors:

[0084] Input: Granular ball model GBs = {centers, radius};

[0085] Output: The nearest neighbor of the granular ball.

[0086]

[0087] Output NN.

[0088] This method is mainly divided into three steps: The first step is to use the natural neighbor search algorithm in the granular ball center list to solve the natural neighbors of each granular ball center.

[0089] The second step,

[0090] According to the formula calculate the distance from each granular ball to the granular ball corresponding to the other granular ball center in the natural neighbors of its center, and select the nearest granular ball as the nearest granular ball of this granular ball.

[0091] The time complexity of the first step of this method is O(NlogN), and the time complexity of the second step is O(λN). Therefore, the total time complexity of this method is O(NlogN), where N is the number of granular balls in the granular ball model, and λ is the natural neighbor eigenvalue of the granular ball center list, and this value is a constant. Therefore, this method can accelerate the process of solving the nearest neighbor of the granular ball.

[0092] As Figure 2 shown, it shows the results of particle ball merging at each stage and the final result after simulating the data stream using the fourclass dataset. Figure 2 Among them, (a)-(d) are the results of particle ball generation after four random samplings of 100 points each, and (e) is the merging result of 4 particle ball models.

[0093] The specific approach of this experiment is to perform 4 random samplings without replacement on the fourclass dataset, with a sample size of 100 each time. Particle ball models are generated and shown on 4 datasets respectively, and then the 4 obtained particle ball models are merged in sequence using the NNGBM algorithm to obtain the final particle ball model. It can be found from the figure that all 4 particle ball models can well describe the data they contain, and the final merged result can also well describe the overall 400 sample points. This intuitively demonstrates the effectiveness of the DSGBG algorithm.

[0094] Experiments were conducted using the UCI datasets HIGGS and codrna to simulate the data stream. The simulation method is to randomly sample a certain number of sample points from the dataset each time without replacement. Among them, 10,000 sample points are randomly sampled without replacement from the HIGGS dataset each time, and 2,000 sample points are randomly sampled without replacement from the codrna dataset each time. Correspondingly, for the HIGGS and codrna datasets, the number of sample points in the test set is ten thousand and two thousand respectively. The information of the datasets HIGGS and codrna is shown in the following table:

[0095]

[0096] The LORE-GBG algorithm, KNN algorithm, Support Vector Machine (SVM) algorithm, and Decision Tree (DT) algorithm were used as controls. For the LORE-GBG algorithm, its parameter purity threshold was set to 0.8. Each time, all the currently obtained data was used as the input of the LORE-GBG algorithm to generate a particle ball model, and then the particle ball KNN algorithm was executed on this particle ball model to obtain the classification accuracy. For the KNN algorithm, its parameter K was set to 5, and each time all the currently obtained data was used as the input to obtain the classification accuracy. For the SVM and DT algorithms, their parameters were set to the default parameters of the corresponding algorithms in the python code library SKlearn. For the DSGBG algorithm, each time the particle ball model obtained in the previous round and the data in the current round were used as the input, and then the obtained particle ball model was executed with the GBKNN algorithm to obtain the classification accuracy, as shown in the following table:

[0097]

[0098] The above table shows the comparison of the classification accuracies of the DSGBG, LORE-GBG, and KNN algorithms after adding data each time on the HIGGS dataset. It can be found that in these 10 rounds, the classification accuracies obtained by the two granule ball generation algorithms are both better than those of the KNN, SVM, and DT algorithms. Among them, the DSGBG algorithm has a better classification accuracy 3 times, and the LORE-GBG algorithm has a better classification accuracy 7 times. Moreover, regardless of which algorithm has a better classification accuracy, the gap is not large.

[0099] As shown in the following table, it shows the comparison of the classification accuracies of the three algorithms after adding data in each round on the codrna dataset. It can be found that on the codrna dataset, the SVM algorithm has the best classification accuracy, and the classification accuracies of the three classical classification algorithms are all better than those obtained by the two granule ball generation algorithms. For the LORE-GBG algorithm and the DSGBG algorithm, the LORE-GBG algorithm is better 5 times, and the DSGBG algorithm is better 5 times. Similarly, the classification accuracy gap between the two algorithms is not large.

[0100]

[0101]

[0102] The following two tables show the comparison of the number of granule balls in the granule ball models obtained by the DSGBG and LORE-GBG algorithms after adding data each time on the HIGGS and codrna datasets. It can be found that in the 10-round experiments on these two datasets, the difference in the number of granule balls in the granule ball models obtained by the two granule ball generation algorithms is not large. This shows that the natural neighbor-based granule ball merging algorithm (NNGBM) has the ability to merge the granule balls that should be merged and does not merge the granule balls that should not be merged, making the final merging result close to directly generating granule balls on the dataset. Therefore, the following conclusions can be drawn: (1) In terms of classification accuracy, the DSGBG algorithm can be comparable to the LORE-GBG algorithm, and it has its own advantages and disadvantages compared with the three classical classification algorithms of KNN, SVM, and DT on the two datasets. However, the LORE-GBG algorithm and the three classical classification algorithms process all the data, while the DSGBG algorithm processes the granule ball model obtained in the previous round and the data in this round. Therefore, relatively speaking, the DSGBG algorithm has a greater efficiency advantage. (2) The difference in the number of granule balls in the granule ball model finally obtained by the DSGBG algorithm and the number of granule balls in the granule ball model obtained by the LORE-GBG algorithm is not large, indicating that the DSGBG algorithm can maintain the classification accuracy of the granule ball model with as few granule balls as possible.

[0103]

[0104]

[0105] In summary, the NNGBM algorithm of the present invention can effectively compress the number of granular balls in the granular ball model while maintaining the classification accuracy of the original granular ball model, thereby improving the quality of the granular ball model. The DSGBG algorithm can efficiently process data streams, and the classification accuracy and the number of granular balls of the granular ball model obtained by it can be comparable to those of the granular ball model generated directly on the overall data set.

[0106] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A data stream-oriented sphere generation algorithm, characterized in that: The following steps are involved: The first step is to generate a granular sphere model using the LORE-GBG algorithm for the current data stream; In the second step, the NNGBM algorithm is used to merge the current granular sphere model with the granular sphere model obtained from the previous data stream. Relying on the NNGBM algorithm's ability to merge two granular sphere models, the DSGBG algorithm can effectively process the data stream.

2. The data stream-oriented sphere generation algorithm according to claim 1, characterized in that: The specific process of the DSGBG algorithm is as follows: Input: data stream S, particle sphere model GBs generated by the previous data stream, purity threshold p; Output: granular sphere model, ①GBs new =LORE-GBG(S); ②GBs=NNGBM(GBs,GBs new ,p) Output GBs.

3. The data stream-oriented sphere generation algorithm according to claim 1, characterized in that: The specific process of the NNGBM algorithm is as follows: Input: granular sphere model GBs, purity threshold p; Output: granular sphere model, ①repeat; ② According to Algorithm 2, all natural neighbor sphere pairs NNGBs in GBs are obtained ③for(GB i ,GB j )in NNGBs do if purity(GB i ∪GB j )>p then Merge Ball GB i and GB j end end ④until(cannot synthesize pellets anymore) ⑤Output GBs.

4. The data stream-oriented sphere generation algorithm according to claim 3, characterized in that: The relevant definitions and formulas of the NNGBM algorithm are as follows: If a sphere GB is given, the following information is known: the radius r of the sphere, the center c of the sphere, the sample distribution {N1, N2, ···, N k }, the total number of ball samples N, where k means that the data set has k categories, N k Represents the number of samples of the kth category in the particle sphere. The natural neighbor correlation of the particle sphere is defined as follows: Given a ball GB i , ball GB j , then GB is easily obtained i , GB j The purity of the combined pellets is given by the following formula: Given a ball GB i , ball GB j According to the formula, the center of the merged sphere is easily obtained as shown in the following formula: According to the above two formulas, the purity and center point of two spheres after merging can be quickly calculated without the need for all relevant data to participate in the calculation, which enables the sphere merging algorithm based on natural neighbors to efficiently merge spheres.

5. The data stream-oriented sphere generation algorithm according to claim 2, characterized in that: The NNGBM algorithm is based on natural neighbors. The fast granular sphere nearest neighbor solution algorithm based on natural neighbors is as follows: Input: granular spherical model GBs = {centers, radius}; Output: The nearest neighbors of the sphere. ① ② Solve the nearest neighbor NaN of centers according to λ; ③for i=1,···,|GBs| do Calculate the distance dist between particle i and the particle corresponding to NaN(centers(i)) / / The nearest neighbor of ball i is the ball with the minimum distance in dist NN(i)=argmin(dist) end Output NN.

6. The data stream-oriented sphere generation algorithm according to claim 5, characterized in that: The natural neighbor search algorithm comprises the following steps: Input: Dataset D; Output: natural neighbor eigenvalue λ, λ nearest neighbor NN. ① Construct a KD tree on D; ②repeat; ③for point p in D do / / Use KD tree to search for p's r nearest neighbor x nb(x)=nb(x)+1 NN r (p)=NN r (p)∪{x} RNN r (x)=RNN r (p)∪{p} end ④r=r+1 ⑤until (the number of points with nb value equal to 0 is 0 or unchanged) ⑥λ=r,NN=NN r ⑦Output λ, NN.