A Knowledge Discovery Method and Device for Incremental Datasets

By designing the EFPT-IKD algorithm of frequent pattern trees and incremental windows, the high-complexity problem of new knowledge discovery on incremental data sets is solved, and efficient knowledge mining and association mining are realized when the data volume increases, reducing time complexity and improving computing efficiency.

CN112925839BActive Publication Date: 2025-07-18CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110107823.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-27
Publication Date
2025-07-18
Estimated Expiration
2041-01-27

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently discover new knowledge on incremental data sets, and faces the problems of high time complexity and high spatial complexity. Especially in the scenario where the amount of data is constantly expanding, the method based on tree data structure requires frequent scanning of the entire data, while the method based on incremental learning has high calculation cost and high requirements for data quality.

Method used

A frequent pattern tree structure and incremental window are designed to maintain new frequent transaction items through the EFPT-IKD algorithm, and dynamically update the frequent pattern tree to realize real-time adjustment of the incremental data set and mining of the association relationship, reducing time complexity.

Benefits of technology

It realizes that new knowledge can be discovered in a timely and accurate manner with the increasing amount of data, reduces time complexity, and improves the efficiency and adaptability of data calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112925839B_ABST
    Figure CN112925839B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a knowledge discovery method for incremental datasets. The knowledge discovery method and device for incremental datasets of the present invention use the EFPT-IKD algorithm and design a tree-shaped data structure - frequent pattern tree that can continuously evolve as the amount of data grows. An incremental window (IW) is set to discover newly added frequent transaction items. The frequent pattern tree is mainly used to store the frequent pattern information in the dataset. Through the incremental window and newly discovered frequent patterns, new knowledge in the incremental dataset is mined, and the newly added frequent patterns are dynamically updated into the original frequent pattern tree, enabling the frequent pattern tree to continuously evolve as the dataset increases. The technical solution provided by the embodiment of the present invention can adapt to application scenarios with continuously expanding data volumes, solve the problems of high time complexity and high space complexity faced by incremental data calculation, and has strong applicability to application scenarios that require analysis of incremental data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a knowledge discovery method, and particularly to a knowledge discovery method and a discovery device on an incremental data set. Background Art

[0002] The Internet of Things, social networks, and the Internet continuously generate new data every moment, and this data needs to be analyzed in a timely manner to mine its time-sensitive value. With the exponential growth of the data volume, its sparsity has become increasingly significant, and newly emerging knowledge and event information are often submerged in a large amount of data. How to extract valuable information from it and discover various hidden potential association relationships between things, including causal relationships, co-variation relationships, coexistence relationships, etc., is a difficult problem in related knowledge discovery research.

[0003] Many researchers have adopted incremental computing to implement data analysis and mining for continuously growing data.

[0004] One type of algorithm is an incremental computing method based on a tree data structure, which mainly uses a tree structure to store the frequent patterns of new and old data and mine the association relationships between the frequent patterns, such as the FUP and FUP2 algorithms. When new data arrives, by adjusting the frequent pattern trees of the new and old data sets, the frequent patterns of the new and old data sets are changed. This algorithm constructs a transaction data set to record information such as the occurrence times of each data transaction. After obtaining the new data set, by calculating the frequent transaction items in the new and old data sets, new frequent patterns are obtained, and then the correlation analysis of the new frequent patterns is performed; the CanTree-Gtree algorithm discovers complete frequent item sets from real-time transactions based on a sliding window. This algorithm uses two tree data structures: CanTree and GTree. One is CanTree, which scans all transactions in the sliding window and uses it as the base tree. A new data structure called GTree (group tree) is used as the projection tree for each data item, and the projection tree is constructed by traversing each node using a top-down tree traversal method. However, the incremental computing method based on the tree data structure needs to scan all the data once for each determination of the frequent pattern. Since the new data is often a small amount of data relative to all the data, it is difficult to discover new frequent patterns using the incremental computing method based on the tree structure.

[0005] Another type is the computational method based on incremental learning. Machine learning and deep learning methods are mostly used for incremental learning, continuously learning new knowledge from new samples and being able to retain most of the knowledge learned previously. Such methods include: parallel incremental wESVM (weighted extreme support vector machine), which is used to combine input data with the original dataset for learning and training. This model can combine the knowledge from subsets of training data through simple matrix addition, enabling it to perform parallel incremental learning by combining the knowledge of data slices at each incremental stage; the ensemble learning method DTEL (diversity and transfer-based ensemble learning method), which uses each saved historical model as an initial model and trains it together with new data through transfer learning. However, the computational method based on incremental learning takes a long time in model training with a large amount of data, and if new data has new features, the model needs to be retrained. Therefore, its model construction and computational cost are generally high, and the requirements for data quality are also high, requiring a large time complexity and space complexity.

[0006] Therefore, in order to adapt to the application scenarios with continuously expanding data volume and solve the problems of high time complexity and high space complexity faced by incremental data calculation, the present invention proposes a knowledge discovery method and a discovery device for incremental datasets, designs a tree-shaped data structure that can evolve as the data volume continuously expands and an incremental window for recording newly added frequent patterns, maintains the frequent transaction items in the newly added dataset, and maintains the timeliness of incremental data calculation. During the data increment process, the present invention makes real-time adjustments to the newly added frequent patterns through the original data frequent pattern tree, the newly added frequent pattern tree, and the incremental sliding window, and timely mines the association relationship of transaction items between the original dataset and the new dataset. At the same time, the present invention solves the problem that the incremental algorithm based on the tree-shaped data structure needs to continuously scan the original data when adjusting the frequent pattern, greatly reducing the time complexity. Summary of the Invention

[0007] To solve the problem that it is difficult to timely and accurately discover new knowledge in the scenario of continuously increasing data volume, the present invention provides a knowledge discovery method for incremental datasets. This method uses the EFPT-IKD algorithm, designs a tree-shaped data structure - frequent pattern tree that can continuously evolve as the data volume continuously grows, sets an incremental window (IW) to discover newly added frequent transaction items. The frequent pattern tree is mainly used to store the frequent pattern information in the dataset. Through the incremental window and the newly discovered frequent patterns, new knowledge in the incremental dataset is mined, and the newly added frequent patterns are dynamically updated into the original frequent pattern tree, enabling the frequent pattern tree to continuously evolve as the dataset increases.

[0008] The technical solution adopted by the present invention is as follows:

[0009] A knowledge discovery method and device for incremental datasets, comprising the following steps:

[0010] A. Based on the frequent transaction item set DB_FI in the original dataset DB, construct the original dataset frequent pattern tree DB_FP-tree, and calculate the association rule set AR(DB_FP-Tree) in DB according to the minimum support min_conf. Let the total association rule set ARSET = AR(DB_FP-Tree), set the upper limit of the incremental window length to m, and initialize the incremental sliding window IW to be empty. At the same time, initialize the incremental frequent pattern tree Idb0_FP-tree.

[0011] B. When the data of the i-th incremental dataset Idb i arrives, store the current data incremental dataset Idb i in the incremental database IDB, initialize the frequent transaction set Idb i _FI of the incremental dataset, scan the data in Idb i , calculate the support of each data item I in Idb i , and perform different operations according to four cases classified by the frequency in the original data and the incremental data.

[0012] C. Append the primary key information (one or more fields in the table, whose values are used to uniquely identify a certain record in the table) of the current incremental data to the end of the queue in the incremental sliding window IW. At this time, the queue length Len(IW) in the incremental sliding window IW is incremented by 1. If Len(IW) <= m (the upper limit m of the incremental window), then dynamically update the frequent pattern tree Idb i-1 _FP-Tree. If Len(IW) > m (the upper limit m of the incremental window), then read the head information in the incremental sliding window IW, transfer the data in the incremental database IDB to the original database DB according to the primary key information, delete the head node of IW, and at the same time update the information of these data to the frequent pattern tree DB_FP-Tree of the original data, and update the node information involved in Idb i _FP-Tree.

[0013] D. After step B is completed, update the incremental frequent pattern tree Idb i _FP-tree according to the finally obtained incremental frequent transaction set Idb i-1 . When updating, sort the transactions in Idb i _FI in descending order of support, and scan the current incremental dataset Idb i again, and update the transaction information in Idb i _FI to Idb i-1 according to the method of constructing the frequent pattern tree.In the _FP-Tree, at this time, Idb i-1 The _FP-tree is updated to Idb i _FP-tree.

[0014] E. Based on the incremental frequent pattern tree Idb i The _FP-tree and the minimum support min_conf are used to calculate the set of association relationships AR(Idb i _FP-Tree) after the i-th data increment, and let the total set of association relationships ARSET = ARSET ∪ AR(Idb i _FP-Tree).

[0015] In step A, the original dataset frequent pattern tree DB_FP-tree is used to store the frequent pattern information in the original dataset; the total set of association relationships ARSET starts to store the association rules of the original data, and ARSET needs to be dynamically adjusted (taking the union) every time incremental data occurs; the upper limit m of the incremental window IW (Incremental Window) is determined by the incremental dataset in the specific situation; the incremental frequent pattern tree Idb0_FP-tree is initialized to a root node.

[0016] In step B, the transaction items in the original data (DB) and the new data (Idb) are divided into 4 cases according to the support degree during the incremental evolution of the dataset and relevant operations are performed:

[0017] i. The transaction item I is still frequent in the original dataset DB and the new dataset Idb i and add the transaction item I to the incremental frequent transaction set Idb i _FI.

[0018] ii. The transaction item I is not frequent in DB but is frequent in Idb i so it is regarded as a newly emerging frequent transaction item,

[0019] and add the transaction item I to Idb i _FI.

[0020] iii. The transaction item I is frequent in DB but not frequent in Idb i so it needs to be discussed in different cases to calculate the global support degree of I:

[0021]

[0022] Among them, Count(I, DB) represents the number of times the transaction item I appears in the original dataset DB, which is obtained by searching the DB_FP-Tree. Count(I, IDB) represents the number of times the transaction item I appears in the incremental database IDB, which is obtained by scanning the incremental frequent pattern tree Idb i _FP-Tree. Len(DB) and Len(IDB) represent the lengths of the dataset DB and IDB respectively. If the support of the transaction item I, support(I) >= min_sup, then add the transaction I to Idb i _FI; otherwise, discard I in DB_FI, delete the corresponding node in the DB_FP-tree, directly connect the parent node and the child node of the node, and at the same time process the association rule records in ARSET that contain the transaction item I, and delete the information of the transaction item I.

[0023] iv. The transaction item I is infrequent in both DB and Idb i and such transactions are discarded.

[0024] In step C, when Len(IW) <= m (the upper limit of the incremental window m), sort the transactions in the incremental frequent transaction set Idb i _FI in descending order of support, and scan the current incremental dataset Idb i again, and update the transaction information in Idb i _FI to the Idb i-1 _FP-Tree according to the method of constructing the frequent pattern tree; when Len(IW) > m, the specific operation of updating the frequent pattern tree Idb i _FP-Tree is to subtract the corresponding value from the count information. If the count is reduced to 0, then delete the node and connect the parent node and the child node of the node.

[0025] In step D, when the first increment occurs, it is necessary to update the initialized incremental frequent pattern tree Idb0_FP-tree and add its leaf nodes; when the i-th increment occurs, it is necessary to adjust the Idb i-1 _FP-tree, sort the transactions in the frequent transaction set Idb i _FI in descending order of support, and scan the current incremental dataset Idb i again, and update the transaction information in Idb i _FI to the Idb i-1 _FP-Tree. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the present invention, the following will briefly introduce the drawings required for the description of the invention content and embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 It is a diagram of four cases of incremental data.

[0028] Figure 2 It is a flowchart of the knowledge discovery method. Detailed implementation manners

[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.

[0030] We use the data of the appeal hotline in a certain province of China for testing. The experimental data is the appeal hotline data from July 11, 2016 to September 20, 2018, involving 64 major problem categories, with a total of 177,835 pieces of data.

[0031] We take the first 30,000 pieces of data in chronological order from all the data as the original data set DB. The average daily data of the appeal hotline in this province is 3,000 (i.e., incremental data). It is known that the subsequent incremental data includes the process of an emergency occurring, developing, breaking out, and disappearing. In the experiment, the minimum support min_sup = 5% and the minimum confidence min_conf = 20% are set.

[0032] Step 1, we preprocess the text data, segment the text based on the jieba library in python, then use the TF-IDF algorithm to calculate the weight of each word, arrange the weights of the keywords in each piece of data in descending order, and take the first 30 words as the keywords of this piece of data based on the average length of each text data. In all the data, each keyword is used as a transaction item, aiming to mine the association relationship between transaction items.

[0033] Step 2, based on the data in DB, we can make a schematic diagram of the association relationship. Each node in the diagram represents a transaction item, and the connection line between nodes indicates that the two transaction items have an association relationship that meets the minimum support min_sup and the minimum confidence min_conf.

[0034] Step 3, in each subsequent data increment process, we use the EFPT-IKD algorithm to mine the correlation between transaction items and record the evolution of the association rule set. The EFPT-IKD algorithm first constructs the DB_FP-tree based on the DB data. Then, each time there is an increment, it judges each transaction in the new dataset, divides it into four cases, performs corresponding operations in the algorithm for each case, records the evolution process of the Idb_FP-tree, and mines the association relationship between frequent transaction items based on the Idb_FP-tree and the transaction records in IW.

[0035] In the results obtained using the EFPT-IKD algorithm, the hollow nodes represent the frequent transaction items before the data increment, and the solid nodes represent the newly emerging frequent transaction items during this data increment process. The connections between the nodes indicate that there is an association relationship between two frequent transaction items that satisfies the minimum support min_sup and the minimum confidence min_conf.

Claims

1. A knowledge discovery method and device for incremental datasets, comprising the following parts: A. Construct an original dataset frequent pattern tree DB_FP-tree based on the frequent transaction item set DB_FI in the original dataset DB, and calculate the association rule set AR(DB_FP-Tree) in DB according to the minimum support min_conf. Let the total association relationship set ARSET = AR(DB_FP-Tree), initialize the incremental sliding window IW, set the upper limit of the window length to m, and initialize the frequent pattern tree Idb0_FP-tree for maintaining the incremental dataset; B. When the i-th incremental data set Idb i arrives, store the current data incremental data set Idb i in the incremental database IDB, initialize the frequent transaction set Idb i _FI of the incremental data set, scan the data in Idb i , calculate the support of each data item I in Idb i , and perform different operations according to the support in 4 cases. On the basis of B, update the Idb i _FP-tree to Idb i+1 _FP-tree; C. Append the primary key information of the current incremental data to the end of the queue within the incremental sliding window IW. At this time, the queue length Len(IW) within the incremental sliding window IW is incremented by 1. If Len(IW) > m, read the information at the head of the queue within the incremental sliding window IW, transfer the data in the incremental database IDB to the original database DB according to the primary key information, delete the head node of IW, and at the same time update the information of these data to the frequent pattern tree DB_FP-Tree of the original data and update Idb i The node information involved in the _FP-Tree, decrement the count information by 1. If the count is decremented to 0, delete the node, connect the parent node and child node of the node, and based on Idb i _FP-tree and min_conf calculate the association relationship set AR(Idbi_FP-Tree) after the i-th data increment, and let the total association relationship set ARSET = ARSET ∪ AR(Idbi_FP-Tree); D. After step B, according to the finally obtained incremental frequent transaction set Idb i _FI constructs / updates the incremental frequent pattern tree Idb i-1 _FP-tree. When updating, sort the transactions in Idb i _FI in descending order of the number of occurrences, and scan the incremental data set Idb again i , and update the transaction information in Idb i _FI to Idb according to the method of constructing the frequent pattern tree i-1 _FP-Tree. At this time, Idb i-1 _FP-tree is updated to Idb i _FP-tree.

2. The knowledge discovery method and discovery device for an incremental data set according to claim 1, wherein In step A, construct DB_FP-tree based on the data in DB, obtain the frequent transaction item set DB_FI in DB, and calculate the association rule set AR(DB) in DB according to the minimum support min_conf. Let the total association relationship set ARSET = AR(DB), and initialize the incremental window IW.

3. A knowledge discovery method and device for incremental data sets according to claim 1, characterized in that In step B, the transaction items in the original data DB and the new data Idb are divided into four cases according to the support degree during the incremental evolution of the dataset and relevant operations are performed: (1) Case1: Transaction I at this time is still frequent in DB + Idb i and add this Transaction I to Idb i _FI; (2) Case2: Transaction I is not frequent in the DB at this time, but it is frequent in Idb i Therefore, it is regarded as a newly emerging frequent transaction item, and this transaction I is added to Idb i _FI; (3) Case3: At this time, transaction I is frequent in the DB but not frequent in Idb i Therefore, it is necessary to discuss the situation separately and calculate the global support degree of I: Among them, Count(I, DB) represents the number of times the transaction item I appears in the original dataset DB, which is obtained by searching the DB_FP-Tree. Count(I, IDB) represents the number of times the transaction item I appears in the incremental database IDB, which is obtained by scanning the incremental frequent pattern tree Idb i _FP-Tree. Len(DB) and Len(IDB) respectively represent the lengths of the datasets DB and IDB. If the support of the transaction item I, support(I) >= min_sup, then the transaction I is added to Idb i _FI; otherwise, I is discarded in DB_FI, the corresponding node in the DB_FP-tree is deleted, the parent node and the child node of the node are directly connected, and at the same time, the record of the transaction I in the ARSET is deleted; (4) Case4: At this time, this transaction I is infrequent both in the DB and in Idb i and such transactions are discarded.

4. A knowledge discovery method and discovery device for an incremental data set according to claim 1, characterized in that, In the said step C, Idb i _FP-tree is updated to Idb i+1 _FP-tree. Based on Idb i _FP-tree and min_conf, calculate the association relationship set AR(db i ) of the i-th incremental data set, and let the total association relationship set ARSET = ARSET ∪ AR(db i ); when the first increment occurs, the first incremental frequent transaction set Idb1_FI constructs an incremental frequent pattern tree; When the i-th increment occurs, the incremental frequent pattern tree is dynamically adjusted.

5. A knowledge discovery method and discovery device for an incremental data set according to claim 1, characterized in that In the said step D, when the i-th increment occurs, if the queue length Len(IW) < m, then update the frequent pattern tree Idb i-1 _FP-Tree, and calculate the association rule set AR at this time, and merge it into ARSET; If the queue length Len(IW) = m, then after updating the data corresponding to the primary key information of the increment window IW to the DB, update the frequent pattern tree Idb i-1 _FP-Tree, and calculate the association rule set AR at this time, and merge it into ARSET.

Citation Information

Patent Citations

  • Astronmical spectral data correlation analysis and system based on constraint frequent mode

    CN101071078A

  • A weighted frequent pattern mining method for data streams based on sliding window

    CN102289507A