A Method for Mining Association Rules of Vertical Data Distribution Based on Weighted Arrays
Through the vertical data distribution association rule mining method based on weighted arrays, the problems of multiple intermediate result sets and large time overhead are solved, and efficient data mining and memory optimization are achieved.
Patent Information
- Application Number
- CN202310244705.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-03-09
AI Technical Summary
When the prior art mines data on large and dense databases with uneven distribution, there are many intermediate result sets, large time overhead, low mining efficiency, and redundant computing and high memory usage problems.
The vertical data distribution association rule mining method based on weighted arrays is adopted. By establishing a vertical data format data set, using the minimum support degree pruning and weighted array for pruning, the frequent item set is gradually obtained, reducing redundant calculations and memory usage.
It improves the efficiency of data mining, reduces the intermediate result set and time overhead, reduces the union and scan operations of candidate frequent item sets, and optimizes the use of memory space.
Smart Images

Figure CN116244354B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and specifically to a method for mining association rules of vertical data distribution based on a weighted array. Background Art
[0002] The dynamic weighted vertical algorithm introduces a dynamic weighted data processing method to the association rule algorithm of vertical data distribution for mining and processing data.
[0003] In the prior art, the invention patent with the application number: CN201711100787.9, titled "Fast Association Rule Mining Method for Dense Databases Based on Vertical Data Distribution", is disclosed, which makes up for the deficiency of the traditional vertical algorithm in mining large dense databases. The smaller the minimum support, the more obvious the advantages of the algorithm, the calculation is simpler, and the occupation of the support set of frequent itemsets in the memory space during the operation of the association rule mining algorithm based on data vertical distribution is reduced.
[0004] However, in the prior art, such as the mining method mentioned in the application number: CN201711100787.9, when mining data, the database and data items show uneven distribution, strong associated data items appear together, and the similarity between transactions is large and frequent, which affects the efficiency of data mining. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for mining association rules of vertical data distribution based on a weighted array to solve the problems of many intermediate result sets, large time overhead, and low mining efficiency in the prior art.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A method for mining association rules of vertical data distribution based on a weighted array includes:
[0008] Step 1: Establish a vertical data format data set and scan it. During the scanning process, minimum support pruning is performed to obtain frequent 1-itemsets.
[0009] Step 2: Calculate a weighted array on the frequent 1-itemsets to obtain a weighted array with a weight of 2. Use the weighted array to prune the frequent 1-itemsets to obtain refined frequent 1-itemsets.
[0010] Step 3: Perform intersection merging on the data items in the refined frequent 1-itemsets. During the merging process, minimum support pruning is used to obtain frequent 2-itemsets.
[0011] Step 4: Continue with weighted array loop mining until the number of refined frequent k-itemsets is 1 or after the intersection of the refined frequent k-itemsets, the number of frequent k + 1-itemsets is 0, and the loop terminates.
[0012] Step 5: Expand the reduced frequent k-itemsets to obtain the final frequent itemsets.
[0013] Further, the first step includes: first establishing a vertically formatted dataset, then scanning the dataset, and pruning during the scanning process using the minimum support threshold to obtain frequent 1-itemsets formed by data items that meet the minimum support.
[0014] Further, in the second step, a weighted array is used to prune data items, and some redundant frequent itemsets are deleted from the frequent 1-itemsets to obtain reduced frequent 1-itemsets.
[0015] Further, in the third step, 2-itemsets with support less than the minimum support are deleted during the merging process.
[0016] Further, the fourth step includes:
[0017] Performing weighted array mining on the frequent 2-itemsets to obtain a weighted array with a weight of 3, using the weighted array to prune data items, and deleting redundant frequent itemsets in the frequent 2-itemsets to obtain reduced frequent 2-itemsets;
[0018] Performing intersection merging on the data items in the reduced frequent 2-itemsets, deleting 3-itemsets with support less than the minimum support during the merging process, and using the minimum support for pruning to obtain frequent 3-itemsets;
[0019] Performing weighted array mining on the frequent 3-itemsets to obtain a weighted array with a weight of 4. Using the weighted array for pruning, deleting the redundant part in the frequent 3-itemsets to obtain reduced frequent 3-itemsets;
[0020] Continue to perform cyclic mining according to the steps of mining reduced frequent 2-itemsets and reduced frequent 3-itemsets until the number of reduced frequent k-itemsets is 1 or the number of frequent k + 1-itemsets obtained after intersecting the reduced frequent k-itemsets is 0, and the loop terminates.
[0021] Further, in the fifth step: According to the calculation method of equivalence classes, the weighted array is expanded in ascending order of weight in combination with the reduced frequent k-itemsets to finally obtain frequent k-itemsets.
[0022] Further, after obtaining the frequent k-itemsets in the fifth step, strong association rules are mined according to the minimum confidence threshold.
[0023] Advantages of the present invention:
[0024] 1. The association rule mining method of the present invention is based on the dynamic weighted vertical algorithm for association rules. After scanning the database, dynamic weighting processing is performed, and then calculations are carried out. At the same time, the intermediate result set is reduced, the union of candidate frequent item sets and the number of scan operations are decreased, and the mining efficiency of the algorithm is improved.
[0025] 2. The association rule mining method of the present invention solves the problem of partial Tidset duplication but non-duplicate data item sets generated by the algorithm's calculation sub-problems, reduces redundant calculation processes, reduces the occupation of memory space, improves the mining performance of the algorithm, reduces the intermediate result set, reduces time overhead, and improves mining efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention will be further described below with reference to the accompanying drawings.
[0027] Figure 1 is a schematic diagram of the vertical data distribution association rule mining method based on a weighted array according to the present invention;
[0028] Figure 2 is a schematic diagram of the database distribution according to the present invention;
[0029] Figure 3 is a schematic diagram of the vertical data distribution association rule mining process based on a weighted array according to the present invention;
[0030] Figure 4 is a schematic diagram of the structure of the mining process of the vertical data distribution association rule according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] A vertical data distribution association rule mining method based on a weighted array, as Figure 1 shown, includes:
[0033] S1: First, a vertical data format data set is established, and then the data set is scanned. During the scanning process, pruning is performed using the minimum support threshold to obtain a frequent 1-item set formed by data items that meet the minimum support.
[0034] S2: Based on the frequent 1-item set, a weighted array is calculated to obtain a weighted array with a weight of 2. The weighted array is used to prune the frequent 1-item set, and some redundant frequent item sets are deleted from the frequent 1-item set to obtain a refined frequent 1-item set.
[0035] S3: Perform intersection merging on the data items in the reduced frequent 1-itemsets. During the merging process, delete the 2-itemsets with support less than the minimum support, and use the minimum support pruning to obtain the frequent 2-itemsets.
[0036] S4: Perform weighted array mining on the frequent 2-itemsets to obtain a weighted array with a weight of 3. Use the weighted array to prune the data items and delete the redundant frequent item sets in the frequent 2-itemsets to obtain the reduced frequent 2-itemsets.
[0037] S5: Perform intersection merging on the data items in the reduced frequent 2-itemsets. During the merging process, delete the 3-itemsets with support less than the minimum support, and use the minimum support pruning to obtain the frequent 3-itemsets.
[0038] S6: Perform weighted array mining on the frequent 3-itemsets to obtain a weighted array with a weight of 4. Use the weighted array for pruning and delete the redundant part in the frequent 3-itemsets to obtain the reduced frequent 3-itemsets.
[0039] S7: According to the steps of mining the reduced frequent 2-itemsets and the reduced frequent 3-itemsets, continue the loop mining until the number of the reduced frequent k-itemsets is 1 or the number of the frequent k + 1-itemsets obtained after the intersection of the reduced frequent k-itemsets is 0, and the loop terminates.
[0040] S8: According to the calculation method of the equivalence class, expand the weighted array in ascending order of the weights, and combine it with the reduced frequent k-itemsets to finally obtain the final frequent item sets.
[0041] S9: For the final frequent items, mine the strong association rules according to the minimum confidence threshold.
[0042] The above mining method has the following effects:
[0043] The above vertical association rule mining algorithm based on the weighted array and the vertical distribution association rule mining algorithm obtain the same results. Moreover, the vertical association rule mining algorithm based on the weighted array uses the weighted array for pruning during the mining process, simplifies the complex calculation process in the vertical distribution association rule mining algorithm, reduces the local redundant frequent item sets in the calculation process, avoids the large memory space and time overhead caused by the calculation of the local redundant frequent item sets, reduces the calculation complexity of the intermediate process, and improves the efficiency of the mining process.
[0044] The relevant definitions in the above mining method are as follows:
[0045] 1. Establish a database D = {i1, i2, i3,..., i n}, where n is a natural number. The transaction T is represented as T = {t1, t2, t3,..., t m}, where m ≤ n and m is a natural number, representing the number of data items included in transaction T. Transaction T is an item set composed of a series of data items with the same unique identifier TID, and TID is a thread control symbol. Each transaction T is a subset of database D, that is If transaction T contains K item sets, it is called a K-item set. The association rule is expressed as of the logical implication relationship, where and X ∪ Y = Φ.
[0046] 2. Support It represents the percentage of the simultaneous occurrence of transaction X and transaction Y in database D.
[0047]
[0048] where P(X ∪ Y) is the probability of the simultaneous occurrence of transaction X and transaction Y in database D.
[0049] 3. Confidence It represents the percentage of the number of the simultaneous occurrence of transaction X and transaction Y in database D to the number of transactions containing transaction X, that is, the conditional probability of Y occurring on X.
[0050]
[0051] 4. Strong association rule: The association rule satisfies that the support is greater than the minimum support and its confidence is greater than the minimum confidence.
[0052] 5. Frequent item set: The frequency of a data item is the number of times the data item appears in the entire database D. If the support of an item set meets the minimum support threshold and the confidence meets the minimum confidence threshold, then the item set is called a frequent item set. The frequent K-item set is denoted as L k .
[0053] Mining association rules includes: Given a database D, the process of mining the rules that meet the minimum support threshold and the minimum confidence threshold in database D can be divided into two processes: one is to mine all frequent item sets that meet the minimum support; the other is to mine strong association rules that meet the minimum confidence from all frequent item sets. Among them, the first process is directly related to the mining efficiency of the algorithm.
[0054] 6. Weighted array (X1, X2, X3,..., X j ) = [1, 2, 3,..., n], indicating that the reTidset phenomenon appears in the database, and the Tidset of data items X1, X2, X3,..., X j is equal, and the weight value is the length of this array.
[0055] The reTidset phenomenon is as follows: When facing a database with a large similarity and frequent uneven data distribution among transactions, the vertical data distribution association rule algorithm will generate records with partially repeated Tidset but non-repeated data item sets when calculating the frequent item sets of sub-problems. We call this type of record the Tidset recurrence record, abbreviated as the reTidset phenomenon.
[0056] Use the weighted array vertical association rule algorithm to mine a vertical format database as Figure 2 shown, set the minimum support to 0.2, and the mining process is as Figure 3 shown. The specific steps of the mining method are as follows:
[0057] In the first step, scan the database D to mine the reduced frequent 1-item sets, prune the Tidset that does not meet the minimum support, and delete the data items I, J, M, N, O, P respectively. During this process, calculate the weighted array and find that the Tidset of data item A is {1, 2, 4, 5, 7, 8, 9, 10}, which is equal to the Tidset of C, and the Tidset of data item B is {1, 7}, which is equal to the Tidset of F, resulting in the reTidset phenomenon. Then establish a weighted array and fill in the data item sets (AC) and (BF), forming weighted arrays [A, C] and [B, F] with a weight ω = 2. Subsequently, delete the records corresponding to C and F in the records, and obtain the reduced frequent 1-item sets A, B, D, E, G.
[0058] In the second step, perform intersection merging on the reduced frequent 1-item sets and prune to mine the reduced frequent 2-item sets through the minimum support. Respectively obtain the data item sets (AB) = {1, 7}, (AD) = {4, 8}, (AE) = {1, 4, 5, 7, 9, 10}, (AG) = {2, 4, 5, 9, 10}, (BD) = {}, (BE) = {1, 7}, (BG) = {}, (DE) = {4}, (DG) = {4}, (EG) = {4, 5, 9, 10}. Delete the data item sets (BD), (BG), (DE), (DG) that do not meet the minimum support threshold of 0.2, and mine the frequent 2-item sets (AB), (AD), (AE), (AG), (BE), (EG).
[0059] In the third step, perform weighted array calculation on the frequent 2-item sets to mine the reduced frequent 2-item sets. Through calculation, it is found that the Tidset of the data item set (AB) is {1, 7}, which is equal to that of (BE). Weight (AB) and (BE) to obtain a weighted array [A, B, E] with a weight ω = 3. Subsequently, delete the (BE) item in the frequent 2-item sets to obtain the reduced frequent 2-item sets (AB), (AD), (AE), (AG), (EG).
[0060] In the third step, the intersection of the reduced frequent 2-itemsets is merged to obtain the frequent 3-itemsets (ABE) and (AEG). By calculation, it is found that the frequent 3-itemsets (ABE) = {1, 7} and (AEG) = {4, 5, 9, 10}, and their support degrees are both greater than the minimum support threshold, so no minimum support pruning is required.
[0061] In the fourth step, the weighted array calculation is performed on the frequent 3-itemsets. Since there is no phenomenon of reTidset, this step is skipped, and the obtained frequent 3-itemsets are all reduced frequent 3-itemsets.
[0062] In the fifth step, the intersection of the reduced frequent 3-itemsets is merged. It is found that the intersection result of (ABE) and (AEG) is (A), and there is no frequent 4-itemset, which meets the termination condition for cyclic mining of reduced frequent k-itemsets, and the frequent itemset mining ends here.
[0063] In the sixth step, referring to the weighted array, the reduced frequent 3-itemsets are expanded in an equivalent class manner. The weighted arrays in ascending order of weights are [AC], [BF], and [ABE]. According to the equivalent class expansion method, first consider the weighted array with w = 2. Replace itemset A with [AC] to obtain the frequent itemsets (ACBE) and (ACEG), replace itemset B with [BF] to obtain the frequent itemsets (ACBFE) and (ACEG). Then consider the weighted array with w = 3. Replace itemset (AB) with [ABE] to obtain the frequent 5-itemset (ACBFE) and the frequent 4-itemset (ACEG).
[0064] In the seventh step, strong association rules are mined. Taking the frequent 5-itemset as an example, let the minimum confidence min_conf = 0.8. The association rules generated by the frequent 5-itemset (ACBFE) with 1-consequent combinations are {ACBF} -> {E} = 1, {ACBE} -> {F} = 1, {ABFE} -> {C} = 1, {CBFE} -> {A} = 1.
[0065] The calculation of the association rule algorithm based on vertical data distribution is as follows:
[0066] The method of calculating support and generating candidate item sets is similar to the Apriori algorithm. The difference is that counting starts during the process of finding candidate sets. The horizontal distribution is a common format in databases, consisting of many transactions T, corresponding to a unique transaction identifier TID. Each transaction is composed of several data items. The vertical distribution of data is defined as follows: Database D consists of a series of data items, and each data item is composed of several transaction identifiers TID in which it appears in the corresponding transaction, that is, including all TIDs of the data item. The list composed of all transaction identifiers is called Tidset. For example, there are two lists containing TIDs, Tidset(A) = (1, 2, 4, 5, 7, 8, 9, 10) and Tidset(B) = (1, 7). The horizontal distribution and vertical distribution of database D are as Figure 2 shown:
[0067] The data format of the horizontal distribution is recorded according to the transaction TID, which conforms to the law of data generated in daily life. The data format of the vertical distribution is recorded according to Tidset, which is more convenient for data processing. The support of frequent item sets can be calculated by performing intersection operations on different Tidsets. When performing union processing on each data item in the vertical distribution, many data item sets composed of common prefixes will be generated. For the convenience of description and calculation, they are divided into different equivalence classes and regarded as different sub-problems for processing. The definition of the equivalence class is given below:
[0068] S a = [a] = {b[k - 1] ∈ L1 | a[1:k - 2] = b[1:k - 2]}
[0069] where a[1:k - 2] = b[1:k - 2] and a[1:k - 2] < b[1:k - 2]. X[i] appearing in the above formula represents the i-th item, and X[i:j] represents all items between the i-th item and the j-th item in this item set.
[0070] For example, let L = {AB, AC, AD, AE, BC, BD, CD, CE}, and the generated equivalence classes S A = [A] = {B, C, D, E}, S B = [B] = {C, D}, S C = [C] = {D, E}. The sets generated by connecting the above three equivalence classes S A 、S B 、S C are {AB, AC, AD, AE}, {BC, BD}, {CD, CE}, and these three item sets are independent of each other. Using equivalence classes can divide the search space into independent sub-spaces. In actual operations, equivalence classes with the same prefix are regarded as an independent sub-problem, and equivalence classes with different prefixes are regarded as each new problem to be solved.
[0071] Use the association rule mining algorithm with vertical data distribution to mine the data distribution as Figure 2 shown, set the minimum support to 0.2, and the mining process is as Figure 4 shown:
[0072] Scan the database to calculate the Tidset of all data items. During the first step, prune the Tidset that does not meet the minimum support, delete data items I, J, M, N, O, P. At the same time, perform the intersection operation on the data item sets A, B, C, D, E, F, G and their respective Tidsets according to the equivalence classes to obtain the frequent 2-itemsets. Divide the frequent 2-itemsets into S A , S B , S C , S E Four equivalence classes. In the second step, perform frequent 3-itemset mining. Perform the intersection operation on the data item sets and Tidsets in the four equivalence classes respectively, and at the same time extract the frequent 3-itemsets that meet the minsup threshold through the Tidset intersection operation, and so on until the frequent k-itemsets that meet the conditions are mined. The results that meet the minsup threshold finally mined from this database are [A, B, C, E, F] and [A, C, E, G].
[0073] The mining process of the vertical data distribution association rules makes full use of the advantages of the vertical data format, uses intersections to calculate the support of data item sets, and at the same time uses the concept of equivalence classes to transform large problems into several independent sub-problems, calculates the frequent item sets of each part of the sub-problems respectively, and finally obtains the global frequent item sets of the entire database by merging.
[0074] Since the intersection calculation has good performance, the mining of vertical data distribution association rules is faster than the Apriori algorithm. However, when this algorithm faces a non-uniform distribution database, for example, the similarity between transactions is high and appears frequently, a large number of frequent item sets will be generated in the intermediate process of equivalence class calculation. These intermediate process frequent item sets will bring high computational redundancy, and these local frequent item sets will occupy a large amount of memory space, resulting in a decline in algorithm performance.
[0075] The results obtained by the dynamic weighted vertical association rule mining algorithm are the same as those of the vertical distribution association rule mining algorithm. The dynamic weighted vertical association rule mining algorithm is more efficient than the vertical distribution association rule mining algorithm and the Apriori algorithm, and has less memory overhead than the vertical distribution association rule mining algorithm.
[0076] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0077] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for mining vertical data distribution association rules based on a weighted array, characterized in that, Including: Step 1: Establish a vertical data format dataset and scan it. During the scanning process, pruning is performed with the minimum support to obtain frequent 1-itemsets. Step 2: Calculate a weighted array on the frequent 1-itemsets to obtain a weighted array with a weight of 2. Use the weighted array to prune the frequent 1-itemsets to obtain reduced frequent 1-itemsets. Step 3: Perform intersection merging on the data items in the reduced frequent 1-itemsets. During the merging process, pruning is performed with the minimum support to obtain frequent 2-itemsets. Step 4: Continue with the weighted array loop mining until the number of reduced frequent k-itemsets is 1 or after the intersection of the reduced frequent k-itemsets, the number of frequent k+1-itemsets obtained is 0, and the loop terminates. Step 5: Expand the reduced frequent k-itemsets to obtain the final frequent itemsets. In step 2, the weighted array is used to prune the data items, and some redundant frequent itemsets are deleted from the frequent 1-itemsets to obtain reduced frequent 1-itemsets. In step 3, 2-itemsets with support less than the minimum support are deleted during the merging process. Step 4 includes: Perform weighted array mining on the frequent 2-itemsets to obtain a weighted array with a weight of 3. Use the weighted array to prune the data items and delete the redundant frequent itemsets in the frequent 2-itemsets to obtain reduced frequent 2-itemsets. Perform intersection merging on the data items in the reduced frequent 2-itemsets. During the merging process, 3-itemsets with support less than the minimum support are deleted, and pruning is performed with the minimum support to obtain frequent 3-itemsets. Perform weighted array mining on the frequent 3-itemsets to obtain a weighted array with a weight of 4. Use the weighted array for pruning and delete the redundant part in the frequent 3-itemsets to obtain reduced frequent 3-itemsets. According to the steps of mining reduced frequent 2-itemsets and reduced frequent 3-itemsets, continue with the loop mining until the number of the reduced frequent k-itemsets is 1 or after the intersection of the reduced frequent k-itemsets, the number of frequent k+1-itemsets obtained is 0, and the loop terminates.
2. The vertical data distribution association rule mining method based on a weighted array according to claim 1, characterized in that Step 1 includes: First, establish a vertical data format dataset, then scan the dataset. During the scanning process, pruning is performed using the minimum support threshold to obtain frequent 1-itemsets formed by data items that meet the minimum support.
3. A method for mining vertical data distribution association rules based on a weighted array according to claim 1, characterized in that, In step 5: According to the calculation method of equivalence classes, the weighted array is expanded in ascending order of weights in combination with the reduced frequent k-itemsets to finally obtain the frequent itemsets.
4. A method for mining vertical data distribution association rules based on a weighted array according to claim 3, characterized in that, For the frequent k-itemsets obtained in step 5, strong association rules are mined according to the minimum confidence threshold.
Citation Information
Patent Citations
Dense database quick association rule mining method based on vertical data distribution
CN107908711A
Method of item-all-weighted positive or negative association model mining between text terms and mining system applied to method
CN103955542A
Probability graph-based transformer state association rule mining method
CN106649479A