Association rule mining method based on user definition

Through user-defined association rule mining methods, combined with binary encoding and Apriori algorithm, heuristic optimization algorithm and value functions are used for directional optimization, solving the problems of low rules quality and waste of computing resources in traditional methods, and achieving efficient and good quality association rule mining.

CN120030066APending Publication Date: 2025-05-23NANTONG MARINE ADVANCED RESEARCH INSTITUTE SOUTHEAST UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411861345.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional association rule mining methods cannot optimize rules based on user needs, resulting in low rules quality and waste of computing resources.

Method used

The user-defined association rule mining method is used to recode the data set in binary, and the initial rules are obtained using the Apriori algorithm. The user sets the terms of interest and not interest, and combines the heuristic optimization algorithm and the value function for directional optimization.

Benefits of technology

It improves the quality and efficiency of association rules, increases the proportion of rules that users are interested in, and reduces the proportion of rules that are not interested in, and is suitable for any text data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030066A_ABST
    Figure CN120030066A_ABST
Patent Text Reader

Abstract

The invention relates to an association rule mining method based on user definition. The method is suitable for all text data sets. According to the method, a user-defined item is closely associated with the preference of a user; the method comprises the following steps: firstly, carrying out binary coding on a data set, setting minimum support degree and confidence degree by a user after coding, and carrying out association rule mining on the data set through an Apriori algorithm to obtain an initial association rule; a user defines interested items and uninterested items according to own preferences, then maps an initial association rule into a matrix according to interested and uninterested principles, and optimizes the initial association rule by using an improved marine predator algorithm; according to the improved marine predator algorithm, a value function is introduced to calculate the value of the rule, and the value function is closely related to a user-defined item, so that directional optimization is realized, and a more satisfactory association rule of the user is obtained. The method has the characteristics of high speed, high stability and good generalization ability, can provide reference for multiple industries, and is simple to operate and convenient to use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an association rule mining method, in particular to a method capable of efficiently mining association rules, specifically to a user-defined association rule mining method. Background Art

[0002] Most traditional association mining algorithms optimize the Apriori algorithm and the FP-growth algorithm to reduce the number of times the algorithm scans the entire text to improve efficiency; or use optimization algorithms to optimize the initial association rules. These optimization algorithms optimize the rules as a whole and indiscriminately delete association rules with low support and confidence to improve efficiency. Although these methods can save time for rule mining to varying degrees, they cannot optimize the rules in the direction required by users, which ultimately leads to low overall quality of the rules.

[0003] Therefore, the user-defined association rule mining method proposed in this patent is aimed at the shortcomings of the above traditional association rule methods, so that the association rules can be optimized in the direction specified by the user during the optimization process, and more association rules that the user is interested in and fewer association rules that the user is not interested in are obtained, thereby improving the mining efficiency and quality of association rules. Summary of the invention

[0004] The purpose of the present invention is to provide a user-defined association rule mining method to address the deficiencies of the prior art, which can avoid the waste of computing resources and obtain association rules of high quality. Moreover, the method can be applied to any text data set.

[0005] The technical solution of the present invention is:

[0006] A method for mining association rules based on user customization is characterized by binary recoding of a data set used for mining association rules; using an Apriori algorithm to mine the data set to obtain initial association rules; the user sets his or her own items of interest and uninterest before optimizing the rules; optimizing the initial association rules by a heuristic optimization algorithm, while adding a value function to guide the optimization direction to be consistent with the user's expected direction; obtaining a final association rule set, in which the proportion of association rules that the user is interested in increases, and the proportion of association rules that the user is not interested in decreases.

[0007] The method comprises the following steps:

[0008] 1) Obtain the data set that needs to be mined for association rules and recode the data set into binary form ’ ; The data set format is as follows:

[0009] T1: 1,3,4,5,7

[0010] T2: 1, 2, 3, 5, 6, 7, 8

[0011] T3: 1, 2, 4, 6, 8

[0012] …

[0013] Tm:2,4,6,8

[0014] The data set consists of m transaction sets; T* represents the *th transaction set; 1, 2, ..., 8 represent items; the data set is encoded into binary, and the binary data set is as follows:

[0015] T1:10111010

[0016] T2: 11101111

[0017] T3:11010101

[0018] …

[0019] Tm:01010101

[0020] 2) Set the minimum support and minimum confidence, perform association rule mining on the data set based on the Apriori algorithm, and obtain the initial association rules.

[0021]

[0022] Among them, Support represents support, Confidence represents confidence; X and Y represent items; m represents the number of transaction sets in the data set; |*| represents the frequency of occurrence of the * item.

[0023] 3) The user sets the item set of interest I = {i 1 ,i 2 ,…,i n} and the set of items of no interest U=={u 1 ,u 2 ,…,u n}.

[0024] 4) Map the initial association rules to the matrix according to the items of interest and disinterest set by the user, in preparation for the next step of using the heuristic optimization algorithm. The mapping satisfies the following conditions: if it is an item of interest, it is represented by 3 in the matrix, the item of disinterest is represented by 0, the neutral item is represented by 2, and the item not in the transaction is represented by 1.

[0025] 5) An improved marine predator algorithm is introduced to perform targeted optimization of association rules. A value function is introduced into the improved marine predator algorithm to calculate the value of the rules, so that the optimization direction is closely related to user interests.

[0026] 6) Obtain the final set of association rules, among which the proportion of association rules that users are interested in (P ir ) increases, and the proportion of association rules that users are not interested in (P ur )reduce.

[0027]

[0028] Among them, N ir 、N ur and N r They represent the number of interesting rules, the number of uninteresting rules, and the total number of optimized rules, respectively.

[0029] Furthermore, the data set should be a text data set; the minimum support, minimum confidence, items of interest and items of no interest are set by the user; the value function in the heuristic optimization algorithm is related to the support, confidence and association rule length.

[0030] Beneficial effects of the present invention:

[0031] The present invention recodes the data set through binary coding, thereby improving the efficiency of computer operations. The Apriori algorithm is used to mine the initial association rules of the data set to obtain a large number of association rules, many of which are irrelevant to what the user wants, resulting in rule redundancy. At this time, the user customizes some items of interest and disinterest, and then uses the optimized marine predator algorithm to optimize the rules. The value function is introduced into the marine predator algorithm to achieve directional optimization, and the direction of optimization is closely related to the user-defined items of interest and disinterest. In the end, the user can obtain a higher proportion of interested association rules and a lower proportion of uninterested association rules. The present invention can be used in any text data set, covering a wide range of fields, including: supermarket sales, infectious disease transmission laws, traffic accidents, etc. The method also has the characteristics of high efficiency and strong stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flow chart of the present invention.

[0033] Figure 2 Schematic diagram of binary encoding of data set in the present invention.

[0034] Figure 3 It is a schematic diagram of mapping items of interest and items of no interest in the present invention.

[0035] Figure 4It is a simulation test result diagram of the method of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0037] The present invention provides a method for mining association rules based on user-defined rules. Figure 1 As shown in Figure 2, this method is applied to mining hidden association rules in text data sets and can be closely related to user needs. Figure 2 The binary code is recoded by the Apriori algorithm; the initial association rules are obtained by mining the data set, such as Figure 1 As shown in , the user should set the minimum support and minimum confidence at this time; the user sets his own items of interest and items of disinterest before rule optimization; after the user sets the items of interest and items of disinterest, as shown in Figure 3 As shown in , the rules are mapped into the matrix to prepare for the next step of optimization. Figure 1 As shown in , the initial association rules are optimized by the heuristic optimization algorithm, and the value function is added to guide the optimization direction to be consistent with the user's expected direction; finally, the final association rule set is obtained, as shown in Figure 4 As shown in the figure, in the final association rule set, the proportion of association rules that users are interested in increases, while the proportion of association rules that users are not interested in decreases.

[0038] In this embodiment, taking the public data set "Connect" as an example, the data set has a total of 67,577 transactions (T), and the largest item is 127.

[0039] The method comprises the following steps:

[0040] 1) Binary encode the data set, as follows Figure 2 For example, in this data set, the first transaction set is

[0041] (1,4,7,10,13,16,19,22,25,28,31,34,37,40,43,46,49,52,55,58,61,64,67,70,73,76,79,82,85,88,91,94,97,100,103,106,109,112,115,118,121,124,127), re-encoded by binary, the encoded binary is

[0042] (1001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001001).

[0043] 2) Set the minimum support and minimum confidence. In this example, the minimum support and minimum confidence are set to 0.97 and 0.97 respectively.

[0044]

[0045] 8092 association rules were obtained, which are Figure 1 At the same time, the support and confidence values ​​of each rule and the length of the rule (the number of items included) can also be obtained.

[0046] 3) The user customizes the items of interest and the items of no interest according to his / her own preferences. In this example, the set of items of interest I={19,55,109} and the set of items of no interest U=={37,75,91} are set.

[0047] 4) According to the items of interest and disinterest set by the user, such as Figure 3 As shown, the rules are mapped to the matrix to prepare for the next step of optimization. For example, in this example, through step 2, an association rule ['109', '72', '75']=>['127', '91'] can be mined from the data set. Figure 3 The mapping principle maps it to the matrix as p=[1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1 ,1,1,1,1,1,1,1,1,1,2,1,1,0,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,0,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,3,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,2] T

[0048] 5) Map all the rules to get the matrix M 127×8092 , the matrix is ​​input into the improved marine predator algorithm to optimize the rules, such as Figure 1 As shown, after the optimization of the marine predator algorithm, the information is updated and its value is calculated using the value function. The value function is designed in this example as:

[0049] Val(R)=Conf(R)×lg(1+Z(R))

[0050] Among them, Conf(R) represents the confidence of the association rule, which can be obtained through step 2, and Z(R) is as follows:

[0051] Z(R)=o(R)×Sup(R)×Length(R)

[0052] Among them, Sup(R) represents the support of the association rule, Length(R) represents the length of the association rule (i.e. the number of items), both of which can be obtained through step 2, o(R) is as follows

[0053]

[0054] Among them, Num_I(R) represents the number of interesting items in the association rule, and similarly Num_U(R) represents the number of uninteresting items in the association rule; λ is a random number between 0 and 0.5; Max_Length represents the rule with the most items in the initial association rule, and Max_Length in this example is 6.

[0055] After the first iteration, the value obtained is used as the threshold. In subsequent iterations, each rule that is higher than the threshold will be stored in the rule set until the value obtained by the iteration is lower than the threshold or the maximum number of iterations is reached. Then the optimization stops and the output rule set is output. This set is the final rule set.

[0056] 6) Obtain the final set of association rules, among which the proportion of association rules that users are interested in (P ir ) increases, and the proportion of association rules that users are not interested in (P ur )reduce.

[0057]

[0058] Among them, N ir 、N ur and N r They represent the number of interesting rules, the number of uninteresting rules, and the total number of optimized rules, respectively.

[0059] In this example, the results are Figure 4 As shown, at the beginning, the proportion of interested rules and the proportion of uninterested rules were about 14% and 20% respectively. After iterations, the proportion of interested rules and the proportion of uninterested rules reached about 22% and 13% respectively.

[0060] The present invention has a simple structure and strong adaptability, can optimize association rules according to user preferences to improve rule quality and efficiency, and also has the following characteristics:

[0061] (1) Simple operation, easy to use, strong stability, and suitable for various text data sets.

[0062] (2) The mining speed is fast and the efficiency is high, and the quality of the mined rules better meets the needs of users.

[0063] The parts not involved in the present invention are the same as the prior art or can be implemented by using the prior art.

Claims

1. A method for mining association rules based on user-defined rules, characterized in that: The following steps are involved: S1, binary recoding of the data set used for association rule mining; S2, use Apriori algorithm to mine the data set to obtain initial association rules; S3, the user sets his / her own items of interest and disinterest before rule optimization; S4, optimize the initial association rules through the improved marine predator algorithm, and add the value function to guide the optimization direction to be consistent with the user's expected direction; S5. A final association rule set is obtained. In the final association rule set, the proportion of association rules that the user is interested in increases, and the proportion of association rules that the user is not interested in decreases.

2. The method for mining association rules based on user-defined rules according to claim 1, characterized in that: In step S1, the data set format is as follows: T1:1,3,4,5,7 T2:1,2,3,5,6,7,8 T3:1,2,4,6,8 …… Tm:2,4,6,8 The data set consists of m transaction sets; T* represents the *th transaction set; 1, 2, ..., 8 represent items; the data set is encoded into binary, and the binary data set is as follows: T1:10111010 T2:11101111 T3:11010101 …… Tm:01010101.

3. The method for mining association rules based on user-defined rules according to claim 2, characterized in that: In step S2, the minimum support and the minimum confidence are set, and the association rules of the data set are mined based on the Apriori algorithm to obtain the initial association rules, which are specifically: Among them, Support represents support, Confidence represents confidence; X and Y represent items; m represents the number of transaction sets in the data set; |*| represents the frequency of occurrence of the * item.

4. The method for mining association rules based on user-defined rules according to claim 3, characterized in that: In step S3, the user sets the item set of interest I = {i1, i2, ..., i n } and the set of uninteresting items U=={u1,u2,…,u n }; Map the initial association rules to the matrix according to the items of interest and disinterest set by the user, in preparation for the use of the heuristic optimization algorithm; the mapping satisfies the following conditions: if it is an item of interest, it is represented by 3 in the matrix, the item of disinterest is represented by 0, the neutral item is represented by 2, and the item not in the transaction is represented by 1.

5. The method for mining association rules based on user-defined rules according to claim 4, characterized in that: In step S4, the improved marine predator algorithm is introduced to perform directional optimization on the initial association rules. The value function is introduced into the improved marine predator algorithm to calculate the value of the rules, so that the optimization direction is closely related to the user's interest; wherein the value function is: Val(R)=Conf(R)×lg(1+Z(R)) Among them, Conf(R) represents the confidence of the association rule, and Z(R) is as follows: Z(R)=o(R)×Sup(R)×Length(R) Among them, Sup(R) represents the support of the association rule; Length(R) represents the length of the association rule, that is, the number of items; o(R) is as follows Among them, Num_I(R) represents the number of interesting items in the association rule, Num_U(R) represents the number of uninteresting items in the association rule; λ is a random number between 0 and 0.5; Max_Length represents the rule with the most items in the original rule; After the first iteration, the obtained value is used as the threshold; when entering the subsequent iterations, each rule that is higher than the threshold will be stored in the rule set until the value obtained from the iteration is lower than the threshold or the maximum number of iterations is reached, then the optimization of the output rule set is stopped. This set is the final rule set.

6. The method for mining association rules based on user definition according to claim 5, characterized in that: In step S5, a final set of association rules is obtained, in which the proportion of association rules that the user is interested in increases, and the proportion of association rules that the user is not interested in decreases; Among them, N ir 、N ur and N r Respectively represent the number of interesting rules, the number of uninteresting rules and the total number of optimized rules, P ir Indicates the proportion of association rules that users are interested in, P ur Indicates the percentage of association rules that users are not interested in.

7. The method for mining association rules based on user-defined rules according to claim 1, characterized in that: The data set is a text data set; the minimum support, minimum confidence, items of interest and items of no interest are set by the user; the value function in the improved marine predator algorithm is related to the support, confidence and association rule length.