A Temporal Constraint-Based Big Data Association Rule Mining Method
By setting time sliding windows and value strategies in big data, building a collection enumeration tree and pruning, the calculation complexity and important rule retention problems of association rule mining under timing constraints are solved, and important correlation rules are effectively mined.
Patent Information
- Application Number
- CN202210797772.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-06
AI Technical Summary
It is difficult for the prior art to efficiently mine important correlation rules in big data under timing constraints, and frequent item set mining algorithms cannot determine project combinations with higher importance.
By setting the time sliding window and value strategy, a collection enumeration tree is built and pruned, the data flow is processed using the time sliding window, element value is calculated, initial element value list and K collection value list are constructed, and the collection enumeration tree is pruned to retain important association rules.
It realizes the efficient mining of important correlation rules in big data under timing constraints, solves the problem of computing complexity, and retains higher value correlation rules.
Smart Images

Figure CN115033622B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data mining, and in particular to a method for mining big data association rules with time series constraints. Background Art
[0002] With the rapid development of information technology, the amount of data generated by all walks of life has increased explosively. However, on the contrary, people cannot refer to sufficient and valuable rules when predicting the industry prospects. Therefore, mining valuable information from massive data has become a current research hotspot. Data mining is an important part of the current research fields of artificial intelligence and databases. Among them, association rule mining is a major branch of the data mining field. Association rule mining aims to mine the potential relationships between different transactions and attributes.
[0003] Traditional association rule mining methods are mainly frequent itemset mining algorithms, such as the Apriori algorithm, etc. The purpose of the frequent itemset mining algorithm is to mine the item combinations that frequently appear in the database and use the frequently appearing item combinations as the association rules in the database. However, the frequent itemset mining algorithm can only find the frequently appearing item combinations and cannot determine the item combinations with higher importance.
[0004] In view of this, how to obtain more important association rules in the database has become an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] The present invention provides a method for mining big data association rules with time series constraints, aiming to (1) realize the mining of big data association rules under time series constraints; (2) retain important association rules in big data; (3) solve the problem of complex calculation in the current association rule pruning strategy.
[0006] To achieve the above object, a method for solving ultra-high definition post-production based on cloud technology resources provided by the present invention includes the following steps:
[0007] S1: Process the data stream to be mined by using a time sliding window to form a database;
[0008] S2: Scan the database, calculate the value of each element in the database, delete the elements with a value less than a predetermined value, and sort the remaining elements according to the value size;
[0009] S3: Scan the database again, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size;
[0010] S4: Construct an initial element value list, and construct a K-set value list in an iterative manner;
[0011] S5: Construct a set enumeration tree, prune the set enumeration tree using the value strategy method, and the finally enumerated set of element sequences is the association rules of the database.
[0012] As a further improvement method of the present invention:
[0013] In the step S1, the time sliding window is used to slide the data stream to be mined to form databases under different time sequence constraints, including
[0014] Set the size of the time sliding window based on time sequence constraints to W, and set the initial time of the time sequence constraint to t1. Input the continuous data stream into the time sliding window in sequence, where represents the transaction data at time t i In a specific embodiment of the present invention, the transaction data includes transaction data of q different elements Among them represents the transaction data item of element m i at time t q In the user transaction scenario, the elements include the type, cost, price, transaction quantity, profit, etc. of the commodity;
[0015] The time sliding window judges the time of the transaction data in the data stream D. If the time of the currently input transaction data item is the initial time t1, the transaction data corresponding to this time is stored in the time sliding window, and the transaction data of subsequent times is stored in the time sliding window in sequence until the time sliding window has no storage space. If there is no transaction data to be input in the data stream, the size of the time sliding window is automatically modified to the number of transaction data in the current window to obtain a full time sliding window; if the time of the currently input transaction data is not the initial time t1, the transaction data of this time is skipped;
[0016] Take the full time sliding window as the database with the initial time t1 of the time sequence and the time sequence range of W;
[0017] By setting different time sliding window sizes and initial times, several databases under different time sequence constraints are obtained, and then data association rules are mined under different time sequence constraint conditions.
[0018] In the step S1, the data in the database is normalized, including:
[0019] Normalize the transaction data items in the database:
[0020]
[0021] Among them:
[0022] x i Represents the transaction data item of element i;
[0023] x i,min Represents the minimum value of the transaction data items in element i;
[0024] x i,max Represents the maximum value of the transaction data items in element i;
[0025] x′ i Represents the normalized transaction data item.
[0026] Calculating the value of each element in the database in the S2 step includes:
[0027] Each element i has an external value ex(m i ), and the external value represents the importance of the element to the user. In the user transaction database, different elements can be different types of commodities, and the external value of the element is the profit of the commodity;
[0028] For the transaction data at different times in the database Transaction data The number of element m i in it is the internal value of element m i in the transaction data ; In the user transaction database, the number of element m i in the transaction data represents the purchase quantity of different commodities at time t i ;
[0029] Calculating the value of different elements in the database:
[0030]
[0031] Where:
[0032] W represents the set of transaction data in the database;
[0033] Delete the elements with the element value value(m i ) < minvalue, and sort the remaining elements according to the value size, where minvalue represents the preset minimum value.
[0034] In the S3 step, adjusting the order of elements in the database, the adjusted order of elements is the sorting order of elements based on the value size, including:
[0035] Scan the database again, adjust the order of elements in the database, and the adjusted order of elements is the sorting order of elements based on the value size, and the greater the element value, the more forward the sorting.
[0036] In the step S4, construct an initial element value list, including:
[0037] Construct an initial element set M = {m1, m2, …, m q}, and calculate the value of the initial element m k in the transaction data :
[0038]
[0039] Where:
[0040] represents the value of the element m k in the transaction data ;
[0041] ex(m k ) represents the external value of the element m k ;
[0042] represents the internal value of the element m k in the transaction data ;
[0043] And calculate the values of other elements except m k in the transaction data :
[0044]
[0045] Where:
[0046] represents the values of other elements except m k in the transaction data ;
[0047] ex(m -k ) represents the external value of other elements except m k ;
[0048] represents the internal values of other elements except m k in the transaction data ;
[0049] Construct the value list of the initial element m k :
[0050]
[0051] Where:
[0052] v mk,i represents the initial element mk In transaction data value;
[0053] Indicates the value of other elements except m k in the transaction data value;
[0054] Repeat the above steps to construct a value list for all elements in the initial element set M = {m1, m2, …, m q}, and sort the value list according to the element sorting order based on the value size.
[0055] The construction of the value list of the K set in the step S4 by an iterative method includes:
[0056] 1) Reconstruct the elements in the database into several element item sets, initialize K = 2, where the element item sets contain distinct elements, and the number of elements in each element item set is K, then the set of element item sets is M = {M1, M2, …, M i , …} = {(m i , m j ), …, (m p , m q )}, M o represents the i-th element item set in the set of element item sets, and the order of the elements in the element item set is strictly in accordance with the element sorting order based on the value size; the external value of the element item set is the sum of the external values of the elements in the element item set, and the internal value of the element item set is the sum of the internal values of the elements in the element item set;
[0057] 2) Calculate the value list of the element item sets in the set of element item sets:
[0058]
[0059] Among them:
[0060] represents the value of the element item set M o in the transaction data value;
[0061] represents the value of other element item sets except M i in the transaction data value;
[0062] 3) K = K + 1;
[0063] 4) Determine whether K is greater than q at this time. If K > q, end the iteration; otherwise, return to step 1), where q is the number of element categories in the transaction data.
[0064] In the S5 step, construct a set enumeration tree and prune the set enumeration tree using the value strategy method, including:
[0065] 1) Take the empty set as the root node of the set enumeration tree. According to the element sorting order based on value size, take the 6 initial elements with the largest value as the child nodes of the set enumeration tree; and initialize K = 2.
[0066] 2) Select an item set with the number of elements being K. If K - 1 prefix values in the selected item set are elements of the set enumeration tree, take the last element of the selected item set as the next layer node of the set enumeration tree.
[0067] 3) K = K + 1;
[0068] 4) Repeat steps 2) - 3) until no next layer node can be constructed.
[0069] 5) Traverse from the initial elements to the leaf nodes of the set enumeration tree. The traversal result is the constructed item set. If there exists an item set whose value in the value list is less than minvalue, and the values of other item sets except the current item set are greater than minvalue, then prune and delete this item set in the set enumeration tree.
[0070] 6) After pruning, traverse from the initial elements to the leaf nodes of the set enumeration tree. The traversal result is the association rule of the database.
[0071] Compared with the prior art, the present invention proposes a method for mining big data association rules with time series constraints. This technology has the following advantages:
[0072] First, in this solution, by setting a time sliding window based on time series constraints, the continuous data stream D is sequentially input into the time sliding window. The time sliding window judges the moment of the transaction data in the data stream D. If the moment of the currently input transaction data item is the initial moment t1, then store the transaction data corresponding to this moment into the time sliding window, and sequentially store the transaction data of subsequent moments into the time sliding window until the time sliding window has no storage space. If there is no transaction data to be input in the data stream, automatically modify the size of the time sliding window to the number of transaction data in the current window to obtain a full time sliding window. If the moment of the currently input transaction data is not the initial moment t1, then skip the transaction data at this moment; take the full time sliding window as the database with the time series initial moment t1 and the time series range of W. By setting different time sliding window sizes and initial moments, obtain several databases under different time series constraints, and then mine the data association rules under different time series constraint conditions, so as to realize the mining of big data association rules under time series constraint conditions.
[0073] Meanwhile, this solution determines the value of different elements according to the internal value and external value of the elements in the database. The external value represents the importance of the element to the user. In the user transaction database, different elements can be different types of goods. The external value of the element is the profit of the goods, and the number of elements in the transaction data is the internal value of the element in the transaction data. In the user transaction database, the number of elements in the transaction data represents the purchase quantity of different goods. Then, the value of the element is the external value of the element multiplied by the sum of the internal values of the element. The elements are sorted according to the value of the element, and an element value list is constructed. The values in the element value list include the value of the initial element in the transaction data and the value of other elements in the transaction data except the current initial element; and according to the construction process of the element value list, an iterative method is used to construct the value list of the element item set, that is, the elements in the database are reconstructed into several element item sets, where the element item set contains mutually distinct elements, and the number of elements in each element item set is K. Then, the set of element item sets is M = {M1, M2, …, M i , …} = {(m i , m j ), …, (m p , m q )}, M i represents the i-th element item set in the set of element item sets. The order of the elements in the element item set is strictly in accordance with the element sorting order based on the value size. The external value of the element item set is the sum of the external values of the elements in the element item set, and the internal value of the element item set is the sum of the internal values of the elements in the element item set. The constructed value list represents the importance of different element item sets.
[0074] According to the constructed value list, the empty set is used as the root node of the set enumeration tree. According to the element sorting order based on the value size, the 6 initial elements with the largest value are used as the child nodes of the set enumeration tree, and K = 2 is initialized; an element item set with the number of elements being K is selected. If K - 1 prefix values in the selected element item set are elements of the set enumeration tree, the last element of the selected element item set is used as the next layer node of the set enumeration tree; K = K + 1; repeat the above steps until the next layer node cannot be constructed; traverse from the initial element to the leaf node of the set enumeration tree, and the traversal result is the constructed element item set. If there is an element item set whose value in the value list is less than the pre-set minimum value minvalue, and the values of other element item sets except the current element item set are greater than minvalue, then the element item set is pruned and deleted in the set enumeration tree; after pruning, traverse from the initial element to the leaf node of the set enumeration tree, and the result of traversing to the leaf node is the association rule of the database, so as to retain the item set with the highest value as the association rule of the big data. Description of the Drawings
[0075] Figure 1 A flowchart showing a method for mining big data association rules with timing constraints provided by an embodiment of the present invention;
[0076] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0077] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0078] Embodiment 1:
[0079] S1: Use a time sliding window to process the data stream to be mined to form a database.
[0080] In the step S1, the time sliding window is used to perform sliding processing on the data stream to be mined to form databases under different timing constraints, including
[0081] Set the size of the time sliding window based on timing constraints to W, and set the initial time of the timing constraint to t1. Input the continuous data stream into the time sliding window in sequence, where represents the transaction data at time t i In a specific embodiment of the present invention, the transaction data includes transaction data of q different elements where represents the transaction data item of element m i at time t; in the user transaction scenario, the elements include the type, cost, price, transaction quantity, profit, etc. of the commodity; q The time sliding window judges the time of the transaction data in the data stream D. If the time of the currently input transaction data item is the initial time t1, the transaction data corresponding to this time
[0082] is stored in the time sliding window, and the transaction data of subsequent times is stored in the time sliding window in sequence until the time sliding window has no storage space. If there is no transaction data to be input in the data stream, the size of the time sliding window is automatically modified to the number of transaction data in the current window to obtain a full time sliding window; if the time of the currently input transaction data is not the initial time t1, the transaction data of this time is skipped; Take the full time sliding window as a database with the initial time t1 of the timing and a timing range of W;
[0083]
[0084] By setting different time sliding window sizes and initial times, several databases under different time series constraints are obtained, and data association rules are mined under different time series constraint conditions.
[0085] S2: Scan the database, calculate the value of each element in the database, delete the elements with values less than the predetermined value, and sort the remaining elements according to the value size.
[0086] Calculating the value of each element in the database in step S2 includes:
[0087] Each element i has an external value ex(m i ), and the external value represents the importance of the element to the user. In the user transaction database, different elements can be different types of commodities, and the external value of the element is the profit of the commodity;
[0088] For the transaction data at different times in the database Transaction data The number of elements m i in it is the internal value of element m i in the transaction data ; In the user transaction database, the number of element m i in the transaction data represents the purchase quantity of different commodities at time t i ;
[0089] Calculate the value of different elements in the database:
[0090]
[0091] where:
[0092] W represents the set of transaction data in the database;
[0093] Delete the elements with element value value(m i ) < minvalue, and sort the remaining elements according to the value size, where minvalue represents the preset minimum value.
[0094] S3: Scan the database again, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size.
[0095] Adjusting the order of the elements in the database in step S3, and the order of the elements after adjustment is the sorting order of the elements based on the value size, includes:
[0096] Scan the database again, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size, and the larger the element value, the more forward the sorting.
[0097] S4: Construct an initial element value list and construct a K-set value list in an iterative manner.
[0098] In the S4 step, constructing the initial element value list includes:
[0099] Construct an initial element set M = {m1, m2, …, m q}, and calculate the value of the initial element m k in the transaction data :
[0100]
[0101] Where:
[0102] represents the value of the element m k in the transaction data ;
[0103] ex(m k ) represents the external value of the element m k ;
[0104] represents the internal value of the element m k in the transaction data ;
[0105] And calculate the values of other elements except m k in the transaction data :
[0106]
[0107] Where:
[0108] represents the value of other elements except m k in the transaction data ;
[0109] ex(m -k ) represents the external value of other elements except m k ;
[0110] represents the internal value of other elements except m k in the transaction data ;
[0111] Construct the value list of the initial element m k :
[0112]
[0113] Wherein:
[0114] represents the initial element m k in the transaction data value;
[0115] represents the value of other elements except m k in the transaction data value;
[0116] Repeat the above steps to construct a value list for all elements in the initial element set M = {m1, m2,..., m q}, and sort the value list according to the element sorting order based on the value size.
[0117] The step S4 constructs the value list of the K set in an iterative manner, including:
[0118] 1) Reconstruct the elements in the database into several element item sets, initialize K = 2, where the element item sets contain distinct elements, and the number of elements in each element item set is K. Then the set of element item sets is M = {M1, M2,..., M i ,...} = {(m i , m j ),...,(m p , m q )}, M i represents the i-th element item set in the set of element item sets. The order of elements in the element item set is strictly in accordance with the element sorting order based on the value size; the external value of the element item set is the sum of the external values of the elements in the element item set, and the internal value of the element item set is the sum of the internal values of the elements in the element item set;
[0119] 2) Calculate the value list of the element item sets in the set of element item sets:
[0120]
[0121] Wherein:
[0122] represents the value of the element item set M i in the transaction data value;
[0123] represents the value of other element item sets except M i in the transaction data value;
[0124] 3) K = K + 1;
[0125] 4) Determine whether K is greater than q at this time. If K > q, end the iteration; otherwise, return to step 1), where q is the number of categories of elements in the transaction data.
[0126] S5: Construct a set enumeration tree, and use the value strategy method to prune the set enumeration tree. The finally enumerated set of element sequences is the association rule of the database.
[0127] In the S5 step of constructing the set enumeration tree and pruning the set enumeration tree using the value strategy method, it includes:
[0128] 1) Take the empty set as the root node of the set enumeration tree. According to the sorting order of elements based on value size, take the 6 initial elements with the largest value as the child nodes of the set enumeration tree; and initialize K = 2;
[0129] 2) Select an item set with the number of elements being K. If K - 1 prefix values in the selected item set are elements of the set enumeration tree, take the last element of the selected item set as the next layer node of the set enumeration tree;
[0130] 3) K = K + 1;
[0131] 4) Repeat steps 2) - 3) until the next layer node cannot be constructed;
[0132] 5) Traverse from the initial elements to the leaf nodes of the set enumeration tree. The traversal result is the constructed item set. If there exists an item set whose value in the value list is less than minvalue, and the values of other item sets except the current item set are greater than minvalue, then prune and delete this item set in the set enumeration tree;
[0133] 6) After pruning, traverse from the initial elements to the leaf nodes of the set enumeration tree. The traversal result is the association rule of the database.
[0134] Example 2:
[0135] This example is basically the same as Example 1, the difference is:
[0136] S1: Use a time sliding window to process the data stream to be mined to form a database.
[0137] In the S1 step, the normalization processing of the data in the database includes:
[0138] Perform normalization processing on the transaction data items in the database:
[0139]
[0140] Among them:
[0141] x i represents the transaction data item of element i;
[0142] x i,min represents the minimum value of the transaction data items in element i;
[0143] x i,max represents the maximum value of the transaction data items in element i;
[0144] x′ i represents the normalized transaction data item.
[0145] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. And the terms "including", "comprising" or any other variant thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including the element.
[0146] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0147] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for mining large data association rules with timing constraints, characterized in that The method includes: S1: Process the data stream to be mined using a time sliding window to form a database; Process the data stream to be mined using a time sliding window through sliding processing to form databases under different temporal constraints, including: Set the size of the time sliding window based on timing constraints to W, and set the initial time of the timing constraints to t1. Input the continuous data stream into the time sliding window in sequence, where represents the transaction data at time t i . The transaction data includes transaction data of q different elements where represents the transaction data item of the qth element m i at time t q . The time sliding window judges the moments of transaction data in the data stream D. If the moment of the currently input transaction data item is the initial moment t1, the transaction data corresponding to this moment is stored in the time sliding window, and then the transaction data of subsequent moments is stored in the time sliding window in sequence until there is no storage space in the time sliding window. If there is no transaction data to be input in the data stream, the size of the time sliding window is automatically modified to the number of transaction data in the current window, obtaining a full time sliding window; if the moment of the currently input transaction data is not the initial moment t1, the transaction data of this moment is skipped; Regarding the full time sliding window as the initial moment t1 of the time series, and the database with a time series range of W; By setting different time sliding window sizes and initial moments, obtain several databases under different temporal constraints, and then mine large data association rules under different temporal constraint conditions; S2: Scan the database, calculate the value of each element in the database, delete the elements with a value less than a predetermined value, and sort the remaining elements according to the value size; Calculate the value of each element in the database, including: Each element m q has an external value ex(m q ), and the external value represents the importance of the element to the user. In the user transaction database, different elements are different types of commodities, and the external value of the element is the profit of the commodity; For the transaction data at different times in the database Transaction data The number of element m in q is the internal value of element m q in the transaction data The internal value In the user transaction database, the number of element m q in the transaction data represents the purchase quantity of different commodities at time t i ; Calculate the values of different elements in the database: Where: W represents the set of transaction data in the database; Delete the elements with value value(m q ) that are less than minvalue, and sort the remaining elements by value, where minvalue represents a preset minimum value; S3: Scan the database again, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size; S4: Construct an initial element value list, and construct an H-set value list through an iterative method; S41: Construct an initial element set M = {m1, m2, …, m q}, and calculate the value of the initial element m k in the transaction data : Where: Represents element m k In the transaction data Value; ex(m k ) represents the external value of element m k ; Represents element m k In the transaction data Internal value; S42: Calculate the value of other elements except m k in the transaction data : Where: Denote the sum of the values of other elements except m k in the transaction data ; ex(m -k ) represents the external value of elements other than m k ; Indicates the internal value of other elements except m k in the transaction data ; S43: Construct the value list of the initial element m k as follows: Where: Represents the initial element m k In the transaction data Value of; Indicates the value of other elements except m k in the transaction data ; Repeat steps S41 - S43 to construct a value list for all elements in the initial element set M = {m1, m2, …, m q}, and sort the value list according to the sorting order of elements based on the value size; Construct an H-set value list through an iterative method, including: A1: Reconstruct the elements in the database into several element item sets, initialize H = 2, where the element item sets contain distinct elements, and the number of elements in each element item set is H, then the set of element item sets with the number of elements being H is M(H): Represents the r-th element item set in the set M(H) of element item sets. The order of elements in the element item set is strictly in accordance with the element sorting order based on value magnitude; the external value of an element item set is the sum of the external values of the elements within the element item set, and the internal value of an element item set is the sum of the internal values of the elements within the element item set; A2: Calculate the value list of the element item sets within the set of element item sets, where the element item set has the following value list: Where: Represents an itemset In transaction data Value; Indicates the value of other item sets in the set of item sets M(H) except in the transaction data ; A3: H = H + 1; A4: Determine whether H is greater than q at this time. If H > q, end the iteration; otherwise, return to step 1), where q is the number of categories of elements in the transaction data; S5: Construct a set enumeration tree, and prune the set enumeration tree using a value strategy method. The finally enumerated set of element sequences is the association rule of the database.
2. The method for mining large data association rules with timing constraints according to claim 1, characterized in that In step S1, perform normalization processing on the data in the database, including: Perform normalization processing on the transaction data items in the database: Where: x q represents the transaction data item of element m q ; x q,min Represents the minimum value of the transaction data item in element m q ; x q,max Represents the maximum value of the transaction data item in element m q ; x' q Represents the transaction data item after normalization processing.
3. A method for mining large data association rules with timing constraints according to claim 1, characterized in that In step S3, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size, including: Scan the database again, adjust the order of the elements in the database, and the order of the elements after adjustment is the sorting order of the elements based on the value size. The greater the element value, the more forward the sorting.
4. A method for mining big data association rules with timing constraints according to claim 1, characterized in that, In step S5, construct a set enumeration tree, and prune the set enumeration tree using a value strategy method, including: 1) Take the empty set as the root node of the set enumeration tree. According to the element sorting order based on value size, take the 6 initial elements with the largest value as the child nodes of the set enumeration tree; and initialize N = 2; 2) Select an element item set with the number of elements being N. If N - 1 prefix values in the selected element item set are elements of the set enumeration tree, then use the last element of the selected element item set as the next layer node of the set enumeration tree; 3) N = N + 1; 4) Repeat steps 2) - 3) until no next layer node can be constructed; 5) Traverse from the initial element to the leaf nodes of the set enumeration tree. The traversal result is the constructed element item set. If there exists an element item set whose value in the value list is less than minvalue, and the values of other element item sets except the current element item set are greater than minvalue, then prune and delete this element item set in the set enumeration tree; 6) After pruning is completed, traverse from the initial element to the leaf nodes of the set enumeration tree, and the traversal result is the association rule of the database.
Citation Information
Patent Citations
Keyword search KSAARM algorithm combining time window and association rules mining
CN109783628A
Association rule mining algorithm based on Boolean matrix reduction
CN111625574A