Mass data related high-utility item set mining method covering negative profit
Through the method of combining multi-level partitioning and pre-evaluation matrix, an efficient and relevant high-efficiency item set mining algorithm framework is built, which solves the efficiency and flexibility of the existing methods when processing massive data, and realizes effective coverage and efficient calculation of negative profit item sets.
Patent Information
- Application Number
- CN202510550861.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing high-efficiency item set mining methods have problems such as lack of combination integrity, low computing efficiency and insufficient flexibility in adapting to scenarios when processing massive data, especially when it contains negative profit item sets.
The multi-level partitioning method is used to divide the database into independent subpartitions, and an efficient and relevant high-efficiency item set mining algorithm framework is constructed through pre-evaluation of matrix and multi-dimensional pruning strategies. This framework supports distributed parallel computing, realizes efficient computing of large-scale data sets, and covers the relevant high-efficiency sets of negative profits.
A complete, efficient and flexible large-scale data processing framework is realized, which significantly reduces memory overhead and computing complexity, can effectively cover the set of high-efficiency items in actual business scenarios, and demonstrates excellent integration, efficiency and flexibility in testing on multiple real data sets.
Smart Images

Figure CN120067176A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and relates to an efficient high-utility itemset mining method related to massive data covering negative profits. Background Art
[0002] High-Utility Itemset Mining (HUIM) is an important branch in the field of data mining, aiming to discover combinations of goods with high profits or high values from transaction data, so as to provide more operational insights for decision-making. Different from traditional frequent itemset mining, HUIM not only considers the occurrence frequency of item sets, but also comprehensively evaluates utility values (such as profit) such as unit price and purchase quantity of goods, which better meets the actual needs of business scenarios. However, existing HUIM methods mainly focus on utility maximization, which may lead to relatively weak correlations in the extracted item sets, or even just accidentally occurring combinations. Such a pattern often misleads decision-making in actual decision-making scenarios.
[0003] To solve this problem, Correlated High-Utility Itemset Mining (CHUIM) has been proposed, aiming to mine item sets with high utility and strong correlations between goods. Although there are already a few methods for CHUIM, with the expansion of the database scale, the implementation based on tree structures and utility lists faces significant memory overhead problems. To address this challenge, researchers have proposed clustering-based techniques and compressed storage methods, but these methods rely on preprocessing steps to eliminate hopeless items through preset thresholds, which forms a computational bottleneck in scenarios where different thresholds need to be explored. In addition, existing methods are based on the maximum greedy strategy, defaulting to the assumption that all goods have positive utility values, which will cause high-utility commodity combinations with negative profits to be missed. The single correlation evaluation index also limits the flexibility of the algorithm in different application scenarios.
[0004] In summary, the existing technologies usually have three technical bottlenecks: integrity, efficiency, and flexibility. Summary of the Invention
[0005] Aiming at the three major technical bottlenecks of the existing Correlated High-Utility Itemset Mining (CHUIM) technology, namely the lack of combination integrity, low computational efficiency, and insufficient flexibility in adapting to scenarios, the present invention proposes a correlated high-utility itemset mining method covering negative profits.
[0006] The technical solution adopted by the present invention is: a method for mining correlated high-utility item sets from massive data covering negative profits, comprising the following steps: Step 1: Use a multi-level partitioning method to divide the original database into multiple independent sub-partitions to form non-overlapping computing units; Step 2: Screen out the candidate 1-item sets whose positive transaction weighted utility meets the minimum utility threshold, and sort these items that meet the threshold constraints in the predefined priority order; Step 3: Construct a two-dimensional pre-evaluation matrix with the sorted candidate items. This matrix stores the upper bounds of utility and support for 1-item sets and 2-item sets; Step 4: Construct the utility list for the candidate set containing positive utility items and negative utility items screened out in Step 2; Step 5: Read each partition from the external memory in turn, judge and prune the partition and the items contained in the partition through the first pruning strategy, and eliminate the partitions whose upper bound of utility is lower than the minimum utility threshold and the items in the partition; Define the items whose upper bound of utility meets the utility threshold as candidate items, construct the utility list for the candidate items in the partition, and record the support, utility, and remaining utility of the items; Step 6: In each partition processed in Step 5, perform depth-first search and item set expansion with the utility list of the candidate set constructed in Step 4 as the initialization; Step 7: Terminate the depth-first search when there are no new candidate item sets to expand, and output all relevant high-utility item sets that meet the conditions.
[0007] Preferably, in Step 1, the multi-level partitioning method is specifically as follows: (1) Sort the items in the transaction according to the predefined priority order of the items; (2) In the first-level partition, split each transaction in the database into a positive item sub-transaction set and a negative item sub-transaction set according to the positive and negative utilities of the items; The second-level partition recursively divides the transaction into sub-transaction sets containing the same prefix items; The third-level partition constructs a dual-index mechanism: one index is arranged in ascending order of item support within the partition, and the other index is arranged in ascending order of the upper bound of utility.
[0008] Preferably, the predefined priority order of the items specifically includes: Let the item universe , represent the set of positive utility items, represent the set of negative utility items; For any item , when using all-confidence or bond correlation index, the priority order of the items is defined as: The positive utility items are sorted in ascending order of positive transaction weighted utility, the negative utility items are sorted in ascending order of support, and the positive utility items always take precedence over the negative utility items.
[0009] Preferably, when using kulc coefficient as the correlation index, the priority order of the items is defined as: Regardless of the positive or negative item utility, they are all sorted in ascending order of support, but the positive utility items still maintain the priority status.
[0010] Preferably, in the third step, the pre-evaluation matrix stores the support and upper bound of utility of the item set in a two-dimensional form, expressed as: ; where , represent candidate items; represents the priority order; represents the upper bound of utility; represents the support.
[0011] Preferably, for the kulc indicator, the utility list constructed in the fourth step contains a support / item number field; for the all-confidence indicator, the utility list contains a maximum support field; for the bond indicator, a bit set structure is used to store the differential support set of the item set.
[0012] Preferably, in the fifth step, the first pruning strategy is a dual pruning mechanism based on the upper bound of utility, including: prefix partition-level pruning. For any item , if its upper bound of utility is lower than the utility threshold , then directly prune the prefix subspace corresponding to this item, that is, terminate the further expansion of all item sets with as the prefix; sub-item-level pruning. In the unpruned prefix subspace , perform a secondary screening on each item : if , then permanently remove the item and all its expanded item sets from the prefix subspace .
[0013] Preferably, in the sixth step, it specifically includes: defining PL as the set of utility lists of item sets that need to be judged and processed in the current iteration round; Traverse all item sets in the PL set in sequence. For one of the item sets , if its utility , then mark the item set as a highly relevant and highly utility item set; According to the second pruning strategy, for the item set , if the sum of its utility and the remaining utility is lower than the utility threshold , that is, , then prune the item set and all its expanded item sets; Combine the two item sets and in the current expansion set to generate a higher-order item set ; According to the third pruning strategy, query the upper bound of the utility of the itemset in the pre-evaluation matrix. If it is lower than the utility threshold then directly prune it; For the itemset generated by expanding according to the predefined item priority order , according to the fourth pruning strategy, if its relevance metric value , is the predefined relevance threshold, then prune this itemset and all its extended item sets; For the high-order itemset generated by the itemset and the itemset , according to the fifth pruning strategy, if its upper bound of relevance , is the predefined relevance threshold, then prune this itemset and all its extended item sets; otherwise, retain this itemset in the new extended candidate set for the next round of expansion.
[0014] Preferably, the upper bound of relevance includes bond coefficient, all-confidence coefficient; where all- confidence the upper bound of relevance of the coefficient is: ; where represents the support; bond the upper bound of relevance of the coefficient satisfies the relational expression: .
[0015] The beneficial effects of the present invention are reflected in the following aspects: (1) A complete, efficient and flexible large-scale data processing framework: The present invention first proposes an algorithm framework for supporting the mining of relevant high-utility item sets covering negative profits, which completely covers the high-utility item sets in actual business scenarios. By reconstructing the upper bound of utility and the relevance pruning mechanism, efficient calculation of large-scale data sets is achieved; (2) Multi-level partition calculation and resource optimization: A multi-level database partition method designed based on commodity attributes and transaction characteristics divides the original data set into independent calculation units and supports distributed parallel calculation. The item set calculation only needs to scan local data, greatly reducing the redundant overhead caused by global traversal; (3) An efficient pruning and pre-evaluation acceleration mechanism: Integrate five innovative pruning strategies and combine with the pre-evaluation matrix to effectively prune hopeless item sets; (4)Multi-dimensional Verification and Flexibility: In the comparative tests of multiple real datasets and synthetic datasets, this framework demonstrates three major advantages: integrity guarantee, efficiency advantage, and flexibility expansion. Description of the Drawings
[0016] Figure 1 Schematic diagram of the research content in the embodiments of the present invention; Figure 2 Schematic diagram of the technical architecture of the high-utility itemset mining method related to massive data covering negative profits proposed in the embodiments of the present invention; Figure 3 Schematic diagram of the multi-level partitioning method proposed by the present invention; Figure 4 Schematic diagram of the data structure of the pre-evaluation matrix PEM proposed by the present invention; Figure 5 Schematic diagram of the variant of the UList structure proposed by the present invention; Figure 6 Based on all-confidence Coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage; Figure 7 Based on bond Coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage; Figure 8 Based on kulc Coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage; Figure 9 Evaluation of the CHIN and FHN algorithms in dealing with negative utilities; where, (a) running time; (b) memory usage; (c) number of candidate item sets. Detailed Implementation Manner
[0017] To facilitate the understanding of the present invention, the present invention will be described in more detail below with reference to the drawings and specific embodiments. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive.
[0018] In a retail scenario, such as Figure 1As shown, the High Utility Itemset Mining (HUIM) algorithm can mine a large number of highly profitable product combinations similar to {steak, coffee, bread} (total utility = $28, correlation bond coefficient = 0.17). However, research has found that such item sets actually often belong to accidental combinations due to insufficient correlation between products. To solve this problem, the Correlated High Utility Itemset Mining (CHUIM) algorithm introduces a correlation threshold constraint of β≥0.4 as a screening mechanism, but this method has obvious defects: when the item set contains "loss baits" (such as bread), the traditional CHUIM method will wrongly eliminate these key combinations based on the "utility always positive" assumption. For example, the {steak, bread} combination creates a total profit utility of $38 and satisfies bond the association strength of the coefficient 0.4, but because bread is a negative utility product (unit profit = $-1), it is misjudged and deleted by the algorithm, resulting in the inability to identify such "losing to make a profit" business strategies. To solve such problems, the present invention proposes a method for mining correlated high utility item sets covering negative profits in retail data.
[0019] Embodiment 1 A method for mining correlated high utility item sets covering negative profits in retail data provided by the present invention has a technical architecture as Figure 2 shown, and specifically includes the following steps: Step 1: Initialization of multi-level partitioning of the database.
[0020] Using the multi-level partitioning method proposed by the present invention, the original database is divided into multiple independent sub-partitions to form non-overlapping computing units. To solve the efficiency bottleneck of the existing method, the present invention proposes an innovative multi-level partitioning method, as Figure 3 shown, and specifically includes the following processes: 1. Pre-definition of the priority order of items: In the search space containing positive and negative utility items, the present invention first defines the priority order of items . Let the item universe , where represents the set of positive utility items, and where represents the set of negative utility items. According to different correlation measurement indicators adopted, the priority order is specifically defined as follows: (1) Priority rule based on all-confidence / bond coefficient For any item , when using all-confidence or bond correlation indicators, the priority order is defined as: positive utility items are sorted in ascending order of positive transaction weighted utility, negative utility items are sorted in ascending order of support, and positive utility items always take precedence over negative utility items, as shown below: (1); Wherein: represents the positive transaction weight utility value of item ; represents the support (frequency of occurrence) of item .
[0021] (2) Based on the priority rule of kulc index, when using kulc coefficient as the correlation index, the priority order is defined as: regardless of the positive or negative utility of the item, sort in ascending order of support, but the positive-utility items still maintain the priority, as shown below: (2).
[0022] 2. Multi-level partitioning strategy In the preprocessing stage, first sort the items in the transaction according to the item priority defined by formulas (1) and (2). In the first-level partitioning, each transaction in the database is split into a positive-item sub-transaction set and a negative-item sub-transaction set based on the positive and negative utilities of the items, to achieve independent processing of positive- and negative-utility items. The second-level partitioning forms a fine-grained partitioning structure by recursively splitting the transaction into sub-transactions containing the same prefix items. The third-level partitioning innovatively constructs a dual-index mechanism: the first index is sorted in ascending order of item support within the partition, and the second index is sorted in ascending order of the upper bound of utility . This dual-index structure enables directly skipping the entire partition corresponding to the items that do not meet the threshold during partition loading, significantly reducing memory occupancy and computational overhead. This hierarchical strategy lays the foundation for the subsequent pruning strategy by avoiding the combined operations of invalid items. in ascending order.
[0023] Step 2: Candidate set generation and sorting.
[0024] Select the candidate items whose positive transaction weights satisfy the minimum utility threshold, and sort these items that meet the threshold according to the defined priority order ( ).
[0025] Step 3: Construct the pre-evaluation matrix (PEM).
[0026] Construct a two-dimensional pre-evaluation matrix PEM with the sorted candidate items. This matrix stores the upper bounds of utility and support of 1-item sets and 2-item sets.
[0027] To improve the computational efficiency, the present invention proposes a pre-evaluation matrix (PEM), whose data structure is as Figure 4 shown. This matrix stores the support degree of item sets and the upper bound of utility in a two-dimensional form, and its mathematical definition is: (3); where , represent candidate items.
[0028] The implementation process of its function is as follows: (1) Data preloading: At the initial stage of mining, scan the data set and calculate the support degrees of all 1-item sets and 2-item sets and the upper bound of utility, and fill them into the PEM matrix according to the corresponding subscripts.
[0029] (2) Fast upper bound determination: When generating a new item set (such as generating {a, b, c} from {a, b} and {a, c}), by querying the support degree of the {b, c} subset in the PEM and the upper bound of utility, obtain the corresponding support degree and upper bound of utility of this item set.
[0030] (3) Early pruning decision: If the upper bound of utility is lower than the preset threshold, directly discard this candidate item set and terminate the expansion of the current item set to avoid subsequent invalid calculations.
[0031] Step Four: Construct a utility list (UList) for the candidate item sets (including positive-utility items and negative-utility items) screened out in Step Two.
[0032] To adapt to the calculation of different correlation metrics, the present invention designs three variants of the UList data structure, as Figure 5 shown, and realizes efficient calculation through field embedding.
[0033] For kulc metric, a support degree / number of items (sup / item) field is added to the UList to directly record the support degree values of each item in the item set. This design avoids the operation of repeatedly scanning the entire database for calculating kulc in the traditional method. For example, when analyzing the {milk, bread} combination, the independent support degrees of the two can be directly obtained from this field, significantly reducing the calculation overhead.
[0034] For all-confidence metric, a maximum support degree ( maxsup ) field is introduced to store the maximum support degree value of all items in the item set. Taking the {steak, coffee} combination as an example, this field will automatically record the higher support degree value between steak and coffee, makingall-confidence The coefficient can be directly obtained through a single division operation, eliminating the redundant calculations that require multiple traversals of transaction data in the traditional method.
[0035] For bond the metrics, a bitset structure is adopted to store the differential support degree set of item sets ( dissup ). When expanding item sets (such as generating {a, b, c} from {a, b} and {a, c}), the dissup intersection of two parent item sets is quickly calculated through bit operations, reducing the bond computational time complexity of the coefficient from O(n) to O(1).
[0036] Through the above structural innovation, the correlation calculation process is simplified to a combination of field queries and basic operations, and it supports the rapid integration of newly added correlation metrics, effectively improving the computational efficiency and scalability of the algorithm.
[0037] Step Five: Partition iterative processing.
[0038] Each partition is sequentially read from the external memory, and the first pruning strategy is used to judge and prune the partition and the items contained in the partition, removing the partitions and the items in the partitions whose upper utility bounds are lower than the utility threshold. For the items whose upper utility bounds meet the utility threshold (becoming candidate items), a UList data structure is constructed for them to record parameters such as the support degree, utility, and remaining utility of the items.
[0039] The first pruning strategy is a dual pruning mechanism based on ptwu . Specifically, it includes: (1) Prefix partition-level pruning: For any item , if its upper utility bound is lower than the preset utility threshold , then the prefix subspace corresponding to this item is directly pruned, that is, the further expansion of all item sets with as the prefix (such as , , , etc.) is terminated.
[0040] (2) Sub-item-level pruning: In the unpruned prefix subspace , each item is secondarily screened. If , then item and all its extended item sets (such as ) are permanently removed from the prefix subspace .
[0041] Step Six: Depth-First Search (DFS) and item set expansion.
[0042] First, define PL as the set of utility lists of the item sets to be judged and processed in the current iteration. PL is initially set to the candidate item set utility list (UList) constructed in Step 4.
[0043] 1. CHUI candidate verification: Traverse all the item sets in PL one by one. For one of the item sets , if its utility , then mark the item set as a highly relevant and highly utility item set (CHUI), being the preset utility threshold.
[0044] 2. Expansion feasibility determination: According to the second pruning strategy, for the item set , if the sum of its utility and the remaining utility is lower than the utility threshold , that is, , then prune this item set and all its expansion item sets (all higher-order item sets with as the prefix).
[0045] 3. Item set combination generation: Combine two item sets and (both with the item set as the prefix) in the current expansion set to generate a higher-order item set .
[0046] 4. Pruning based on PEM: According to the third pruning strategy, query the upper bound of the utility of the item set in the PEM matrix . If it is lower than the utility threshold , then directly prune it.
[0047] 5. Relevance filtering: Includes upper bound pruning of relevance and joint upper bound pruning of relevance. Only retain the item sets that meet the relevance threshold to the new expansion candidate set for the next round of expansion.
[0048] (1) Upper bound pruning of relevance adopts the fourth pruning strategy, specifically: For the item set generated by expanding according to the predefined item priority order , if its relevance metric value , ( being the preset relevance threshold), then prune this item set and all its expansion item sets.
[0049] This pruning strategy calculates in real time during the item set expansion process, which includes bond coefficient,all- confidence Coefficient, kulc Coefficient and other indicators. If , it indicates that the correlation between the item set and its extended item sets does not meet the threshold constraint, then terminate the further extension of to avoid generating all item sets prefixed with .
[0050] (2) Pruning the upper bound of correlation adopts the fifth pruning strategy, specifically: for the high-order item set generated by the item set and the item set , define the upper bound of the correlation of its all-confidence coefficient as: (4); Among them, is the support of the item set. This upper bound predicts the upper bound of the coefficient of the item set by comparing the extreme values of the support of the sub-item set and the joint item set. all-confidence
[0051] For the upper bound of the correlation of the bond coefficient, it inherits the boundary property of formula (4) and satisfies the relationship: (5).
[0052] Therefore, the of formula (4) can be directly used as the upper bound of the correlation of the bond coefficient. Based on this property, if , the algorithm terminates the extension of the item set and its extended item sets.
[0053] Step Seven: Recursion Termination and Result Output.
[0054] Terminate the DFS when there are no new candidate item sets to extend, and output all CHUI sets that meet the conditions.
[0055] An overview of each pruning strategy is shown in Table 1.
[0056] Table 1 Summary and Overview of Pruning Strategies .
[0057] Example Two: Performance Evaluation Since the method for dealing with datasets containing negative utility is proposed for the first time in this invention, there is no existing algorithm for direct comparison. To evaluate its performance, this invention designs two comparative experiments: (1) Compare with existing high-utility itemset mining algorithms on a dataset containing only positive utilities; (2) Set the correlation threshold to 0 and compare with existing high-utility itemset mining algorithms (for handling negative utility data).
[0058] Experiment 1: Evaluation of the method of the present invention (CHIN algorithm) and existing methods (CHUIM algorithm) The present invention evaluates the performance of the CHIN algorithm and compares it with five latest relevant high-utility itemset mining algorithms. The comparative experiments are based on three widely used correlation metrics: kulc coefficient, all-confidenc e coefficient, and bon d coefficient. To ensure a fair comparison, the experiments are conducted only on datasets containing only positive utilities because the existing comparative algorithms do not support CHUIM with negative utilities.
[0059] Existing methods (CHUIM algorithm) include: FCHM: P. Fournier-Viger, Y. Zhang, J. C.-W. Lin, D.-T. Dinh, and H. Bac Le, “Mining correlated high-utility itemsets using various measures,” Logic Journal of the IGPL, vol. 28, no. 1, pp. 19–32, 2020; ECHUM: D. Dharavath Ramesh, K. K. Sethi, and A. Rathore, “Positive correlation based efficient high utility pattern mining approach,” in Proceedings of Sixteenth International Conference on Information Processing, Data Science and Computational Intelligence, 2022, p. 273; GMCHM: N. M. Hung, T. NT, and B. Vo, “A general method for mining high-utility itemsets with correlated measures,” Journal of Information and Telecommunication, vol. 5, no. 4, pp. 536–549, 2021。
[0060] Table 2 shows the performance of CHIN in terms of running time (RT), candidate number (CN), and peak memory usage (MU). The experiments were conducted on four datasets (Retail, Mushroom, Chainstore, and Chicage-Cremes, and the datasets were sourced from the SPMF open-source data mining platform http: / / www.philippe-fournier-viger.com / spmf / ), and the parameter , was set to 0.1% of the total utility of the corresponding dataset. The experimental results show that CHIN is significantly superior to the existing CHUIM algorithm under all three correlation metrics. For example, on the Retail dataset, the running time of CHIN under the kulc metric is 1.66 seconds, while that of ECHUM is 16.11 seconds. Meanwhile, the candidate number and memory consumption are significantly reduced. This benefits from its advanced multi-level partitioning and effective pruning strategies, which reduce the candidate generation and computational costs, thereby reducing the memory usage and I / O overhead.
[0061] Table 2: Evaluation of CHIN and CHUIM algorithms on real datasets .
[0062] As Figure 6 and Figure 7 shown, under the all-confidence and bond coefficient metrics, the running time and memory usage of all algorithms on the Retail dataset decrease significantly as the minimum utility threshold increases. For example, in Figure 7 (a), the running time of FCHM and CHIN shortens from dozens of seconds to a few seconds, while in Figure 7 (b), the memory usage drops from approximately one thousand MB to several hundred MB. At lower values (such as ), the memory overhead of GMCHM is relatively high, while CHIN and FCHM maintain higher efficiency. As Increases, the gap between algorithms narrows, but CHIN remains competitive in terms of running time and memory usage.
[0063] For kulc coefficient, Figure 8 shows a similar trend: both the running time and memory usage of CHIN and ECHUM decrease significantly as increases. However, CHIN is generally superior to ECHUM, especially at lower values, indicating its higher efficiency in processing large-scale databases. When is set to 0.1, 0.5, and 0.9, the experimental results always show similar trends. Increasing results in a decrease in the number of discovered item sets, thus reducing the computational cost, and CHIN shows advantages in most scenarios.
[0064] In summary, these experiments verify the robust performance advantages of CHIN proposed in the present invention under different correlation metrics, making it a flexible and efficient method for mining correlated high-utility item sets.
[0065] Experiment 2: Evaluation of CHIN and FHN The present invention studies the performance of CHIN and FHN (J. C.-W. Lin et al., “FHN: an efficient algorithm for mining high-utility itemsets with negative unit profits,” Knowledge-Based Systems, vol. 111, pp. 283–298, 2016.) through the correlation threshold β = 0. Under this condition, the correlation constraint is removed, allowing both algorithms to focus on mining high-utility item sets.
[0066] The present invention Figure 9 compares CHIN with the widely used FHN on the Mushroom_negative dataset. The results show that CHIN is superior to FHN in terms of running time, memory usage, and the number of candidate item sets. Specifically, due to its advanced partitioning method and pruning strategy, CHIN significantly reduces the running time and the number of generated candidate item sets, thus reducing the computational cost and memory consumption.
[0067] In addition, by processing prefix-based partitioning, CHIN achieves a lower I / O cost than FHN, which is particularly beneficial when dealing with large-scale datasets. These findings indicate that even when the correlation threshold is zero, CHIN still maintains a robust performance advantage over FHN, highlighting its efficiency in mining high-utility item sets with negative utilities.
Claims
1. A method for mining high-utility itemsets related to massive data covering negative profits, characterized in that: The following steps are involved: Step 1: Use a multi-level partitioning method to divide the original database into multiple independent sub-partitions to form non-overlapping computing units; Step 2: Screen out the candidate 1-item sets whose weighted utilities of positive transactions meet the minimum utility threshold, and sort these items that meet the threshold constraint according to a predefined priority order; Step 3: Use the sorted candidate items to construct a two-dimensional pre-evaluation matrix, which stores the utility upper bound and support upper bound of 1-item sets and 2-item sets; Step 4: Construct a utility list for the candidate set containing positive utility items and negative utility items selected in step 2; Step 5: Read each partition from the external memory in turn, and use the first pruning strategy to judge and prune the partitions and the items contained in the partitions, and remove the partitions and items in the partitions whose utility upper bounds are lower than the minimum utility threshold; define the items whose utility upper bounds meet the utility threshold as candidate items, build a utility list for the candidate items in the partition, and record the support, utility, and residual utility of the items; Step 6: In each partition processed in step 5, perform depth-first search and item set expansion using the candidate set utility list constructed in step 4 as the initialization; Step 7: When there are no new candidate item sets to expand, terminate the depth-first search and output all relevant high-utility item sets that meet the conditions.
2. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 1 is characterized in that: In step 1, the multi-level partitioning method is specifically as follows: (1) Sort the items in the transaction according to the predefined priority order of the items; (2) In the first-level partitioning, each transaction in the database is split into a positive sub-transaction set and a negative sub-transaction set based on the positive and negative utilities of the items. The second-level partitioning recursively divides the transaction into sub-transaction sets containing items with the same prefix. The third-level partitioning constructs a dual index mechanism: one index is sorted in ascending order by the support of the items in the partition, and the other index is sorted in ascending order by the upper bound of the utility.
3. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 2 is characterized in that: The priority order of the predefined items specifically includes: , represents the set of positive utility items, represents a set of negative utility items; for any item , when using all-confidence or bond When using the relevance index, the priority order of items is defined as follows: positive utility items are sorted in ascending order of positive transaction weighted utility, and negative utility items are sorted in ascending order of support, and positive utility items always take precedence over negative utility items.
4. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 3 is characterized in that: When using kulc When the coefficient is used as a relevance indicator, the priority order of the items is defined as follows: regardless of whether the item utility is positive or negative, they are sorted in ascending order of support, but items with positive utility still retain priority.
5. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 1, characterized in that: In step 3, the pre-evaluation matrix stores the support and utility upper bound of the item set in a two-dimensional form, which is expressed as: ; in, , represents a candidate; Indicates the order of priority; represents the upper bound of utility; Indicates support.
6. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 1, characterized in that: The utility list constructed in step 4 is kulc Indicators, the utility list contains support / number of items fields; for all-confidence Indicator, the utility list contains the maximum support field; for bond Indicator, uses a bit set structure to store the difference support set of item sets.
7. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 1, characterized in that: In step 5, the first pruning strategy is a double pruning mechanism based on the utility upper bound, including: prefix partition level pruning, for any item , if its utility upper bound Below the utility threshold , then directly prune the prefix subspace corresponding to the item , that is, to terminate the Further expansion of all item sets of the prefix; sub-item level pruning, in the unpruned prefix subspace For each item Perform secondary screening: If , then from the prefix subspace Permanently remove items from and all its extensions.
8. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 1, characterized in that: The step six specifically includes: defining PL as a utility list set of item sets that need to be judged and processed in the current iteration round; Traverse all item sets in the PL set in turn, and for one of the item sets If its effectiveness , then the marked itemset It is a set of highly relevant and high-utility items; According to the second pruning strategy, for the item set If its effectiveness Residual Utility The sum is below the utility threshold ,Right now , then prune the item set and all its extended item sets; Expand the two itemsets in the current set and Combining to generate higher-order itemsets ; According to the third pruning strategy, query the item set in the pre-evaluation matrix The utility upper bound of Then prune directly; For expanding the generated itemsets according to the predefined item priority order , according to the fourth pruning strategy, if its correlation metric value , is the preset correlation threshold, then the item set is pruned and all its extended item sets; For the item set and item sets Generated high-order itemsets , according to the fifth pruning strategy, if its correlation upper bound , is the preset correlation threshold, then the item set is pruned and all its expanded item sets; otherwise, the item set is retained in the new expanded candidate set for the next round of expansion.
9. The method for mining high-utility itemsets related to massive data covering negative profits according to claim 8, characterized in that: The upper bound of the correlation includes bond coefficient, all-confidence Coefficient; where all-confidence The upper bound of the correlation of the coefficient is: ; in, Indicate support; bond The upper bound of the correlation of the coefficients satisfies the relationship: 。
Citation Information
Patent Citations
High-utility item set mining method containing negative utility
CN110471960A
High-efficiency high-occupancy-ratio item set mining algorithm based on suffix division in mass data
CN114528332A
Periodic utility item set mining method, system and device under incremental database
CN118377815A
Pattern mining method, high-utility item-set mining method and relevant device
WO2018059298A1