An Efficient Utility Itemset Mining Method for Massive Data Covering Negative Profits

The method addresses the limitations of existing high-utility itemset mining by employing multi-level database partitioning and pruning strategies to efficiently identify and cover negative profit items, enhancing computational efficiency and flexibility in large-scale data processing.

CN120067176BActive Publication Date: 2025-07-15HARBIN INST OF TECH AT WEIHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510550861.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-15
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing high-efficiency item set mining methods have problems such as large memory overhead, low computing efficiency and insufficient flexibility in adapting to scenarios when dealing with large-scale databases. Especially when dealing with negative profit product combinations, important combinations are easily missed. The existing methods are based on a single correlation evaluation indicator, resulting in insufficient flexibility in algorithms in different application scenarios.

Method used

The multi-level partitioning method is used to divide the database into independent subpartitions, and a two-dimensional pre-evaluation matrix and multi-level pruning strategy are constructed. Combined with depth-first search, the calculation process is optimized through predefined priority order and multi-dimensional pruning strategy to support the mining of high-efficiency item sets of negative profits.

Benefits of technology

It realizes efficient computing of large-scale data sets, reduces memory overhead and computing costs, improves the flexibility and integrity of the algorithm, can effectively cover the set of efficient items in actual business scenarios, and supports distributed parallel computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067176B_ABST
    Figure CN120067176B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data mining and relates to an efficient utility itemset mining method related to massive data covering negative profits. The method includes: initializing multi-level partitioning of a database; generating and sorting candidate items; constructing a pre-evaluation matrix; constructing a utility list for candidate items that meet the utility constraints; performing iterative processing on partitions; performing depth-first search and itemset expansion; terminating recursion and outputting results. The present invention for the first time proposes an algorithm framework that supports the mining of relevant high-utility itemsets covering negative profits, and completely covers the high-utility itemsets in actual business scenarios. Through reconstructing the upper bound of utility and the correlation pruning mechanism, efficient calculation of large-scale data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining, and relates to an efficient high-utility itemset mining method related to massive data covering negative profits. Background Art

[0002] High-Utility Itemset Mining (HUIM) is an important branch in the field of data mining, aiming to discover commodity combinations with high profits or high values from transaction data, so as to provide more operational insights for decision-making. Different from traditional frequent itemset mining, HUIM not only considers the occurrence frequency of item sets, but also comprehensively evaluates utility values (such as profits) such as commodity unit prices and purchase quantities, which is more in line with the actual needs of business scenarios. However, existing HUIM methods mainly focus on utility maximization, which may lead to relatively weak correlations in the extracted item sets, or even just accidentally occurring combinations. Such a pattern often misleads decision-making in actual decision-making scenarios.

[0003] To solve this problem, Correlated High-Utility Itemset Mining (CHUIM) has been proposed, aiming to mine item sets with high utility and strong correlations between commodities. Although there are already a few methods for CHUIM, with the expansion of the database scale, the implementation based on tree structures and utility lists faces significant memory overhead problems. To address this challenge, researchers have proposed clustering-based techniques and compressed storage methods, but these methods rely on preprocessing steps to eliminate hopeless items through preset thresholds, which forms a computational bottleneck in scenarios where different thresholds need to be explored. In addition, existing methods are based on the maximum greedy strategy, defaulting to the assumption that all commodity utility values are positive, which will cause high-utility commodity combinations with negative profits to be missed. The single correlation evaluation index also limits the flexibility of the algorithm in different application scenarios.

[0004] In summary, the existing technologies usually have three technical bottlenecks: integrity, efficiency, and flexibility. Summary of the Invention

[0005] Aiming at the three major technical bottlenecks existing in the existing correlated high-utility itemset mining technology (CHUIM), namely the lack of combination integrity, low computational efficiency, and insufficient flexibility in adapting to scenarios, the present invention proposes a correlated high-utility itemset mining method covering negative profits.

[0006] The technical solution adopted by the present invention is: an efficient high-utility itemset mining method related to massive data covering negative profits, comprising the following steps:

[0007] Step 1: Use a multi-level partitioning method to divide the original database into multiple independent sub-partitions to form non-overlapping computing units;

[0008] Step 2: Screen out candidate 1-item sets whose positive transaction weighted utility meets the minimum utility threshold, and sort the items that meet the threshold constraints in the predefined priority order;

[0009] Step 3: Construct a two-dimensional pre-evaluation matrix with the sorted candidate items. This matrix stores the upper bounds of utility and support for 1-item sets and 2-item sets;

[0010] Step 4: Construct a utility list for the candidate set containing positive utility items and negative utility items screened out in Step 2;

[0011] Step 5: Read each partition from the external storage in turn, judge and prune the partition and the items contained in the partition through the first pruning strategy, and eliminate the partitions whose upper bound of utility is lower than the minimum utility threshold and the items in the partition; Define the items whose upper bound of utility meets the utility threshold as candidate items, construct a utility list for the candidate items in the partition, and record the support, utility, and remaining utility of the items;

[0012] Step 6: In each partition processed in Step 5, perform a depth-first search and item set expansion with the candidate set utility list constructed in Step 4 as the initialization;

[0013] Step 7: Terminate the depth-first search when there are no new candidate item sets to expand, and output all relevant high-utility item sets that meet the conditions.

[0014] Preferably, in Step 1, the multi-level partitioning method is specifically as follows:

[0015] (1) Sort the items in the transaction according to the predefined priority order of the items;

[0016] (2) In the first-level partition, split each transaction in the database into a positive item sub-transaction set and a negative item sub-transaction set according to the positive and negative utilities of the items; The second-level partition recursively divides the transaction into sub-transaction sets containing the same prefix items; The third-level partition constructs a dual-index mechanism: one index is arranged in ascending order of item support within the partition, and the other index is arranged in ascending order of the upper bound of utility.

[0017] Preferably, the predefined priority order of the items specifically includes: Let the item universe , represents the set of positive utility items, represents the set of negative utility items; For any item , when using all-confidence or bond correlation index, the priority order of the items is defined as: The positive utility items are sorted in ascending order of positive transaction weighted utility, the negative utility items are sorted in ascending order of support, and the positive utility items always take precedence over the negative utility items.

[0018] Preferably, when using kulc as the correlation index, the priority order of items is defined as follows: regardless of the positive or negative utility of the item, they are sorted in ascending order of support, but the items with positive utility still maintain their priority.

[0019] Preferably, in the third step, the pre-evaluation matrix stores the support and upper bound of utility of the item set in a two-dimensional form, expressed as:

[0020] ;

[0021] where, , represent candidate items; represents the priority order; represents the upper bound of utility; represents the support.

[0022] Preferably, for the kulc index, the utility list constructed in the fourth step contains the support / item number field; for the all-confidence index, the utility list contains the maximum support field; for the bond index, the differential support set of the item set is stored using a bit set structure.

[0023] Preferably, in the fifth step, the first pruning strategy is a dual pruning mechanism based on the upper bound of utility, including: prefix partition-level pruning. For any item , if its upper bound of utility is lower than the utility threshold , then directly prune the prefix subspace corresponding to this item, that is, terminate the further expansion of all item sets with as the prefix; sub-item-level pruning. In the unpruned prefix subspace , perform a secondary screening on each item : if , then permanently remove the item and all its expanded item sets from the prefix subspace .

[0024] Preferably, in the sixth step, it specifically includes: defining PL as the set of utility lists of item sets that need to be judged and processed in the current iteration round;

[0025] Traverse all item sets in the PL set in sequence. For one item set , if its utility , then mark the item set as a highly correlated and highly utility item set;

[0026] According to the second pruning strategy, for the item set , if its utility The sum with the remaining utility is lower than the utility threshold , that is , then prune this itemset and all its extended item sets;

[0027] Combine two item sets in the current extended set and to generate a higher-order item set ; According to the third pruning strategy, query the upper bound of the utility of the item set in the pre-evaluation matrix. If it is lower than the utility threshold , then directly prune it;

[0028] For the item set generated by expanding according to the predefined item priority order, according to the fourth pruning strategy, if its correlation metric value , is the predefined correlation threshold, then prune this item set and all its extended item sets;

[0029] For the higher-order item set generated by the item set and the item set , according to the fifth pruning strategy, if its upper bound of correlation , is the predefined correlation threshold, then prune this item set and all its extended item sets; otherwise, retain this item set in the new extended candidate set for the next round of expansion.

[0030] Preferably, the upper bound of the correlation includes bond coefficient, all-confidence coefficient; among them, all- confidence The upper bound of the correlation of the coefficient is:

[0031] ;

[0032] Among them, represents the support degree;

[0033] bond The upper bound of the correlation of the coefficient satisfies the relationship:

[0034] .

[0035] The beneficial effects of the present invention are reflected in the following aspects:

[0036] (1) Complete, efficient, and flexible large-scale data processing framework: The present invention first proposes an algorithm framework for mining relevant high-utility item sets that supports covering negative profits, completely covering high-utility item sets in actual business scenarios. By reconstructing the upper bound of utility and the correlation pruning mechanism, efficient computation of large-scale data sets is achieved;

[0037] (2) Multi-level partitioning calculation and resource optimization: A multi-level database partitioning method designed based on commodity attributes and transaction characteristics divides the original data set into independent computing units and supports distributed parallel computing. Item set calculation only needs to scan local data, greatly reducing the redundant overhead brought by global traversal;

[0038] (3) Efficient pruning and pre-evaluation acceleration mechanism: Integrate five innovative pruning strategies and combine with a pre-evaluation matrix to effectively prune hopeless item sets;

[0039] (4) Multi-dimensional verification and flexibility: In comparative tests of multiple real data sets and synthetic data sets, this framework demonstrates three major advantages: integrity guarantee, efficiency advantage, and flexibility expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the research content in the embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of the technical architecture of the method for mining high-utility item sets related to massive data covering negative profits proposed in the embodiment of the present invention;

[0042] Figure 3 It is a schematic diagram of the multi-level partitioning method proposed by the present invention;

[0043] Figure 4 It is a schematic diagram of the data structure of the pre-evaluation matrix PEM proposed by the present invention;

[0044] Figure 5 It is a schematic diagram of the variant of the UList structure proposed by the present invention;

[0045] Figure 6 It is for all-confidence coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage;

[0046] Figure 7 It is for bond coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage;

[0047] Figure 8 It is for kulc coefficient-based algorithm performance evaluation; where, (a) running time; (b) memory usage;

[0048] Figure 9 Evaluation of the CHIN and FHN algorithms in handling negative utility; among them, (a) running time; (b) memory usage; (c) number of candidate item sets. Detailed implementation mode

[0049] For the convenience of understanding the present invention, the present invention will be described in more detail below with reference to the accompanying drawings and specific embodiments. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive.

[0050] In a retail scenario, such as Figure 1 As shown, the high-utility item set mining (HUIM) algorithm can mine a large number of high-profit product combinations similar to {steak, coffee, bread} (total utility = $28, correlation bond coefficient = 0.17). However, research has found that such item sets actually often belong to accidental combinations due to insufficient correlation between products. To solve this problem, the related high-utility item set mining (CHUIM) algorithm introduces a correlation threshold constraint of β≥0.4 as a screening mechanism, but this method has obvious defects: when the item set contains "loss baits" (such as bread), the traditional CHUIM method will erroneously eliminate these key combinations based on the "utility is always positive" assumption. For example, the {steak, bread} combination creates a total profit utility of $38 and satisfies bond the association strength of the coefficient 0.4, but because bread is a negative-utility product (unit profit = $-1), it is misjudged and deleted by the algorithm, resulting in the inability to identify such "losing to make a profit" business strategies. To solve such problems, the present invention proposes a method for mining related high-utility item sets covering negative profits in retail data.

[0051] Example 1 A method for mining related high-utility item sets covering negative profits provided by the present invention has a technical architecture as Figure 2 shown, and specifically includes the following steps:

[0052] Step 1: Initialization of multi-level partitioning of the database.

[0053] Using the multi-level partitioning method proposed by the present invention, the original database is divided into multiple independent sub-partitions to form non-overlapping computing units. To solve the efficiency bottleneck of the existing method, the present invention proposes an innovative multi-level partitioning method, as Figure 3 shown, and specifically includes the following processes:

[0054] 1. Pre-definition of the priority order of items:

[0055] In the search space containing positive and negative utility items, the present invention first defines the priority order of items . Let the item universe , where represents the set of positive utility items, where represents the set of negative utility items. According to different correlation measurement indicators adopted, the priority order is specifically defined as follows:

[0056] (1) Priority rule based on all-confidence / bond coefficient

[0057] For any item , when all-confidence or bond correlation index is adopted, the priority order is defined as: positive utility items are sorted in ascending order of positive transaction weighted utility, negative utility items are sorted in ascending order of support, and positive utility items always take precedence over negative utility items, which is expressed as follows:

[0058] (1);

[0059] Where: represents the positive transaction weighted utility value (Positive Transaction WeightUtility) of item ; represents the support (frequency of occurrence) of item .

[0060] (2) Priority rule based on kulc index

[0061] When kulc coefficient is adopted as the correlation index, the priority order is defined as: regardless of the positive or negative utility of the item, they are sorted in ascending order of support, but positive utility items still maintain their priority, which is expressed as follows:

[0062] (2).

[0063] 2. Multi-level partitioning strategy

[0064] In the preprocessing stage, first, the items in the transaction are sorted according to the item priority defined by formulas (1) and (2). In the first-level partitioning, each transaction in the database is split into a positive item sub-transaction set and a negative item sub-transaction set based on the positive and negative utilities of the items, realizing the independent processing of positive and negative utility items. The second-level partitioning forms a fine-grained partitioning structure by recursively splitting the transaction into sub-transactions containing the same prefix items. The third-level partitioning innovatively constructs a dual-index mechanism: the first index is based on the item support within the partition Sorted in ascending order, the second index According to the upper bound of utility Sorted in ascending order. This dual-index structure enables directly skipping the entire partition corresponding to the items that do not meet the threshold during partition loading, significantly reducing memory occupancy and computational overhead. This hierarchical strategy lays the foundation for the implementation of subsequent pruning strategies by avoiding the combined operations of invalid items.

[0065] Step 2: Candidate set generation and sorting.

[0066] Filter out the candidate items whose positive transaction weighted utility meets the minimum utility threshold, and sort these threshold-meeting items in the defined priority order ( ).

[0067] Step 3: Construct the pre-evaluation matrix (PEM).

[0068] Construct a two-dimensional pre-evaluation matrix PEM with the sorted candidate items. This matrix stores the upper bounds of utility and support degrees of 1-item sets and 2-item sets.

[0069] To improve the computational efficiency, the present invention proposes a pre-evaluation matrix (PEM), whose data structure is as Figure 4 shown. This matrix stores the support degrees of item sets and upper bounds of utility in a two-dimensional form, and the mathematical definition is:

[0070] (3);

[0071] Among them, , represent candidate items.

[0072] The implementation process of its function:

[0073] (1) Data preloading: At the initial stage of mining, scan the data set and calculate the support degrees and upper bounds of utility of all 1-item sets and 2-item sets, and fill them into the PEM matrix according to the corresponding subscripts.

[0074] (2) Quick upper bound determination: When generating a new item set (such as generating {a, b, c} from {a, b} and {a, c}), by querying the support degree and upper bound of utility of the {b, c} subset in the PEM, obtain the support degree and upper bound of utility corresponding to this item set.

[0075] (3) Early pruning decision: If the upper bound of utility is lower than the preset threshold, directly discard this candidate item set and terminate the expansion of the current item set to avoid subsequent invalid calculations.

[0076] Step 4: Construct a utility list (UList) for the candidate item sets (including positive-utility items and negative-utility items) screened out in Step 2.

[0077] To adapt to the calculation of different correlation metrics, the present invention designs three variants of the UList data structure, as Figure 5 shown, to achieve efficient calculation through field embedding.

[0078] For kulc the metric, a support / number of items (sup / item) field is added to the UList to directly record the support values of each item in the item set. This design avoids the operation of repeatedly scanning the entire database in the traditional method. For example, when analyzing the combination of {milk, bread}, the independent support values of both can be directly obtained from this field, significantly reducing the calculation overhead. kulc For

[0079] the metric, a maximum support ( all-confidence ) field is introduced to store the maximum support value of all items in the item set. Taking the combination of {steak, coffee} as an example, this field will automatically record the higher support value between steak and coffee, enabling the maxsup coefficient to be directly obtained through a single division operation, eliminating the redundant calculation of repeatedly traversing the transaction data in the traditional method. all-confidence For

[0080] the metric, a bitset structure is used to store the differential support set ( bond ) of the item set. When expanding the item set (such as generating {a, b, c} from {a, b} and {a, c}), the dissup intersection of the two parent item sets can be quickly calculated through bit operations, reducing the time complexity of calculating the dissup coefficient from O(n) to O(1). bond Through the above structural innovation, the correlation calculation process is simplified to a combination of field queries and basic operations, and it supports the rapid integration of new correlation metrics, effectively improving the calculation efficiency and scalability of the algorithm.

[0081] Step 5: Process iteratively by partition.

[0082] Read each partition from the external memory in sequence, and perform judgment and pruning on the partition and the items contained in the partition through the first pruning strategy, eliminating the partition and the items in the partition whose upper utility bounds

[0083] are lower than the utility threshold. For the items whose upper utility bounds meet the utility threshold (becoming candidate items), construct a UList data structure for them and record parameters such as the support, utility, and remaining utility of the items. For the items whose upper utility bounds

[0084] The first pruning strategy is to adopt a dual pruning mechanism based on ptwu . Specifically, it includes:

[0085] (1) Prefix partition-level pruning: For any item , if its upper bound of utility is lower than the preset utility threshold , then directly prune the prefix subspace corresponding to this item, that is, terminate the further expansion of all item sets prefixed with (such as , , , etc.).

[0086] (2) Sub-item-level pruning: In the unpruned prefix subspace , perform a secondary screening on each item . If , then permanently remove the item and all its extended item sets (such as ) from the prefix subspace .

[0087] Step Six: Depth-First Search (DFS) and item set expansion.

[0088] First, define PL as the set of utility lists of item sets that need to be judged and processed in the current iteration round. PL is initially set to the candidate item set utility list (UList) constructed in Step 4.

[0089] 1. CHUI candidate verification: Traverse all item sets in PL in sequence. For one item set , if its utility is greater than or equal to the preset utility threshold , then mark the item set as a highly relevant and highly useful item set (CHUI),

[0090] 2. Expansion feasibility determination: According to the second pruning strategy, for the item set , if the sum of its utility and the remaining utility is lower than the utility threshold , that is, , then prune this item set and all its extended item sets (all higher-order item sets prefixed with ).

[0091] 3. Item set combination generation: Combine the two item sets and (both prefixed with the item set ) in the current expansion set to generate a higher-order item set .

[0092] 4. Pruning based on PEM: According to the third pruning strategy, query the upper bound of the utility of the item set in the PEM matrix . If it is lower than the utility threshold , then directly prune it.

[0093] 5. Correlation filtering: It includes upper bound pruning of correlation and upper bound pruning of joint correlation. Only retain the item sets that meet the correlation threshold to the newly expanded candidate set for the next round of expansion.

[0094] (1) The upper bound pruning of correlation adopts the fourth pruning strategy, specifically: for the item sets generated by expanding in the predefined item priority order , if its correlation metric value ( is the predefined correlation threshold), then prune this item set and all its expanded item sets.

[0095] This pruning strategy calculates in real time during the item set expansion process , which includes bond coefficient, all- confidence coefficient, kulc coefficient and other indicators. If , it indicates that the correlation of the item set and its expanded item sets does not meet the threshold constraint, then terminate the further expansion of to avoid generating all item sets with as the prefix.

[0096] (2) The upper bound pruning of joint correlation adopts the fifth pruning strategy, specifically: for the high-order item set generated by the item set and the item set , define the upper bound of the correlation of its all-confidence coefficient as:

[0097] (4);

[0098] Among them, is the support of the item set. This upper bound predicts the upper bound of the coefficient of the item set all-confidence by comparing the extreme values of the support of the sub-item set and the joint item set.

[0099] For the upper bound of the correlation of the bond coefficient, it inherits the boundary property of formula (4) and satisfies the relationship:

[0100] ​ (5).

[0101] Therefore, the upper bound of the correlation of the coefficients in formula (4) can be directly used. As bond the upper bound of the correlation of the coefficients. Based on this property, if , the algorithm terminates the expansion of the extended item set of the item set its extended item set.

[0102] Step Seven: Recursion Termination and Result Output.

[0103] When there are no new candidate item sets to expand, terminate the DFS and output all CHUI sets that meet the conditions.

[0104] An overview of each pruning strategy is shown in Table 1.

[0105] Table 1 Summary and Overview of Pruning Strategies

[0106] .

[0107] Example Two: Performance Evaluation

[0108] Since the method for dealing with datasets containing negative utilities is proposed for the first time in the present invention, there is no existing algorithm for direct comparison. To evaluate its performance, the present invention designs two comparative experiments:

[0109] (1) On the dataset containing only positive utilities, compare with the existing high-utility item set mining algorithms;

[0110] (2) Set the correlation threshold to 0 and compare with the existing high-utility item set mining algorithms (for dealing with negative utility data).

[0111] Experiment 1: Evaluation of the Method of the Present Invention (CHIN Algorithm) and the Existing Method (CHUIM Algorithm)

[0112] The present invention evaluates the performance of the CHIN algorithm and compares it with five latest related high-utility item set mining algorithms. The comparative experiment is based on three widely used correlation metrics: kulc Coefficient, all-confidenc e coefficient, and bon d coefficient. To ensure a fair comparison, the experiment is only carried out on the dataset containing only positive utilities because the existing comparative algorithms do not support CHUIM with negative utilities.

[0113] Existing methods (CHUIM algorithm) include: FCHM: P. Fournier-Viger, Y. Zhang, J. C.-W. Lin, D.-T. Dinh, and H. Bac Le, “Mining correlated high-utility itemsets using various measures,” Logic Journal of the IGPL, vol. 28, no. 1, pp. 19–32, 2020;

[0114] ECHUM: D. Dharavath Ramesh, K. K. Sethi, and A. Rathore, “Positive correlation based efficient high utility pattern mining approach,” in Proceedings of Sixteenth International Conference on Information Processing, Data Science and Computational Intelligence, 2022, p. 273;

[0115] GMCHM: N. M. Hung, T. NT, and B. Vo, “A general method for mining high-utility itemsets with correlated measures,” Journal of Information and Telecommunication, vol. 5, no. 4, pp. 536–549, 2021.

[0116] Table 2 shows the performance of CHIN in terms of running time (RT), candidate item volume (CN), and peak memory usage (MU). The experiments were conducted on four datasets (Retail, Mushroom, Chainstore, and Chicage-Cremes, and the datasets were sourced from the SPMF open-source data mining platform http: / / www.philippe-fournier-viger.com / spmf / ), with parameters , Set to 0.1% of the total utility of the corresponding dataset. The experimental results show that CHIN significantly outperforms the existing CHUIM algorithm under three correlation metrics. For example, on the Retail dataset, CHIN has a running time of 1.66 seconds under the kulc metric, while ECHUM is 16.11 seconds, and at the same time, the number of candidates and memory consumption are significantly reduced. This benefits from its advanced multi-level partitioning and effective pruning strategy, which reduces the candidate generation and computational cost, thereby reducing the memory usage and I / O overhead.

[0117] Table 2: Evaluation of CHIN and CHUIM algorithms on real datasets

[0118] .

[0119] As Figure 6 and Figure 7 shown, under the all-confidence and bond coefficient metrics, the running time and memory usage of all algorithms on the Retail dataset decrease significantly as the minimum utility threshold increases. For example, in Figure 7 (a), the running time of FCHM and CHIN shortens from dozens of seconds to several seconds, while in Figure 7 (b) the memory usage drops from about one thousand MB to several hundred MB. At lower values (such as ), the memory overhead of GMCHM is relatively high, while CHIN and FCHM maintain higher efficiency. As increases, the gap between the algorithms narrows, but CHIN still remains competitive in terms of running time and memory usage.

[0120] For the kulc coefficient, Figure 8 shows a similar trend: the running time and memory usage of CHIN and ECHUM both decrease significantly as increases. However, CHIN is generally better than ECHUM, especially at lower values, indicating its higher efficiency in processing large-scale databases. When is set to 0.1, 0.5, and 0.9, the experimental results always show similar trends. Increasing will result in a reduction in the number of discovered item sets, thus reducing the computational cost, and CHIN shows advantages in most scenarios.

[0121] In summary, these experiments verify the robust performance advantages of CHIN proposed in the present invention under different correlation metrics, making it a flexible method for mining correlated high-utility item sets.

[0122] Experiment 2: Evaluation of CHIN and FHN

[0123] The present invention studies the performance of CHIN and FHN (J. C.-W. Lin et al., “FHN: an efficient algorithm for mining high-utility itemsets with negative unit profits,” Knowledge-Based Systems, vol. 111, pp. 283–298, 2016.) through a correlation threshold β = 0. Under this condition, the correlation constraint is removed, allowing both algorithms to focus on high-utility itemset mining.

[0124] The present invention Figure 9 compares CHIN with the widely used FHN on the Mushroom_negative dataset. The results show that CHIN outperforms FHN in terms of running time, memory usage, and the number of candidate itemsets. Specifically, due to its advanced partitioning method and pruning strategy, CHIN significantly reduces the running time and the number of generated candidate itemsets, thus reducing the computational cost and memory consumption.

[0125] In addition, by processing prefix-based partitioning, CHIN achieves a lower I / O cost than FHN, which is particularly beneficial when dealing with large-scale datasets. These findings indicate that even when the correlation threshold is zero, CHIN still maintains a robust performance advantage over FHN, highlighting its efficiency in mining high-utility itemsets with negative utilities.

Claims

1. An efficient utility itemset mining method related to massive data covering negative profits, characterized in that, Including the following steps: Step 1: Use a multi-level partitioning method to divide the original database into multiple independent sub-partitions, forming non-overlapping computing units; Step 2: Screen out candidate 1-item sets whose positive transaction weighted utility meets the minimum utility threshold, and sort these items that meet the threshold constraints in the predefined priority order; Step 3: Construct a two-dimensional pre-evaluation matrix with the sorted candidate items. This matrix stores the upper bounds of the utility and support of 1-item sets and 2-item sets; Step 4: Construct a utility list for the candidate set containing positive utility items and negative utility items screened out in Step 2; Step 5: Read each partition from the external storage in turn, judge and prune the partition and the items contained in the partition through the first pruning strategy, and eliminate the partitions and the items in the partition whose upper bound of utility is lower than the minimum utility threshold; Define the items whose upper bound of utility meets the utility threshold as candidate items, construct a utility list for the candidate items in the partition, and record the support, utility, and remaining utility of the items; Step 6: In each partition processed in Step 5, perform a depth-first search and item set expansion with the utility list of the candidate set constructed in Step 4 as the initialization; Step 7: Terminate the depth-first search when there are no new candidate item sets to expand, and output all relevant high-utility item sets that meet the conditions.

2. The efficient utility itemset mining method related to massive data covering negative profit according to claim 1, characterized in that In Step 1, the multi-level partitioning method is specifically as follows: (1) Sort the items in the transaction according to the predefined priority order of the items; (2) In the first-level partition, split each transaction in the database into a positive item sub-transaction set and a negative item sub-transaction set according to the positive and negative utilities of the items; The second-level partition recursively divides the transaction into sub-transaction sets containing the same prefix items; The third-level partition constructs a dual-index mechanism: one index is sorted in ascending order of item support within the partition, and the other index is sorted in ascending order of the upper bound of utility.

3. The efficient utility itemset mining method related to massive data covering negative profit according to claim 2, characterized in that, The priority order of the predefined items specifically includes: setting the item universe , represents the set of positive utility items, represents the set of negative utility items; for any item , when using all-confidence or bond the correlation index, the priority order of the items is defined as: the positive utility items are sorted in ascending order of positive transaction weighted utility, the negative utility items are sorted in ascending order of support, and the positive utility items always take precedence over the negative utility items.

4. The efficient utility itemset mining method related to massive data covering negative profits according to claim 3, characterized in that, When using kulc the coefficient as the correlation index, the priority order of items is defined as follows: regardless of the positive or negative item utility, they are sorted in ascending order of support, but the positive-utility items still maintain the priority status.

5. The method for mining high-utility item sets related to massive data covering negative profits according to claim 1, wherein In Step 3, the pre-evaluation matrix stores the support and the upper bound of utility of the item set in a two-dimensional form, expressed as: ; Among them, , represent candidate items; represents the priority order; represents the upper bound of utility; represents the support degree.

6. The efficient utility itemset mining method related to massive data covering negative profit according to claim 1, characterized in that, The utility list constructed in step four above, for kulc the metric, the support / item number field is included in the utility list; for all-confidence the metric, the maximum support field is included in the utility list; for bond the metric, a bit set structure is adopted to store the differential support set of the item set.

7. The method for mining high-utility item sets related to massive data covering negative profits according to claim 1, characterized in that, In the fifth step, the first pruning strategy is a dual pruning mechanism based on the upper bound of utility, including: prefix partition-level pruning. For any item , if its upper bound of utility is lower than the utility threshold , then directly prune the prefix subspace corresponding to this item , that is, terminate the further expansion of all item sets with as the prefix; sub-item-level pruning. In the unpruned prefix subspace , perform a secondary screening for each item : If , then permanently remove the item and all its extended item sets from the prefix subspace .

8. The method for mining high-utility item sets related to massive data covering negative profits according to claim 1, characterized in that, In Step 6, it specifically includes: Define PL as the set of utility lists of the item sets that need to be judged and processed in the current iteration round; Traverse all item sets in the PL set in sequence, for one of the item sets , if its utility , then mark the item set as a highly relevant and highly utility item set; According to the second pruning strategy, for the item set , if its utility and the remaining utility sum is lower than the utility threshold , that is , then prune this item set and all its extended item sets; Combine two item sets in the current extended set and to generate a higher-order item set ; According to the third pruning strategy, query the upper bound of the utility of the item set in the pre-evaluation matrix. If it is lower than the utility threshold then directly prune; For the item set generated by expanding according to the predefined item priority order , according to the fourth pruning strategy, if its relevance metric value , is the preset relevance threshold, then prune this item set and all its extended item sets; For the high-order item set generated by item set and item set , according to the fifth pruning strategy, if its upper bound of relevance , is the preset relevance threshold, then prune this item set and all its extended item sets; otherwise, retain this item set in the new extended candidate set for the next round of extension.

9. The high-utility itemset mining method related to massive data covering negative profits according to claim 8, characterized in that The upper bound of the correlation includes bond coefficient,[[]] all-confidence coefficient; among which, all-confidence the upper bound of the correlation of the coefficient is: ; Among them, represents the support degree; bond The upper bound of the correlation of coefficients satisfies the relation: 。

Citation Information

Patent Citations

  • High-utility item set mining method containing negative utility

    CN110471960A

  • Periodic utility item set mining method, system and device under incremental database

    CN118377815A