Periodic clustering frequent pattern mining method for time-ordered periodic transaction data

CN118861115BActive Publication Date: 2026-09-29INNER MONGOLIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410619094.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-18
Publication Date
2026-09-29
Estimated Expiration
2044-05-18

AI Technical Summary

Technical Problem

[0006]本发明提出了面向时间有序周期事务数据的周期聚簇频繁模式挖掘方法,解决在时间有序周期事务数据集中挖掘周期聚簇频繁模式问题

Benefits of technology

[0013]本发明的有益效果:本发明Naive算法通过使用IPCFPM-list结构,可高效得到项集对应周期时间嵌套集合。同时,对于不可能成为聚簇频繁模式的周期,可删除对应周期索引和位向量,减少了对这些周期做交集和重复聚簇频繁模式判断,进而减少了时空开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861115B_ABST
    Figure CN118861115B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing and data mining, and particularly relates to a period clustering frequent pattern mining method for time-ordered period transaction data, comprising the following steps: firstly, using an Apriori-TID algorithm to judge whether each period of an item set is frequent; finding a time set corresponding to each frequent period, and using a DBSCAN clustering algorithm on the time set of each frequent period and judging whether clustering is satisfied, if clustering is satisfied, a corresponding clustering occurrence interval is obtained; performing similarity calculation between the clustering occurrence intervals of all periods, and judging whether the item set is a period clustering frequent pattern; performing the above judgment process on all item sets, and all period clustering frequent patterns can be found, and the method is a Naive algorithm. The present application proposes a period clustering frequent pattern mining algorithm for time-ordered period transaction data, and solves the problem of mining period clustering frequent patterns in a time-ordered period transaction data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data processing and data mining technology, and in particular to a method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data. Background Technology

[0002] Frequent itemset mining is a hot research topic in data mining, aiming to discover frequently occurring itemsets in transactional datasets. Originally designed for basket analysis, it has been widely applied to many other data mining tasks, including association rule mining, sequential pattern mining, clustering, and classification. Current research on frequent itemset mining can be broadly categorized into two types based on whether the data in the transactional dataset is temporally ordered: those for unlabeled transactional data and those for time-ordered transactional data. Frequent itemset mining for unlabeled transactional data includes research on basic frequent itemsets and derived types such as maximal frequent itemsets, frequent closed itemsets, and high-performance itemsets. In the field of frequent itemset mining for time-ordered transactional datasets, related research mainly focuses on the distribution of frequent itemsets within the dataset. For example, periodic patterns are used to mine itemsets that occur regularly in a transaction dataset; local periodic patterns are used to mine itemsets that occur regularly within certain time intervals; partial periodic patterns are used to mine itemsets that occur regularly in a transaction dataset but are not necessarily consecutive; rare periodic patterns are used to mine itemsets that occur infrequently but regularly in a transaction dataset; canonical frequent patterns are used to mine itemsets that appear in a transaction dataset at fixed intervals; and periodic clustering patterns are used to find events in an event sequence where the occurrence positions are clustered and the interval between any two adjacent clusters satisfies the periodic characteristic. However, overall, this field has not yet conducted research on the problem of frequent itemset mining where the same clustering state occurs in multiple periods, which makes it impossible to directly use existing methods to effectively handle many related applications.

[0003] Existing technologies propose local periodic pattern mining algorithms that can discover itemsets exhibiting periodic behavior within undefined time intervals. Two novel metrics are used to evaluate the periodicity and frequency of patterns within time intervals: maximum overflow period (a metric that allows detection of time intervals with variable lengths) and minimum duration (a metric that ensures these occurrence time intervals have a minimum duration). Existing technologies also propose periodic clustering pattern mining algorithms, which can mine events in event sequences where the occurrence locations are clustered and the interval between any two adjacent clusters satisfies periodic characteristics.

[0004] Local periodic patterns describe patterns that occur periodically within a certain time period and are constrained by a minimum duration condition. Frequent periodic clustering patterns, however, seek patterns that occur frequently within a period or at a specific point in time, exhibiting a clustered occurrence. The pattern does not require a duration constraint, and frequent periodic clustering patterns require similar clustering intervals across multiple periods. Periodic clustering patterns describe only one itemset and do not consider the occurrence of multiple itemsets. Furthermore, items whose clustering locations within a period do not exhibit periodicity are discarded. In comparison, frequent periodic clustering patterns focus on all itemsets that satisfy the frequent periodic clustering condition, and within each period satisfying the frequent clustering pattern, there may be multiple clusters; frequent periodic clustering patterns do not require the clustering locations to have periodicity.

[0005] Therefore, existing technologies cannot solve the problem of mining frequent patterns of periodic clusters in time-ordered periodic transaction datasets. Summary of the Invention

[0006] This invention proposes a method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which solves the problem of mining frequent patterns of periodic clusters in time-ordered periodic transaction datasets.

[0007] The technical solution adopted in this invention is: a method for mining frequent patterns of periodic clustering in time-ordered periodic transaction data, comprising the following steps: Step 1: First, use the Apriori-TID algorithm to determine whether each period of the itemset is frequent; Step 2: Find the set of occurrence times corresponding to each frequent cycle, and use the DBSCAN clustering algorithm on the set of occurrence times in each frequent cycle to determine whether they are clustered. If they are clustered, the corresponding cluster occurrence interval is obtained. Step 3: Perform similarity calculations between cluster occurrence intervals in all periods. If a cluster occurrence interval in a certain period is similar to the cluster occurrence intervals in other periods, and the total number of periods containing similar cluster occurrence intervals is greater than a given threshold, then the cluster occurrence interval is the periodic occurrence interval of the itemset, and the itemset is the periodic clustering frequent pattern. Step 4: By performing the above judgment process on all itemsets, all frequent patterns of periodic clustering can be found. This method is called the Naive algorithm.

[0008] As a further improvement to this invention, the existence rules of unnecessary operations in the Naive algorithm are summarized in the form of Lemma 1, Lemma 2, Theorem 1, and Theorem 2. Based on these rules, redundant operations in the Naive algorithm are effectively reduced. Lemma 1: For itemsets x 'and x correspondPTS x’ and PTS x Medium-term index is u members PTS x’ u and PTS x u Considering PTS x’ u and PTS x u After division C x’ u and C x u ,if x’ ⊂ x Then we have | C x u |≤| C x’ u |Established; Lemma 2: Let minCF = minSup ×| PTDS |×(1- maxNR ), | PTDS |for PTDS Includes the total number of transactions, given itemsets x about PTS x Medium-term index is u members PTS x u ,right PTS x u After division C x u We have the following conclusion: if | C x u |< minCF ,So x Not a frequent clustering pattern; Theorem 1: For itemsets x 'about PTS x’ Medium-term index is u members PTS x’ u If | C x'u |< minCF Then any superset x ⊃ x 'exist u The clustering pattern is not frequent in any of the cycles; Theorem 2: For itemsets x ',make cfPidset x’ Indicates satisfaction minCF The set of periodic indexes, | cfPidset x’ | indicates the total number of periodic indexes included. cfPidset x’ | / | pidSet |< minPR Then any superset x ⊃ x Neither of them exhibits a frequent clustering pattern.

[0009] As a further improvement to this invention, the data structure uses... ts-list The structure uses bit vectors to represent the set of times that occur over all periods. ts-list It is a structure of ( x , bitvec x hash table of ) bitvec x express x For the corresponding bit vector, let x corresponding ts-list Recorded as ts-list x ,use ts-list ( x )express bitvec x And use an array to store the mapping between bit vector positions and timestamps, denoted as bvTimeMap The specific implementation of the algorithm can be solved in the following three stages: Phase 1: Mapping Itemsets ts-list pass bvTimeMap The mapping is done to the set of occurrence times across all periods, and then the time set for each period is obtained using PTDS; next, the occurrence time set across all periods for each itemset is frequently processed. minSup Filter to obtain all frequent cycles of the itemset; Phase Two: For itemsets that satisfy frequent periodicity, a threshold based on the neighborhood radius is established. Eps and density threshold minPts and maximum noise ratio threshold maxNR Filtering ensures that the period satisfies the clustering characteristics, resulting in a dataset that includes all intervals where frequent clustering patterns occur; Phase 3: Finally, based on the minimum interval similarity threshold... minSMand minimum periodicity threshold minPR Filter the dataset of frequent clustering patterns for each itemset to ensure that the resulting itemsets have periodic occurrence intervals, meaning that the occurrence intervals occur at least in other... minPR Similar occurrence intervals exist across -1 cycles. By performing the above operation on each itemset, all periodic clustering frequent patterns can be obtained.

[0010] As a further improvement to the present invention, the second stage uses the DBSCAN algorithm. Eps and minPts All are input parameters for the DBSCAN algorithm.

[0011] As a further improvement of the present invention, the two occurrence intervals of the third stage. v and w Similarity is if and only if any of the following conditions are met: v.begin ≤ w.begin and v.end ≥ w.end , recorded as v ∩ w ≤ minSM The intervals intersect; | v.begin - w.begin |≤ minSM And| v.end - w.end |≤ minSM , recorded as w ⊆ v The interval contains.

[0012] As a further improvement to the present invention, in ts-list Based on the structure, it was proposed IPCFPM-list Data structures: Consider l Itemset x , x corresponding IPCFPM-list Represented as IPCFPM-list x , IPCFPM-list x It is a two-level hash table structure, the outer hash structure is ( x , bitvec x ),in, x It is a hash table key ; bitvec x It is a hash table value ,and bitvec x It is itself a hash structure, represented as ( pid , bitvec xpid ),in, pid ∈ pidSet x , pid for key ,correspond bitvec x pid for value , bitvec x pid Indicates a periodic index pid The period corresponds to the bit vector representation of the time set.

[0013] The beneficial effects of this invention: The Naive algorithm of this invention uses... IPCFPM-list The structure can efficiently obtain nested sets of periodic times corresponding to itemsets. Furthermore, for periods unlikely to be frequent clustering patterns, the corresponding period index and bit vector can be deleted, reducing the need to perform intersection and repeated frequent clustering pattern checks on these periods, thereby reducing time and space overhead. Attached Figure Description

[0014] Figure 1 The IPCFPM-list corresponding to {a}, {f} and {a,f} in the method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data of this invention; Figure 2 This invention relates to bvTimeMapSet, a method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data. Figure 3 This is a comparison chart of the running time of the algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to the present invention when minSup changes; Figure 4 This is a comparison chart of the running time of the algorithm for the periodic clustering frequent pattern mining method for time-ordered periodic transaction data of this invention when Eps changes; Figure 5 This is a comparison chart of the running time of the algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to the present invention when minPts changes; Figure 6 This is a comparison chart of the running time of the algorithm for the periodic clustering frequent pattern mining method for time-ordered periodic transaction data of this invention when maxNR changes; Figure 7 This is a comparison chart of the running time of the algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to the present invention when minSM changes; Figure 8 This is a comparison chart of the running time of the algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to the present invention when minPR changes; Figure 9 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares the space overhead and discovers the number of patterns when minSup changes. Figure 10 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares spatial overhead and discovers the number of patterns when Eps changes. Figure 11 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares the space overhead and discovers the number of patterns when minPts changes. Figure 12 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares spatial overhead and discovers the number of patterns when maxNR changes. Figure 13 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares spatial overhead and discovers the number of patterns when minSM changes. Figure 14 This invention relates to an algorithm for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, which compares spatial overhead and discovers the number of patterns when minPR changes. Detailed Implementation

[0015] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments described herein are merely illustrative and are not intended to limit the scope of this application.

[0016] This invention provides a method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data, comprising the following steps: Step 1: First, use the Apriori-TID algorithm to determine whether each period of the itemset is frequent; Step 2: Find the set of occurrence times corresponding to each frequent cycle, and use the DBSCAN clustering algorithm on the set of occurrence times in each frequent cycle to determine whether they are clustered. If they are clustered, the corresponding cluster occurrence interval is obtained. Step 3: Perform similarity calculations between cluster occurrence intervals in all periods. If a cluster occurrence interval in a certain period is similar to the cluster occurrence intervals in other periods, and the total number of periods containing similar cluster occurrence intervals is greater than a given threshold, then the cluster occurrence interval is the periodic occurrence interval of the itemset, and the itemset is the periodic clustering frequent pattern. Step 4: By performing the above judgment process on all itemsets, all frequent patterns of periodic clustering can be found. This method is called the Naive algorithm.

[0017] This invention summarizes the patterns of unnecessary computations in the Naive algorithm through Lemmas 1 and 2, and Theorems 1 and 2, effectively reducing redundant computations in the Naive algorithm based on these patterns: Lemma 1: For itemsets x 'and x correspond PTS x’ and PTS x Medium-term index is u members PTS x’ u and PTS x u Considering PTS x’ u and PTS x u After division C x’ u and C x u ,if x’ ⊂ x Then we have | C x u |≤| C x’ u |Established; Proof: Because x’ ⊂ x According to the Apriori property, any containing x’ The matters also necessarily include x , can be obtained PTS x u ∈ PTS x' u Therefore, ∀ t ∈ PTS x u They all t ∈ PTS x’ u According to the DBSCAN algorithm, cluster formation is based on... Eps and minPts ,and C xu and C x’ u The clusters in the middle are respectively composed of TS x and TS x' The timestamps in the cluster are used to construct the ∀ cluster. c 2∈ C x u ,∃cluster c 1∈ C x' u , making c 2∈ c 1. Therefore, we can obtain | c 2|≤| c 1|, and thus we can obtain | C x u |≤| C x' u |. Q.E.D. Lemma 2: Let minCF = minSup ×| PTDS |×(1- maxNR ), | PTDS |for PTDS Includes the total number of transactions, given itemsets x about PTS x Medium-term index is u members PTS x u ,right PTS x u After division C x u We have the following conclusion: if | C x u |< minCF ,So x Not a frequent clustering pattern; Proof: Because | C x u |< minCF ,according to minCF = minSup ×| PTDS |×(1- maxNR ), we can get | C x u |< minSup ×|PTDS |×(1- maxNR ), assuming x In the periodic index u The frequency of these cycles indicates the support level within that cycle. Sup x u ≥ minSup ,in Sup x u =| PTS x u | / | PTDS Therefore, we can obtain | C x u |< Sup x u ×| PTDS |×(1- maxNR Expanding the parentheses on the right side of the comparison operator in the above expression and recombine the left and right sides of the comparison operator, we get... Sup x u ×| PTDS |-| C x u |> maxNR × Sup x u ×| PTDS By replacing the left and right sides of the comparison operator, we get | Noi x u |> maxNR ×| PTS x u By simplifying both sides of the comparison operator, we can obtain... NR x u > maxNR , because when NR x u > maxNR hour, x In the periodic index u The periodicity does not show a frequent clustering pattern. Q.E.D. Theorem 1: For itemsets x 'about PTS x’ Medium-term index is u members PTS x’ u If | C x'u |< minCF Then any superset x ⊃ x 'exist u The clustering pattern is not frequent in any of the cycles; Proof: Because | C x' u |< minCF According to Lemma 2, we can obtain x 'Not a frequent clustering pattern.' Because... x ⊃ x According to Lemma 1, we can obtain | C x u |≤| C x' u Therefore, it can be deduced that | C x u |< minCF According to Lemma 2, we can obtain x exist u The clustering pattern is not frequent over a given period. Therefore, if x 'If it is not a frequent clustering pattern, then any superset x ⊃ x 'exist u The clustering pattern is not frequent in any of the cycles. Q.E.D. Theorem 2: For itemsets x ',make cfPidset x’ Indicates satisfaction minCF The set of periodic indexes, | cfPidset x’ | indicates the total number of periodic indexes included. cfPidset x’ | / | pidSet |< minPR Then any superset x ⊃ x Neither of them exhibits a frequent clustering pattern.

[0018] Proof: Because | cfPidset x’ | / | pidSet |< minPR , x '⊂ x According to Lemma 1, we know that | C x u |≤| C x’ u | can be deduced| cfPidset x |≤|cfPidset x’ Therefore, there is | cfPidset x | / | pidSet |< minPR .for x It can be known that ( ITsup x v +1) / | pidSet |≥ minPR This is a necessary condition for the frequent pattern of periodic clustering, because ( ITsup x v +1)≤| cfpidSet x |, then ( ITsup x v +1) / | pidSet |< minPR Any superset x ⊃ x None of them exhibit a frequent, periodic clustering pattern. Q.E.D. The data structure used in this invention ts-list The structure uses bit vectors to represent the set of times that occur over all periods. ts-list It is a structure of ( x , bitvec x hash table of ) bitvec x express x For the corresponding bit vector, let x corresponding ts-list Recorded as ts-list x ,use ts-list ( x )express bitvec x And use an array to store the mapping between bit vector positions and timestamps, denoted as bvTimeMap The specific implementation of the algorithm can be solved in the following three stages: Phase 1: Mapping Itemsets ts-list pass bvTimeMap The mapping is done to the set of occurrence times across all periods, and then the time set for each period is obtained using PTDS; next, the occurrence time set across all periods for each itemset is frequently processed. minSup Filter to obtain all frequent cycles of the itemset; Phase Two: For itemsets that satisfy frequent periodicity, a threshold based on the neighborhood radius is established. Eps and density threshold minPts and maximum noise ratio threshold maxNRFiltering ensures that the period satisfies the clustering characteristics, resulting in a dataset that includes all intervals where frequent clustering patterns occur; Phase 3: Finally, based on the minimum interval similarity threshold... minSM and minimum periodicity threshold minPR Filter the dataset of frequent clustering patterns for each itemset to ensure that the resulting itemsets have periodic occurrence intervals, meaning that the occurrence intervals occur at least in other... minPR Similar occurrence intervals exist across -1 cycles. By performing the above operation on each itemset, all periodic clustering frequent patterns can be obtained.

[0019] The DBSCAN algorithm was used in stage two of this invention. Eps and minPts All are input parameters for the DBSCAN algorithm.

[0020] The two occurrence intervals of stage three in this invention v and w Similarity is if and only if any of the following conditions are met: v.begin ≤ w.begin and v.end ≥ w.end , recorded as v ∩ w ≤ minSM The intervals intersect; | v.begin - w.begin |≤ minSM And| v.end - w.end |≤ minSM , recorded as w ⊆ v The interval contains.

[0021] In this invention ts-list Based on the structure, it was proposed IPCFPM-list Data structures: Consider l Itemset x , x corresponding IPCFPM-list Represented as IPCFPM-list x , IPCFPM-list x It is a two-level hash table structure, the outer hash structure is ( x , bitvec x ),in, x It is a hash table key ; bitvec x It is a hash table value ,and bitvec xIt is itself a hash structure, represented as ( pid , bitvec x pid ),in, pid ∈ pidSet x , pid for key ,correspond bitvec x pid for value , bitvec x pid Indicates a periodic index pid The period corresponds to the bit vector representation of the time set. Example

[0022] Table 1 is a time-ordered cyclical transaction dataset describing customer shopping basket patterns in retail stores from June to August 2023. a , b , c , d , e and f These represent different products, and the Timestamp indicates the time the transaction occurred. If we use 3 as the frequency threshold to mine frequent itemsets from this dataset, we can use two frequent 1-itemsets { a}and{ f For example, these occurrences are frequent at periodic indices 1, 2, and 3, and will all be mined. However, sometimes we might only want to get results like { f This refers to frequent itemsets that occur in concentrated or clustered locations across multiple periods, with similar clustering intervals (e.g., zongzi leaves and glutinous rice around the Dragon Boat Festival, mooncakes around the Mid-Autumn Festival, and promotional items). It also includes the locations where these itemsets cluster together across multiple periods (e.g., {...}). f The periodic clustering location is from the 1st to the 2nd of each month, which allows for a better understanding of customer needs and improvement of marketing plans.

[0023] 1 2 ,, June 2, 2023 1 3 ,,, June 2, 2023 1 4 ,, June 9, 2023 2 5 ,,,, July 1, 2023 2 6 ,,, July 1, 2023 2 7 ,,,, July 2, 2023 2 8 ,, July 8, 2023 2 9 ,, July 9, 2023 3 10 ,,, August 1, 2023 3 11 ,,, August 1, 2023 3 12 ,, August 2, 2023 3 13 ,, August 2, 2023 3 14 August 7, 2023 3 15 ,, August 8, 2023 Consider the table shown in Table 1 PTDS , with a monthly cycle, 1 itemset { a}and{ f}, and their corresponding combinations { a , f}correspond IPCFPM-list . IPCFPM-list {a} =((1,1011),(2,10101),(3, 010101)), IPCFPM-list{f} =((1,1110),(2, 11100),(3,111000)), IPCFPM-list {a,f} =((1, 1010),(2,10100),(3,010000)).

[0024] bvTimeMapSet Structure as Figure 2 As shown, this is a hash structure. bvTimeMap Based on this, a periodic extension was performed. The periodic index is... key Its corresponding array is value Used to map itemsets IPCFPM-list The mapping is a nested set of periodic occurrence times.

[0025] With 1 item set { a For example,} IPCFPM-list {a} pass bvTimeMapSet Mapping yields PTS {a} = { PTS {a} 1, PTS {a} 2, PTS {a} 3}={[6 / 1,6 / 2,6 / 9],[7 / 1,7 / 2,7 / 9], [8 / 1,8 / 2,8 / 8]}.

[0026] The Naive algorithm uses... IPCFPM-list The structure can efficiently obtain nested sets of periodic times corresponding to itemsets. Furthermore, for periods unlikely to be frequent clustering patterns, the corresponding period index and bit vector can be deleted, reducing the need to perform intersection and repeated frequent clustering pattern checks on these periods, thereby reducing time and space overhead.

[0027] (I) Experimental Environment and Dataset All algorithms used in the experiment were implemented in Python and ran on a Windows 11 22H2 system with an Intel(R) Core(TM) i7-10700 CPU @ 2.90GHz. Two real-world datasets were used: MBA and Value-Inc. MBA data spanned from December 2010 to December 2011, containing 4130 different products and 20640 transactions; Value-Inc data spanned from February 2018 to February 2019, containing 3407 different products and 25898 transactions. The datasets can be found from [link to dataset]. and get.

[0028] (II) Performance Evaluation of IPCFPM Algorithm (1) Algorithm running time analysis First, let's analyze... minSup The effect of changing the value of on the algorithm's running time; other parameter values ​​are as follows: Eps =5, minPts =8, maxNR =0.1, minSM =0, minPR =0.6 algorithm in minSup Comparison of runtime when changes occur, such as Figure 3 As shown.

[0029] Depend on Figure 3 As can be seen, on both datasets, with minSup As the number of algorithms increased, the running time of all algorithms gradually decreased. This is because all three algorithms used the frequent conditional pruning of the Apriori-TID algorithm. minSup The larger the value, the fewer frequent the periodicity of the itemsets, resulting in fewer candidate itemsets. This reduces the algorithm's time complexity and search space, thus reducing the algorithm's runtime. Furthermore, it can be observed on both datasets that, in the same... minSup Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive - NoPrune < Naive. It can also be observed that using... IPCFPM-list The Naive-NoPrune algorithm is better than the Naive algorithm because... IPCFPM-list The structure can reduce unnecessary time consumption during the mining process. The IPCFPM algorithm reduces the running time by 62.4% compared to the Naive algorithm.

[0030] analyze Eps The impact of changes in the value of on the algorithm's running time; other parameter values ​​are as follows: minSup =0.0004, minPts =8, maxNR =0.1, minSM =0, minPR =0.6, the algorithm in Eps Comparison of runtime when changes occur, such as Figure 4 As shown.

[0031] Depend on Figure 4 As can be seen, on both datasets, with Eps As the value increases, the running time of the Naive algorithm and the Naive-NoPrune algorithm increases slightly and gradually, while the running time of the IPCFPM algorithm increases significantly and gradually. This is because the optimization strategy of the IPCFPM algorithm is based on... minCF ,and EpsThe larger the number of timestamps, the more likely they are to cluster together, and the more frequently clustered the itemsets become, or the more likely they are to satisfy Theorem 1. This results in a larger number of candidate itemsets, increasing time complexity and search space, and consequently, the algorithm's runtime. In contrast, the Naive algorithm and the Naive-NoPrune algorithm only address this by frequently pruning... Eps Increasing the value does not change the algorithm's search space. Although the itemsets exhibit more frequent clustering patterns, it only increases the algorithm's time overhead in the periodic intervals. Therefore, the increase in running time for the Naive and Naive-NoPrune algorithms is not significant. Furthermore, it can be observed on both datasets that... Eps Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive, Naive-NoPrune < Naive. Taking the MBA dataset as an example, when... Eps When the value is 5, the IPCFPM algorithm reduces the running time by 54.2% compared to the Naive algorithm.

[0032] analyze minPts The impact of changes in the value of on the algorithm's running time; other parameter values ​​are as follows: minSup =0.0004, Eps =5, maxNR =0.1, minSM =0, minPR =0.6, the algorithm in minPts Comparison of runtime when changes occur, such as Figure 5 As shown.

[0033] Depend on Figure 5 As can be seen, on both datasets, with minPts As the value increases, the running time of the Naive algorithm and the Naive-NoPrune algorithm decreases slightly, while the running time of the IPCFPM algorithm decreases significantly. This is because the IPCFPM algorithm optimization strategy is based on... minCF ,and minPts Larger itemsets can cluster more frequently, resulting in fewer timestamps and less periodicity, thus reducing the number of candidate itemsets, decreasing time complexity and search space, and consequently reducing algorithm runtime. In contrast, the Naive and Naive-NoPrune algorithms only perform frequent pruning. While they reduce the number of itemsets and increase periodicity, this only reduces the time overhead during the periodicity calculation interval, so the Naive algorithm's runtime reduction is not significant. Furthermore, it can be observed on both datasets that... minPts Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive - NoPrune < Naive. Taking the MBA dataset as an example, when minPtsWhen the value is 6, the IPCFPM algorithm reduces the running time by 54.1% compared to the Naive algorithm.

[0034] analyze maxNR The impact of changes in the value of on the algorithm's running time; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, minSM =0, minPR =0.6, the algorithm in maxNR Comparison of runtime when changes occur, such as Figure 6 As shown.

[0035] Depend on Figure 6 As can be seen, on both datasets, with maxNR As the value increases, the running time of both the Naive and Naive-NoPrune algorithms increases slightly and gradually. On the Value-Inc dataset, the running time of the IPCFPM algorithm increases significantly and gradually. On the MBA dataset, the running time of the IPCFPM algorithm first increases significantly and then increases slightly. This is because the optimization strategy of the IPCFPM algorithm is based on... minCF ,and maxNR The larger the dataset, the more noise is allowed, the more frequently clustered itemsets there are, and the more candidate itemsets there are, increasing time complexity and search space, thus increasing the algorithm's runtime. In contrast, the Naive and Naive-NoPrune algorithms only perform frequent pruning. Although the itemsets are more frequently clustered, this only increases the time overhead in the periodicity calculation interval, so the increase in runtime for the Naive algorithm is not significant. Furthermore, it can be observed on both datasets that... maxNR Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive - NoPrune < Naive. Taking the MBA dataset as an example, when maxNR When the value is 0, the IPCFPM algorithm reduces the running time by 66.7% compared to the Naive algorithm.

[0036] analyze minSM The impact of changes in the value of on the algorithm's running time; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, maxNR =0, minPR =0.6, the algorithm in minSM Comparison of runtime when changes occur, such as Figure 7 As shown.

[0037] Depend on Figure 7 As can be seen, on both datasets, with minSM As the value increases, the running time of all algorithms increases slightly and gradually. This is because... minSM When the value changes, it only affects the time overhead of the algorithm during the periodic interval. minSM The larger the value, the more similar occurrence intervals there are. More occurrence intervals can become periodic occurrence intervals, thus increasing the search space and consequently increasing the algorithm's running time. Simultaneously, it can be observed on both datasets that, in the same... minSM Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive - NoPrune < Naive. Taking the MBA dataset as an example, when minSM When the value is 0, the IPCFPM algorithm reduces the running time by 62.3% compared to the Naive algorithm.

[0038] analyze minPR The impact of changes in the value of on the algorithm's running time; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, maxNR =0, minSM =0.6, the algorithm in minSM Comparison of runtime when changes occur, such as Figure 8 As shown.

[0039] Depend on Figure 8 As can be seen, on both datasets, with minPR As the number of algorithms increases, the runtime of all algorithms gradually decreases because... minPR The larger the dataset, the more periods are required for the itemsets to satisfy the frequent clustering pattern. Therefore, fewer itemsets are considered as candidate itemsets, thus reducing time complexity and search space, and consequently reducing the algorithm's runtime. Furthermore, it can be observed on both datasets that, in the same... minPR Under these values, the running times of the three algorithms always satisfy the following relationship: IPCFPM < Naive - NoPrune < Naive. Taking the MBA dataset as an example, when minPR When the value is 0.5, the IPCFPM algorithm reduces the running time by 62.4% compared to the Naive algorithm.

[0040] (2) Algorithm space overhead analysis First, let's analyze... minSup The impact of changing the value of on the algorithm's space overhead; other parameter values ​​are as follows: Eps =5, minPts =8, maxNR =0.1, minSM =0, minPR =0.6 algorithm in minSup Comparison of spatial overhead during change and the number of patterns discovered, such as Figure 9 As shown.

[0041] Depend on Figure 9 As can be seen, on both datasets, with minSup As the number of discovered patterns increases, the space overhead and the number of patterns discovered by all three algorithms gradually decrease. This is because all three algorithms use frequent conditional pruning. minSup The larger the value, the fewer frequent the periodicity of the itemsets, resulting in fewer candidate itemsets. This reduces the search space of the algorithm, thereby reducing its space overhead and the number of patterns discovered. Furthermore, it can be observed on both datasets that, in the same... minSup Under the given values, the space cost of the algorithm always satisfies the following relationship: IPCFPM < Naive < Naive-NoPrune. However, the space cost of the Naive-NoPrune algorithm is always greater than that of the Naive algorithm, contrary to our theoretical analysis. This is because... IPCFPM-list In the two-level hash structure, the keys and the mapping between keys and values ​​also occupy a certain amount of space, and there is a fragmentation problem, causing the Naive-NoPrune algorithm to use more storage space than the Naive algorithm. Taking the MBA dataset as an example, when minSup When the value is 0.0004, the space overhead of the IPCFPM algorithm is reduced by 16.9% compared to the Naive algorithm.

[0042] analyze Eps The impact of varying values ​​of on the algorithm's space overhead; other parameter values ​​are as follows: minSup =0.0004, minPts =8, maxNR =0.1, minSM =0, minPR =0.6, the algorithm in Eps Comparison of spatial overhead and number of patterns discovered during changes Figure 10 As shown.

[0043] Depend on Figure 10 As can be seen, on both datasets, with Eps As the value increases, the space overhead and the number of discovered modes in all algorithms gradually increase. This is because the optimization strategies of the IPCFPM algorithm are all based on... minCF The Naive and Naive-NoPrune algorithms, through frequent pruning, show no significant increase in space consumption compared to the IPCFPM algorithm. Furthermore, it can be observed on both datasets that... Eps Under these values, the algorithm's space cost always satisfies the following relationship: IPCFPM < Naive < Naive-NoPrune. Taking the MBA dataset as an example, when... Eps When the value is 1, the space overhead of the IPCFPM algorithm is reduced by 16.8% compared to the Naive algorithm.

[0044] analyze minPts The impact of varying values ​​of on the algorithm's space overhead; other parameter values ​​are as follows: minSup =0.0004, Eps =5, maxNR =0.1, minSM =0, minPR =0.6, the algorithm in minPts Comparison of spatial overhead during change and the number of patterns discovered, such as Figure 11 As shown.

[0045] Depend on Figure 11 As can be seen, on both datasets, with minPts As the value increases, the space overhead and the number of discovered patterns in all algorithms gradually decrease. This is because the optimization strategy of the IPCFPM algorithm is based on... minCF The Naive and Naive-NoPrune algorithms only reduce space consumption significantly by frequently pruning. Furthermore, it can be observed on both datasets that... minPts Under these values, the algorithm's space cost always satisfies the following relationship: IPCFPM < Naive < Naive-NoPrune. Taking the MBA dataset as an example, when... minPts When the value is 6, the space overhead of the IPCFPM algorithm is reduced by 13.6% compared to the Naive algorithm.

[0046] analyze maxNR The impact of varying values ​​of on the algorithm's space overhead; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, minSM =0, minPR =0.6, the algorithm in maxNR Comparison of spatial overhead and number of patterns discovered during changes Figure 12 As shown.

[0047] Depend on Figure 12 As can be seen, on both datasets, with maxNR As the value increases, the space overhead and the number of discovered modes in all algorithms gradually increase. This is because the optimization strategies of the IPCFPM algorithm are all based on... minCF The Naive and Naive-NoPrune algorithms only perform frequent pruning, so their space consumption increase is not significant. Furthermore, it can be observed on both datasets that... maxNRAt any given value, the Naive-NoPrune algorithm consistently has the highest space overhead, followed by the Naive algorithm. The other three algorithms have almost identical space overhead. Based on specific numerical comparisons, the space overhead of the algorithms always follows this relationship: IPCFPM < Naive < Naive-NoPrune. Taking the MBA dataset as an example, when... maxNR When the value is 0, the space overhead of the IPCFPM algorithm is reduced by 18.8% compared to the Naive algorithm.

[0048] analyze minSM The impact of varying values ​​of on the algorithm's space overhead; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, maxNR =0, minPR =0.6, the algorithm in minSM Comparison of spatial overhead during change and the number of patterns discovered, such as Figure 13 As shown.

[0049] Depend on Figure 13 As can be seen, on both datasets, with minSM As the value increases, the space overhead of all algorithms and the number of patterns discovered gradually increase. This is because... minSM When the value changes, it only affects the interval in which the algorithm calculates the period. minSM The larger the value, the more similar occurrence intervals there are. More occurrence intervals can become periodic occurrence intervals, thus increasing the search space and consequently increasing the algorithm's space overhead and the number of patterns discovered. Simultaneously, it can be observed on both datasets that, in the same... minSM Under these values, the algorithm's space cost always satisfies the following relationship: IPCFPM < Naive < Naive-NoPrune. Taking the MBA dataset as an example, when... minSM When the value is 0, the space overhead of the IPCFPM algorithm is reduced by 16.6% compared to the Naive algorithm.

[0050] analyze minPR The impact of varying values ​​of on the algorithm's space overhead; other parameter values ​​are as follows: minSup =0.0004, Eps =5, minPts =8, maxNR =0, minSM =0.6, the algorithm in minSM Comparison of spatial overhead during change and the number of patterns discovered, such as Figure 14 As shown.

[0051] Depend on Figure 14 As can be seen, on both datasets, with minPRAs the value increases, the space overhead of all algorithms and the number of patterns discovered gradually decrease. This is because... minPR The larger the dataset, the more periods are required for the itemsets to satisfy the frequent clustering pattern, resulting in fewer itemsets as candidate itemsets. This reduces the search space, thereby reducing the algorithm's space consumption and the number of patterns discovered. Furthermore, it can be observed on both datasets that, in the same... minPR Under these values, the algorithm's space cost always satisfies the following relationship: IPCFPM < Naive < Naive-NoPrune. Taking the MBA dataset as an example, when... minPR When the value is 0.5, the space overhead of the IPCFPM algorithm is reduced by 17.1% compared to the Naive algorithm.

[0052] (3) Summary Simulation experiments based on two real-world datasets demonstrate the effectiveness of the IPCFPM algorithm. Compared with the Naive algorithm, the ICFPM algorithm significantly improves in terms of time and space efficiency, making it an effective technique for mining frequent patterns of periodic clustering.

[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for mining frequent patterns of periodic clustering in time-ordered periodic transaction data, characterized by: The process includes the following steps: Step 1: First, use the Apriori-TID algorithm to determine whether each period of the itemset is frequent; Step 2: Find the set of occurrence times corresponding to each period that is frequent, and use the DBSCAN clustering algorithm on the set of occurrence times of each frequent period to determine whether they are clustered. If they are clustered, the corresponding cluster occurrence interval is obtained; Step 3: Perform similarity calculations between the cluster occurrence intervals of all periods. If the cluster occurrence interval of a certain period is similar to the cluster occurrence intervals of other periods, and the total number of periods containing similar cluster occurrence intervals is greater than a given threshold, then the cluster occurrence interval is the periodic occurrence interval of the itemset, and the itemset is the periodic clustering frequent pattern; Step 4: By performing the above judgment process on all itemsets, all periodic clustering frequent patterns can be found. This method is the Naive algorithm. The data structure uses a ts-list structure, which uses bit vectors to represent the set of times occurring over all periods. A ts-list is a hash table with structure (x, bitvecx), where bitvecx represents the bit vector corresponding to x. Let ts-listx represent the ts-list corresponding to x, and ts-list(x) represent bitvecx. An array, denoted as bvTimeMap, is used to store the mapping between bit vector positions and timestamps. The specific algorithm implementation can be solved in the following three stages: Phase 1: Map the itemset to the corresponding ts-list using bvTimeMap to the set of times that occur in all periods. Then, obtain the time set for each period based on PTDS. Next, perform frequent minSup filtering on the time set of occurrences in all periods of each itemset to obtain all frequent periods of the itemset. Phase 2: The set of occurrence times corresponding to the frequent period of itemsets is filtered based on the neighborhood radius threshold Eps, the density threshold minPts, and the maximum noise ratio threshold maxNR to ensure that the period satisfies the clustering characteristics, resulting in a dataset containing all the frequent occurrence intervals of clustering patterns. Phase 3: Finally, the dataset of frequent clustering pattern intervals corresponding to the itemsets is filtered based on the minimum interval similarity threshold minSM and the minimum periodicity occurrence threshold minPR to ensure that the obtained itemsets contain periodic occurrence intervals, that is, the occurrence intervals have similar occurrence intervals at least in the other minPR-1 periods. By performing the above operation on each itemset, all periodic clustering frequent patterns can be obtained.

2. The method for mining frequent clustering patterns of time-ordered periodic transaction data according to claim 1, characterized in that: By summarizing the patterns of unnecessary operations in the Naive algorithm through Lemmas 1 and 2, and Theorem 1 and Theorem 2, the existence rules of unnecessary operations in the Naive algorithm are effectively reduced based on these rules: Lemma 1: For itemsets x 'and x correspond PTS x’ and PTS x Medium-term index is u members PTS x’ u and PTS x u Considering PTS x’ u and PTS x u After division C x’ u and C x u ,if x’ ⊂ x Then we have | C x u |≤| C x’ u |Established; Lemma 2: Let minCF = minSup ×| PTDS |×(1- maxNR ), | PTDS |for PTDS Includes the total number of transactions, given itemsets x about PTS x Medium-term index is u members PTS x u ,right PTS x u After division C x u We have the following conclusion: if | C x u |< minCF ,So x Not a frequent clustering pattern; Theorem 1: For itemsets x 'about PTS x’ Medium-term index is u members PTS x’ u If | C x' u |< minCF Then any superset x ⊃ x 'exist u The clustering pattern is not frequent in any of the cycles; Theorem 2: For itemsets x ',make cfPidset x’ Indicates satisfaction minCF The set of periodic indexes, | cfPidset x’ | indicates the total number of periodic indexes included. cfPidset x’ | / | pidSet |< minPR Then any superset x ⊃ x Neither of them exhibits a frequent clustering pattern.

3. The method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to claim 1, characterized in that: Phase two uses the DBSCAN algorithm. Eps and minPts All are input parameters for the DBSCAN algorithm.

4. The method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to claim 3, characterized in that: The three stages occur in two intervals. v and w Similarity is if and only if any of the following conditions are met: v.begin ≤ w.begin and v.end ≥ w.end , recorded as v ∩ w ≤ minSM, Intersecting intervals; | v.begin - w.begin |≤ minSM And| v.end - w.end |≤ minSM , recorded as w ⊆ v, The interval contains.

5. The method for mining frequent patterns of periodic clusters in time-ordered periodic transaction data according to claim 3, characterized in that: exist ts-list Based on the structure, it was proposed IPCFPM-list Data structures: Consider l Itemset x , x corresponding IPCFPM-list Represented as IPCFPM-list x , IPCFPM-list x It is a two-level hash table structure, the outer hash structure is ( x , bitvec x ),in, x It is a hash table key ; bitvec x It is a hash table value ,and bitvec x It is itself a hash structure, represented as ( pid , bitvec x pid ),in, pid ∈ pidSet x , pid for key ,correspond bitvec x pid for value , bitvec x pid Indicates a periodic index pid The period corresponds to the bit vector representation of the time set.