Data stream-oriented efficient pattern mining heterogeneous parallel acceleration method
Through the ARM+FPGA architecture and TWU pruning strategy, combined with dynamic and static task allocation, the problem of existing technologies not considering the utility value of data on heterogeneous data stream architectures is solved, and the computing performance and iteration efficiency of high-utility item set mining are improved.
Patent Information
- Application Number
- CN202510893685.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-14
AI Technical Summary
When existing technologies accelerate data mining algorithms on heterogeneous architectures oriented to data streams, they do not consider the actual utility value of the data, resulting in information loss.
Adopting ARM+FPGA architecture, the single item utility is obtained by scanning the original database, and a calculation auxiliary structure is constructed. The TWU pruning strategy and four parallel IP cores are used, combined with a dynamic and static task allocation strategy, to calculate the high utility of single items and item sets.
The computational performance and iteration efficiency of the high-utility itemset mining algorithm on heterogeneous architectures are improved, storage resource constraints and data access efficiency are met, and high-utility pattern mining is achieved.
Smart Images

Figure CN120780756A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining and heterogeneous algorithm acceleration, and in particular relates to a high-utility pattern mining heterogeneous parallel acceleration method for data streams. Background Art
[0002] With the rapid development of the information age, efficiently mining valuable information from massive, multi-source, and complex data has become a critical research topic in data analysis, giving rise to the field of data mining. High-utility itemset mining, a cutting-edge area of data mining, aims to extract itemsets with greater practical utility from datasets and extract valuable insights. In recent high-utility itemset mining algorithms, the construction of the utility list structure often consumes the majority of computational time. Furthermore, with the continuous growth of data size, traditional high-utility itemset mining methods face challenges such as high computational complexity, high storage overhead, and insufficient real-time performance. This is particularly true in IoT edge node scenarios with limited computing resources, where efficient preprocessing and mining of massive data has become an urgent need.
[0003] Existing data mining algorithms designed to replicate or accelerate data streams on heterogeneous architectures only consider the frequency of itemsets. The basic approach is to deploy and accelerate the recurring frequency calculation module within the algorithm on an FPGA (Field Programmable Gate Array) platform, while placing the flow control and itemset iteration command modules on an ARM (Advanced RISC Machines) or PC (Personal Computer) platform. During algorithm execution, the ARM or PC platform scans the stream-structured database only once and then calls the FPGA platform's frequency calculation module to mine highly frequent itemsets in the heterogeneous architecture.
[0004] In 2015, Bustio et al. ("Bustio-Martínez, Lázaro and Hernández-León, Raudel and Cumplido, René, et al. A Hardware-Based Approach for Frequent Itemset Mining in Data Streams[C] / / 2015") used a contraction tree to implement FIM for data streams. This method extracts frequent itemsets with a single database traversal and zero latency. In this model, each path corresponds to an itemset, and each node records the frequency of the current itemset. The following year, Bustio et al. expanded on their previous research in the paper "Bustio-Martínez, Lázaro and Hernández-León, Raudel and Cumplido, René, et al. Frequent Itemsets Mining in Data Streams Using Reconfigurable Hardware[C] / / 2020 IEEE 38th International Conference on Computer Design (ICCD). 2016:32-45." They added a data stream preprocessing stage to detect the top-k frequent items and proposed an algorithm suitable for a sliding window model. By selecting the top-k frequent items, this method significantly improves the utilization of the contraction tree. In the paper "Hoang, Trong-Thuc and Nguyen, Xuan-Thuan and Nguyen, Hong-Thu, et al. FPGA-based frequent items counting using matrix of equality comparators[C / OL] / / 2017IEEE 60th International Midwest Symposium on Circuits and Systems (MWSCAS). 2017:285-288. DOI: 10.1109 / MWSCAS.2017.8052916," Hoang et al. proposed a data stream single-item mining architecture that does not rely on space-saving techniques. The core idea is to utilize matrix-based equality comparators to directly compare all input items, completing the count in a single clock cycle and significantly improving computational speed. Furthermore, the method uses DMA storage to accelerate computation.
[0005] The above methods use frequent itemset mining algorithms when porting or accelerating data stream-oriented data mining algorithms on heterogeneous architectures. These algorithms only consider the frequency of itemsets in the database, not their actual utility value, leading to information loss in some application scenarios. Summary of the Invention
[0006] To address the above-mentioned problems in the prior art, the present invention provides a data stream-oriented high-utility pattern mining heterogeneous parallel acceleration method. The technical problem to be solved by the present invention is achieved through the following technical solutions: The present invention provides a data stream-oriented high-utility pattern mining heterogeneous parallel acceleration method, comprising: Step 1: Scan the original database, which includes multiple transactions, obtain the utility of each item in each transaction and construct a calculation auxiliary structure; Step 2: Calculate the residual utility of the single item in each transaction according to the calculation auxiliary structure, and store the utility and residual utility of the single item in each transaction in the DDR; Step 3: Prune the constructed TWU value hash table according to the TWU pruning strategy to obtain a pruned TWU value hash table; Step 4: Using four parallel IP cores, based on the utility and residual utility of the items stored in the DDR in each transaction, a polling method is used to calculate the high utility of each item in the pruned TWU value hash table, and the high-utility item is determined based on the calculation result of the high utility of the item. Step 5: Divide the four parallel IP cores into a parallel first IP core group and a second IP core group, each IP core group includes two parallel IP cores, and adopt a dynamic and static combined task allocation strategy to allocate all extended item sets in the pruned TWU value hash table to the first IP core group and the second IP core group. Calculate the high utility of the item set based on the utility and residual utility of the single item stored in the DDR in each transaction, and determine the high-utility item set based on the calculation result of the high utility of the item set.
[0007] Compared with the prior art, the present invention has the following beneficial effects: This heterogeneous parallel acceleration method for high-utility pattern mining in data streams fully considers storage resource constraints, data access efficiency, and computational task balance. Through hybrid task scheduling, storage optimization, efficient data exchange mechanisms, and a customized parallelizable hotspot code IP core, it effectively improves the computational performance and iteration efficiency of the PFUM-Miner (high-utility itemsets mining based on prefix-free and utility-matrix) high-utility itemset mining algorithm on heterogeneous architectures, laying a solid foundation for subsequent heterogeneous porting research on high-utility itemset mining.
[0008] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 This is a flow chart of a heterogeneous parallel acceleration method for high-utility pattern mining for data streams provided by an embodiment of the present invention; Figure 2 This is a task allocation diagram of a data stream-oriented high-utility pattern mining heterogeneous parallel acceleration method provided by an embodiment of the present invention; Figure 3 It is a data structure initialization diagram provided by an embodiment of the present invention; Figure 4 This is an example diagram of calculating the residual utility of a single item provided by an embodiment of the present invention; Figure 5 1 is a schematic diagram of the architecture of a parallel heterogeneous accelerator provided by an embodiment of the present invention; Figure 6 This is a schematic diagram of a register-to-four-way IP core command segmentation provided by an embodiment of the present invention; Figure 7 This is a schematic diagram of a burst control module provided by an embodiment of the present invention; Figure 8 This is a schematic diagram of a storage process for the results of a single high-utility calculation process provided by an embodiment of the present invention; Figure 9 1 is a schematic diagram of an expanded order of an itemset provided by an embodiment of the present invention; Figure 10 This is a schematic diagram of a parallel computing task allocation framework provided by an embodiment of the present invention; Figure 11is a state transition relationship and condition schematic diagram provided by an embodiment of the present application; Figure 12 is a comparison schematic diagram of algorithm running time and speedup ratio in the relatively sparse test1 and test2 databases, wherein the left graph is the test1 database, and the right graph is the test2 database; Figure 13 is a comparison schematic diagram of algorithm running time and speedup ratio in the most dense Connect and Chess databases, wherein the left graph is the Connect database, and the right graph is the Chess database; Figure 14 is a comparison schematic diagram of algorithm running time and speedup ratio in the Mushroom and T20I20D10K databases with medium sparsity, wherein the left graph is the Mushroom database, and the right graph is the T20I20D10K database; Figure 15 is a comparison schematic diagram of algorithm iteration number and compression ratio in the relatively sparse test1 and test2 databases, wherein the left graph is the test1 database, and the right graph is the test2 database; Figure 16 is a comparison schematic diagram of algorithm iteration number and compression ratio in the most dense Connect and Chess databases, wherein the left graph is the Connect database, and the right graph is the Chess database; Figure 17 is a comparison schematic diagram of algorithm iteration number and compression ratio in the Mushroom and T20I20D10K databases with medium sparsity, wherein the left graph is the Mushroom database, and the right graph is the T20I20D10K database; Figure 18 is a proportion of a lookup table resource and a register resource in the overall resource of a heterogeneous accelerator in an architecture provided by an embodiment of the present application. DETAILED DESCRIPTION
[0010] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined invention purpose, a data flow-oriented efficient pattern mining heterogeneous parallel acceleration method according to the present application is described in detail below in combination with the drawings and specific embodiments.
[0011] The aforementioned and other technical contents, features, and effects of the present invention are clearly presented in the following detailed description of the specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a deeper and more specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are provided for reference and illustration purposes only and are not intended to limit the technical solutions of the present invention.
[0012] The embodiment of the present invention provides a data stream-oriented high-utility pattern mining heterogeneous parallel acceleration method, see Figure 1 , Figure 1 This is a flow chart of a method for high-utility pattern mining and heterogeneous parallel acceleration for data streams provided by an embodiment of the present invention. The method for high-utility pattern mining and heterogeneous parallel acceleration for data streams of this embodiment is implemented using an ARM+FPGA architecture. Figure 1 As shown, the following steps may be included: Step 1: Scan the original database, which includes multiple transactions, obtain the utility of each item in each transaction and build a calculation auxiliary structure.
[0013] In this example, the ARM architecture reads a txt file containing the original database on an SD (Secure Digital) card. The FATFS library function is used to read rows (i.e., a transaction in the original database). Therefore, the FATFS library must be mounted first, and only after successful installation can the next step be performed.
[0014] Optionally, the data stream-oriented window size of the acceleration method is 1024, that is, the number of transactions included in the original database of the embodiment.
[0015] In this embodiment, the data in the txt file is read by line (transaction), and each transaction includes multiple items, the utility of the items in the transaction, and the total utility of the transaction in the line.
[0016] In this embodiment, the calculation auxiliary structure includes a transaction utility vector, a TWU (Transaction Weighted Utilization) value hash table, and a position vector structure.
[0017] The transaction utility vector is used to record the total transaction utility of each transaction in the original database. In this embodiment, the transaction utility vector is recorded as tV.
[0018] The TWU value hash table is used to record the TWU value of each item in the original database, and the TWU value hash table is sorted in ascending order according to the TWU value of the item. In this embodiment, the TWU value hash table is denoted as twuM, and the TWU value of an item is the cumulative total transaction utility of all transactions in the original database that include the item.
[0019] The position vector structure is used to record the base address in DDR (Double Data Rate SDRAM) corresponding to each item, divided according to the order in which each item appears in the original database. A consecutive 1024 addresses in DDR constitute the address range corresponding to each item, and the base address is the starting point of the address range corresponding to each item. In this embodiment, the position vector structure is denoted by 1M.
[0020] Step 2: Based on the calculation auxiliary structure, calculate the residual utility of the single item in each transaction, and store the utility and residual utility of the single item in each transaction in the DDR.
[0021] In this embodiment, step 2 includes: Step 2.1: For each transaction, calculate the residual utility of each item in the current transaction based on the total utility of the current transaction and the utility of each item in the current transaction, according to the sorting order of each item in the TWU value hash table; Step 2.2: For each address range in the DDR corresponding to an item, an address in the address range stores the utility and remaining utility of the current item in a transaction. For transactions that do not include the current item, the corresponding address remains 0. The 32-bit portion of each address is split into the upper 16 bits and the lower 16 bits. The upper 16 bits store the utility of the current item in the current transaction, and the lower 16 bits store the remaining utility of the current item in the current transaction.
[0022] For example, the calculation process can be expressed as follows: call the total utility of the current transaction, subtract the utility value of the first item in the TWU value hash table included in this transaction from this value, and obtain the residual utility of this item, which is stored in the corresponding position in the DDR. Then, subtract the utility value of the second item in the TWU value hash table included in the transaction from this residual utility to obtain the residual utility of the second item. This process is repeated in this way. After obtaining the total utility of each transaction, the utility values of all items included in the transaction should be subtracted in turn to obtain the corresponding residual utility, until the calculation of all transactions and items is completed.
[0023] Step 3: According to the TWU pruning strategy, prune the constructed TWU value hash table to obtain a pruned TWU value hash table.
[0024] In this embodiment, step 3 includes: comparing the TWU value of each single item in the TWU value hash table with a preset threshold, removing single items in the TWU value hash table that are less than the threshold, and obtaining a pruned TWU value hash table.
[0025] According to the set TWU pruning strategy, if the TWU value of a single item is less than the threshold, then the single item and its extended itemset have no possibility of becoming a high-utility itemset. To reduce useless calculations, all single items with TWU values less than the threshold are discarded.
[0026] Step 4: Using four parallel IP cores, based on the utility and residual utility of the items stored in the DDR in each transaction, a polling method is used to calculate the high utility of each item in the pruned TWU value hash table. The high-utility item is determined based on the calculation result of the high utility of the item.
[0027] In this embodiment, four parallel IP (Intellectual Property) cores are designed to perform single utility calculations. The ARM side, as the main processor, is responsible for configuring the IP cores and analyzing the returned calculation results.
[0028] Specifically, each IP core is provided with a first BRAM (Block Random Access Memory) and a second BRAM. Accordingly, step 4 includes: Step 4.1: Call four parallel IP cores and use polling to parallelly calculate the single-item high utility of four consecutive single items in the pruned TWU value hash table. In each round of single-item high utility calculation, the IP core is configured using the ARM side. If the number of single items in the last calculation is less than four, the required number of IP cores is called.
[0029] In this embodiment, four consecutive items in the TWU value hash table are calculated in parallel each time, and the next four consecutive items are selected in the next calculation. If the number of items in the last calculation does not meet the requirement of four, the required number of IP cores are called. Regardless of whether they are called in the last round, the calculation result in the current result register is returned to the IP command.
[0030] Because the IP core is called to perform single-item utility calculations during this process, each IP core needs to be configured. In this embodiment, the configuration details of the IP core configured by the ARM side during each round of single-item high-utility calculations include: the base address of the item read from the DDR, the user-preset threshold, the IP core call enable, and the end-of-calculation flag set to be set if the current item is the last item calculated by the IP core.
[0031] The raised end calculation flag can be understood as a signal to tell the IP core that this is the last round of calculation, so that the IP core returns all previous results after the calculation is completed.
[0032] Step 4.2: After each round of calculation, the calculation result of the single high utility is stored in the result register corresponding to the IP core in a left-shifted manner. When the result register is full, the data in it is written to the first BRAM. When the IP core receives the end calculation flag, after the last round of calculation, the data in the result register is written to the first BRAM, and the validity flag is stored in the second BRAM. For the IP core that is not called in the last round of calculation, the data in the corresponding result register is written to the first BRAM, and the validity flag is stored in the second BRAM.
[0033] Because the data bit width at each BRAM address is 32 bits, to maximize bit width utilization, each IP core stores calculation results in groups of 32. Specifically, each time the IP core calculates a single item, it stores the result as the least significant bit at the current address in the corresponding result register. The original 32-bit data is shifted left sequentially, and the most significant bit is discarded. Each group of 32 data is written to the BRAM. If the last data arrives, the current data is written to the BRAM regardless of whether the group of 32 is complete.
[0034] Exemplarily, after each round of calculation, the 1-bit calculation result of the single high utility is stored in the result register corresponding to the IP core in a left-shift manner. After 32 rounds of single high utility calculation, the 32-bit data containing 32 single high utilities are written into the first BRAM. When the IP core receives the end calculation flag, regardless of whether the data in the current result register has updated the 32-bit calculation result, the data in the result register will be written into the first BRAM after the last round of calculation, and the validity flag will be stored in the second BRAM; if the IP core has not been called for calculation when the end calculation flag is received, the data in the result register at this time will be written into the first BRAM, and the validity flag will be stored in the second BRAM.
[0035] Step 4.3: The ARM side determines whether the calculation of all items in the pruned TWU value hash table is completed based on the pruned TWU value hash table pointer. If the parallel calculation of the current round is completed but there are still items in the pruned TWU value hash table that have not been calculated, return to step 4.1; if the parallel calculation of the current round is completed and all items in the pruned TWU value hash table are calculated, execute step 4.4; Step 4.4: The ARM side reads the contents of the first BRAM and the second BRAM according to the validity flag, analyzes the calculation result of the single item high utility, and determines and records the single item with high utility.
[0036] In the embodiment, the utility of the single item is calculated by accumulating the utility of the single item in all transactions, comparing the accumulated sum with the threshold set by the user, and if the accumulated sum is greater than or equal to the threshold, the result of the single item utility calculation of the single item is 1, that is, the single item is a high-utility single item, otherwise the result of the single item utility calculation of the single item is 0, that is, the single item is not a high-utility single item.
[0037] Step 5: The four parallel IP cores are divided into a parallel first IP core group and a second IP core group, each IP core group including two parallel IP cores, all the expanded item sets in the pruned TWU value hash table are distributed to the first IP core group and the second IP core group by using a dynamic-static combined task distribution strategy, the high utility of the item set is calculated according to the utility and the remaining utility of the single item in each transaction stored in the DDR, and the high-utility item set is determined according to the calculation result of the item set utility.
[0038] In the embodiment, the item set utility is still calculated by the designed four parallel IP cores. Specifically, step 5 includes: Step 5.1: The four parallel IP cores are called, the first expanded item set of the first single item in the pruned TWU value hash table is distributed to the first IP core group to calculate the item set utility, the expanded item sets of the other single items except the first single item in the pruned TWU value hash table are distributed to the second IP core group to calculate the item set utility, and inside the first IP core group and the second IP core group, a dynamic task distribution method is used, when any IP core completes the calculation of the item set under the current branch, the next branch of the item set calculation task is taken according to the order of the item set in the IP core group until the iteration of all item sets in the IP core group ends, and in each calculation of the item set utility, the IP core is configured by the ARM end.
[0039] In the embodiment, the configuration content of the IP core configured by the ARM end in each calculation of the item set utility includes: the base address of the single item read by the IP core from the DDR, the threshold set by the user, the IP core call enable, and the loc signal and the itr signal for assisting the calculation of the item set utility.
[0040] Step 5.2: After each calculation, the calculation result of the item set utility is stored in the second BRAM, and according to the calculation result of the item set utility, the utility and the remaining utility of the item set with high utility and expansion potential are stored in the first BRAM.
[0041] Step 5.3: After each calculation, use the ARM side to read the contents of the second BRAM and parse the calculation results of the item set's high utility. Determine and record the high-utility item set. If all items in the pruned TWU value hash table are expanded and calculated, the algorithm ends. Otherwise, expand the item set with expansion potential and return to step 5.1.
[0042] In this embodiment, the prefix-free item set expansion method is used to expand the item sets of all individual items in the pruned TWU value hash table. For the specific prefix-free item level expansion method, please refer to the Chinese patent document "High-Utility Item Set Mining Method and Device" with patent application number 202411207964.3 and publication number CN119179727A, which will not be described in detail here. It should be noted that during the expansion process, if the current extended item set {Pxy} contains m individual items, its connected item set should include {Px} and {y}, and the extended item set that each IP core is responsible for calculating is expanded according to the expansion method maintained by pointers and vectors proposed in the CN119179727A patent. In this embodiment, the utility and residual utility of {y} are obtained from the DDR, while the utility and residual utility of {Px} are obtained from the first BRAM. The storage structure of the first BRAM can be considered a regular design, that is, the zeroth group of utility values must correspond to a single item, the first group of utility values corresponds to an item set containing two single items, and so on. Since the utility and residual utility of {Px} must be stored in the (m-2)th group in the first BRAM, then: loc = sizeof({Px}-1) = (m - 2), which is the location of {Px} in the first BRAM; itr = sizeof({Pxy}-1) = (loc + 1) = (m - 1), that is, when {Pxy} has expansion potential, it should be stored in the (m-1)th group of the first BRAM.
[0043] In this embodiment, the calculation result of the itemset high utility is a 3-bit data, including a 1-bit judgment result of the itemset high utility, a 1-bit judgment result of the itemset expansion potential, and a 1-bit validity flag. The calculation of the itemset high utility is performed by summing the utility of the itemset in all transactions and comparing the sum with a user-set threshold. If the sum is greater than or equal to the threshold, the calculation result of the itemset high utility is 1, indicating that the itemset is a high-utility itemset. Otherwise, the calculation result of the itemset high utility is 0, indicating that the itemset is not a high-utility itemset. The calculation of the itemset expansion potential is performed by summing the utility and residual utility of the itemset in all transactions and comparing the sum with a user-set threshold. If the sum is greater than or equal to the threshold, the calculation result of the itemset expansion potential is 1, indicating that the itemset has expansion potential. Otherwise, the calculation result of the itemset expansion potential is 0, indicating that the itemset does not have expansion potential.
[0044] It should be noted that after the calculation of the high utility of each item set is completed, a judgment is first made based on the calculation result of the high utility of the item set, and the utility and residual utility of the item set with expansion potential are stored in the first BRAM. At the same time, the prefix-free item set expansion method is used to expand it and then return to step 5.1 for the next round of calculation until all single items in the pruned TWU value hash table are expanded and calculated.
[0045] It is worth noting that in this embodiment, after the IP core is configured using the ARM side, when the IP core receives an instruction to read data from the DDR, it performs the first burst according to the configured base address of the single item read from the DDR, and automatically enables a new burst read and updates the base address after the burst transmission is completed, until the data in the address range corresponding to a single item in the DDR is read four times in a burst.
[0046] The heterogeneous parallel acceleration method for high-utility pattern mining in data streams, developed in this embodiment, fully considers storage resource constraints, data access efficiency, and computational task balance. Through hybrid task scheduling, storage optimization, efficient data exchange mechanisms, and a customized parallelizable hotspot code IP core, it effectively improves the computational performance and iteration efficiency of the PFUM-Miner algorithm for high-utility itemset mining on heterogeneous architectures, laying a solid foundation for subsequent heterogeneous porting research on high-utility itemset mining.
[0047] In order to facilitate understanding of the heterogeneous parallel acceleration method for high-utility pattern mining of data streams of the present invention, a set of examples are provided to illustrate the method of the present invention. For specific algorithm flow and task allocation, see Figure 2 The task allocation diagram of the high-utility pattern mining heterogeneous parallel acceleration method for data flow is shown.
[0048] The original database is shown in Table 1 below, and the minimum utility threshold value, i.e., the threshold value preset by the user, is 20.
[0049] Table 1. Original database
[0050] First, the original database is scanned by row (by transaction), and the utility of the single item included in the transaction is stored in the DDR utility matrix according to the read order, while tV, twuM and lM are constructed. Since the accelerator faces the data stream and the sliding window size is 1024, in the DDR, the adjacent 1024 addresses are used to store the utility and the remaining utility of a single item in all transactions, and the upper 16 bits of 32-bit data under each address are used to store the utility of the single item in the transaction, and the lower 16 bits store the remaining utility of the single item. Therefore, the specific operation is shown in Table 2. Figure 3 , Figure 3 is a data structure initialization diagram provided by an embodiment of the application.
[0051] In Figure 3 , the scanning mode of transaction T1 is shown, first record its transaction total utility, then record the utility of the single item in the transaction in turn according to the order of the single item read, the TWU value of the single item and the storage order of each single item in the DDR. Since the DDR in the PS (Processing System) end of the FPGA stores the operating system, application program and other necessary data, the actual available space is much smaller than the theoretical value, and the design of discarding the first row record of the total utility in the original utility matrix is abandoned to maximize the utilization rate of the storage unit. Therefore, after scanning the original database, the above four data structures store data as shown in Tables 2, 3, 4 and 5.
[0052] Table 2. Utility matrix in DDR stores the utility of single item in transaction
[0053] Table 3. Transaction utility vector
[0054] Table 4. TWU value hash table
[0055] Table 5. Position vector structure
[0056] The threshold value preset by the user is 20, so no single item is discarded in the TWU pruning operation. If there is, the utility value of the discarded single item will also delete the utility in the transaction in Table 2.
[0057] Next, the residual utility of the item needs to be calculated in ascending order of the TWU value, as shown in the following formula: Figure 4 The following table 6 shows an example of the calculation of the residual utility of an item according to an embodiment of the present application. The final utility matrix in DDR stores the utility and residual utility of the item in the transaction as shown in table 6.
[0058] Table 6. Utility matrix in DDR stores the utility and residual utility of the item in the transaction
[0059] In the design process of the heterogeneous accelerator, the PL (Programmable Logic) IP core of FPGA is selected to calculate the item and the item set in parallel, which is efficient. The ARM is used as the main processor, the AXI-HP (Advanced eXtensible Interface-High-performance Purpose) protocol is used to realize the read and write of the DDR unit, and the AXI-GP (Advanced eXtensible Interface-General Purpose) protocol is used to realize the read and write of the BRAM unit and the control of the parallel IP core. The architecture of the parallel heterogeneous accelerator is shown in the following table 7. Figure 5
[0060] The AXI-GP channel for controlling the four parallel IP cores contains eight registers, each of which contains 32 bits. These registers are used to distinguish different start configuration commands issued by the ARM. Among them: 1) slave_register_0-slave_register_3: four parallel IP core address control registers, used to determine the current IP core reading DDR data base address; 2) slave_register_4: threshold configuration register, which is convenient for calling the minimum utility threshold set by the user in the IP core calculation process; 3) slave_register_5 and slave_register_6: iteration control registers, used for state machine control of the IP core in the high utility calculation process of the item set and BRAM address control of the IP core; 4) slave_register_7: mode configuration and start register, containing IP core call, end data flag and IP core function selection three functions.
[0061] The following table 7 shows the names and bit widths of all functional signals.
[0062] Table 7. Register command function
[0063] In addition, since a four-way parallel computing IP core is designed, in order to prevent conflicts or confusion during command transmission, it is necessary to configure corresponding commands for each of the four IP cores. The command allocation of the four IP cores in the eight registers is as follows: Figure 6 shown.
[0064] The maximum burst length of the AXI protocol is 256, but the data stream window size designed by the present invention is 1024. If a strategy is adopted to feed back to the control end after each burst to request the next burst to be enabled, the overhead of the number of interactions will be greatly increased at the expense of running time. Therefore, a continuous burst control module is designed. When the IP core receives an instruction from the PS end to read data from the DDR, it will perform the first burst according to the DDR base address read in the corresponding register (slave register 0-4), and automatically enable a new burst read and update the base address after each burst transmission is completed until all four burst reads are completed. This mechanism greatly improves the efficiency of data transmission. In order to facilitate the IP to read and write DDR, a burst control module will be calculated and configured for each IP core. Figure 7 Taking the burst control module shown as an example, when the first burst address is 0x00000000, the four consecutive burst read addresses are: 0x00000000, 0x00000400, 0x00000800, and 0x00000C00.
[0065] According to the above design, we continue to calculate the high utility item sets of the original database in Table 1. The parallel IP core should calculate the high utility of the individual items {f}, {g}, {d}, and {b} in the first round, sorted by TWU value. Therefore, the register configurations are as follows: slave_register_0, slave_register_1, slave_register_2, and slave_register_3 are equal to 0x00005000, 0x00006000, 0x00003000, and 0x00001000, respectively, based on the storage locations of the individual items in DDR. Slave_register_4 remains equal to 0x00000014 throughout the calculation process, based on the user-set threshold. Because the IP core currently being called is for the individual utility calculation, slave_register_5 and slave_register_6 are both equal to 0. Similarly, because the IP core currently being called is for the individual utility calculation and this is not the last round of calculation, slave_register_7 is equal to 0x11110000. When the configuration is complete and calculation is enabled, pull high the read_txn signal of the four-way IP core. At this time, slave_register_7 is equal to 0x11111111, and the IP core starts calculation.
[0066] In the IP core for single-item high-availability calculation, its input and output signals are shown in Table 8 below.
[0067] Table 8. IP core input and output signals for single high-availability calculation
[0068] To clearly and concisely illustrate the calculation process, the example raw database used contains only five transactions. However, in the actual calculation process, the IP core still continuously reads 1024 transactions through the AXI-HP channel. When the axi_rready and axi_rready signals are simultaneously valid, the current clock data is valid. To more clearly illustrate the calculation process, for example, when the input signals axi_rready and axi_rready are simultaneously valid for 1024 clock cycles, data_in is 0x00010000 in the first 256 clock cycles, 0x00020000 in the second 256 clock cycles, 0x00010000 in the third 256 clock cycles, and 0x00020000 in the fourth 256 clock cycles. During these 1024 cycles, the IP core slices the lower 16 bits of data_in and stores them in registers. The first burst reads 256 1s (the utility value) and accumulates them into the 32-bit sum register (sum). The second burst reads 256 2s. The third burst also reads 1, and the fourth burst reads 2. After 1024 clock cycles, sum = 0x600 and is cleared to 0 on the 1025th clock, ready for the next calculation. During data transfer, the comparison and accumulation operations continue, and sum is constantly compared with MUTIL. If it is greater than or equal to MUTIL, the result (result) is asserted and held for any 1024 clock cycles. Finally, at the 1025th clock, the calculation is complete, signaling (done). The calculation result is presented as a single bit per cycle. At this point, the high-utility item in the result is marked as 1, while the low-utility item is marked as 0.
[0069] According to the above process, the high utilities of the four parallel computation items {f}, {g}, {d}, and {b} are all 0, except for {b}, which is 1, and are stored in the corresponding to_bram_reg registers. In the second round of computation, since only items {a}, {e}, and {c} remain, only the first three IP cores are invoked, and this is the final round. Because the fourth computation IP core still needs to output calculation results, the computation is configured so that slave_register_0, slave_register_1, and slave_register_2 are equal to 0x00000000, 0x00004000, and 0x00002000, respectively, based on the DDR storage locations of the respective items. The slave_register_3 configuration is ignored; therefore, slave_register_7 is equal to 0x33330000, and when the computation is started, slave_register_7 is equal to 0x33331111.
[0070] According to the original database, the high utility of each item {a}, {e}, and {c} in this round of parallel computing is 0. All computing is completed, and the computing data is read.
[0071] See Figure 8 The storage process of the results in the single-item high-utility calculation process shown in the figure, in order to optimize storage and transmission efficiency, after each single-item calculation cycle, the comparison result is stored as the lowest bit in the result storage register (to_bram_reg), and the register is serially shifted left by the single-bit result in the clock cycle. After 32 calculation cycles, the register stores 32 bits of data containing 32 single-item judgment results. Figure 8 The right side shows the data format in to_bram_reg after 32 complete calculation cycles. After the module completes the utility calculation and judgment of 32 individual items, it outputs 32 bits of data containing the 32 calculation results to the first BRAM, the calculation result BRAM. During this process, the BRAM address signal increments after each storage to prevent data overwriting.
[0072] If the calculation reaches the last item, the current to_bram_reg stored in the item judgment result may not meet the 32 requirements. When the IP core detects the LAST signal going high (the end-of-calculation flag, indicating that the current item is the last one calculated by the IP core) and the calculation is complete, it transfers the currently stored calculation data to the first BRAM and sends a 1 to the second BRAM (the calculation result valid BRAM) (indicating that the current calculation is complete and all previously transmitted judgment results are valid). For example, in the example database in Table 1, when the IP core calculates only two or one item, except for IP core #4, whose to_bram_reg = 0x00000001 indicates that only the last item calculated by the IP core is valid, all other IP cores have to_bram_reg = 0x00000000. Although IP core #4 does not perform any actual calculations during the final round of calculation, if LAST = 1, it will transfer 0x00000001 to the first BRAM and 1 to the second BRAM. At this time, the other three IP cores will also write to_bram_reg into the first BRAM after the last round of calculation and update the validity of the second BRAM.
[0073] After calculating the high utility of all individual items within the IP core, the ARM side reads the results from each BRAM channel via the AXI-GP interface and verifies their validity. When analyzing and recording the high utility of each individual item, the data read from the BRAM of IP core 1 is first called. Since two parallel calculations are performed, the data is left-shifted 30 bits, and then left-shifted again to extract the highest bit. If this bit is 0, {f}, which ranks first in the TWU value ranking, is not high utility. Next, the BRAM data read from IP cores 2, 3, and 4 is the same as that of IP core 1, resulting in {g}, {d}, and {b}, which are low utility, low utility, and high utility, respectively. Next, IP core 1 is returned to the first channel and left-shifted again to extract the highest bit, resulting in {a}, which is low utility. Similarly, the BRAM data read from IP cores 2 and 3 is read again, resulting in {e} and {c}, which are both low utility.
[0074] In summary, the lower the address of the calculation result in each BRAM, the higher the number of bits, and the higher the order of the corresponding calculation item. Furthermore, only when the last set of to_bram_reg data does not hold 32 bits is it necessary to first shift left and discard invalid bits.
[0075] Next, according to the ascending order of TWU value and the "depth first" strategy, the itemsets will be sorted in the order Figure 9The order shown is used for expansion. This process calls the IP core on the FPGA side to perform parallel computing. The present invention proposes a task allocation strategy that combines static and dynamic operations, including two-layer static and dynamic allocation methods. First, static task division is adopted to divide the expanded item set of the first single item and all expanded item sets of other single items into two groups, and the four IP cores for calculation are divided into two groups, each responsible for the calculation of the two groups of expanded item sets. Between the two IP cores in each group, a dynamic task allocation method is adopted, and tasks are not fixedly allocated. When an IP core completes the branch judgment of the current item set, it will take the next task branch in the order within the group until all iterations are completed.
[0076] See Figure 10 The parallel computing task allocation framework shown in the figure assumes that the database has only five items, and their TWU values are sorted in the order {a, b, c, d, e}. A static task partitioning strategy is first used to assign the expanded itemset of {a} to the first IP core group for computation, while the expanded itemsets of {b}, {c}, {d}, and {e} are assigned to the second IP core group for computation. However, the uncertainty of the pruning strategy can lead to uneven workloads across different computation tasks, and the results of the itemset computation determine how the subsequent itemsets are expanded. Therefore, a dynamic allocation strategy is adopted within the group. If IP core number 1 in the first IP core group completes the computation of {ab} and its expanded itemset, and IP core number 2 has not yet completed the computation of {ac} and its expanded itemset, IP core number 1 will be assigned to the expanded computation of {ad} and its expanded itemset. If IP core number 3 in the second IP core group completes the computation of {b} and its expanded itemset, and IP core number 4 has not yet completed the computation of {c}, IP core number 3 will be assigned to the computation of {d}, and so on.
[0077] The startup settings for each IP core in each IP core group must include setting the required individual items' storage addresses in DDR via registers slave_register_0-slave_register_3. For example, if the currently expanded item set is {Pxy}, then according to the prefix-free expansion scheme {Pxy} = {Px} + {y}, the utility row for {y} must be read from DDR; slave_register_4 maintains the previously set threshold. slave_register_5 and slave_register_6 must be configured based on the iteration item set of each IP core. If the currently expanded item set {Pxy} contains m individual items, its connected item set should include {Px} and {y}. In the present invention, the utility value of {y} is obtained from DDR, while the utility value of {Px} is obtained from the first BRAM, the iteration item set utility storage BRAM. Assuming that an expansion based on {f} requires {f} to be stored in the first BRAM first, this ensures consistency in subsequent operations. Therefore, the storage structure in the first BRAM can be regarded as a regular design as shown in Table 9 below, that is, the zeroth group of utility values must correspond to a single item, the first group corresponds to an item set containing two single items, and so on.
[0078] Table 9. Storage range of each set
[0079] This paper proposes a stack optimization search strategy that uses only 1M to record the storage order of individual items in DDR cells. It dynamically obtains the required {Px} location for the currently expanded item set {Pxy} and the storage location of {Pxy} in the first BRAM when there is potential for expansion, namely the itr and loc signal values in Table 10 below. These are defined as follows: loc = sizeof({Px}) -1= (m - 2); itr = sizeof({Pxy}) -1= (loc + 1) = (m - 1).
[0080] According to Table 9, when expanding the itemset {fgdbaec}, loc = 5 and itr = 6; when expanding the itemset {fgdb}, loc = 2 and itr = 3. In summary, slave_register_5 and slave_register_6 are dynamically configured based on the iterative itemset for each calculation, ensuring the correct storage location of read utility data and {Pxy}. The IP core does not require a tail data signal, and when calling this IP core, func_switch = 0, so slave_register_7 equals 0x00000000. When configuration is complete and calculation is enabled, pull the read_txn signal of the four IP cores high. At this point, slave_register_7 equals 0x00001111, and the IP core begins calculation.
[0081] In the IP core for item set high utility calculation, its input and output signals are shown in Table 10 below.
[0082] Table 10. IP core input and output signals for itemset high-availability calculation
[0083] Since the IP core needs to perform different calculations or storage operations based on the input and intermediate calculation results, a state machine is introduced as a control module. The design state of the present invention includes four states, and their functions are: 1. IDLE state: This is the idle state. In this state, no calculation or storage operations are performed, and the system waits for triggering. 2. R_S (Read and Save) state: In this state, the module actively reads the 1024 utility value combinations of the specified single item in the sliding window at this time from the DDR and stores the data in the specified location of the first BRAM; 3. R2_C (Read two storage and Calculate) state: In this state, the module reads 1024 utility combinations of the specified items (itemsets) from the first BRAM and DDR simultaneously, and expands a new itemset based on these two items (itemsets). It connects the utility matrix of the new itemset, determines the high utility of the current expanded itemset and the necessity of expansion based on the new utility matrix, and updates the calculation results to the calculation result BRAM, that is, the second BRAM; 4. S (Save) state: Store the utility matrix of the new item set to the first BRAM.
[0084] See Figure 11 The state transition relationship and conditions shown intuitively demonstrate the triggering conditions and state transition relationship of the state machine control process in the core of the IP core during the high-utility calculation of the item set.
[0085] First, when the IP core is not called or the trigger condition is not met, it remains in the IDLE state. In this state, the IP core does not perform any calculation or storage operations.
[0086] Initially, the IP core is in the IDLE state. If both axi_rready and axi_rvalid are asserted simultaneously when reading utility values, the IDLE state transitions. At this point, if itr = 0, indicating a write command to the BRAM storing the iterated item set's utility, write position 0. This means the utility matrix row for the first single item must be stored in BRAM group 0 to prepare for subsequent item set expansion. Therefore, the IP core transitions to the R_S state and, during the AXI protocol clock cycle that reads valid data, stores this data in BRAM group 0. If both axi_rready and axi_rvalid become invalid, and the IP core's built-in 1024-bit counter reaches zero, indicating that all 1024 32-bit data items have been stored in BRAM, the state transitions back to IDLE. For example, if data_in = 0x00010001 is input during 1024 consecutive clock cycles of axi_rready and axi_rvalid signals, the IP core in the R_S state simultaneously pulls high both the we and en signals of the iterative item set utility storage BRAM for 1024 clock cycles with axi_rready and axi_rvalid signals pulled high. For the BRAM address signal, addr = itr * 0x1000 on the first clock cycle and increments by 0x4 on each clock cycle. The BRAM dout signal is directly connected to data_in, storing all utility data for 1024 clock cycles. That is, dout = data_in = 0x00010001 for 1024 clock cycles, and state = IDLE after the storage is complete.
[0087] In the IDLE state, when utility values are read, axi_rready and axi_rvalid are both valid, and this is not the first read (itr!=0). This means that when the next store operation is executed, the write position of the iterative itemset utility storage BRAM is not 0. According to the proposed depth-first prefix-free itemset expansion, all utility matrix rows stored thereafter should correspond to itemsets rather than individual items. This means that the module needs to expand and evaluate the previous items (itemsets) before storing them. Under this condition, the module transitions to the R2_C state.
[0088] During the 1024 clock cycles when both axi_rready and axi_rvalid are valid, the module reads the utility row of item set 0 in the BRAM according to the loc signal and concatenates it with the single-item utility row transmitted by the AXI channel to form a new utility matrix. Assume that the first 768 values of the data_in and din signals are 0x00010001, and the last 256 are 0x00100001. This means that the first 768 utility values of both items (itemsets) are 0x00010001, and the last 256 are 0x00100001. The IP core separates the utility value and residual utility values of these two pairs of 1024 values and accumulates them into registers.
[0089] The calculation of the high efficiency of an item set is the same as the calculation of the high utility of a single item, but the result output method is different. Since each set of calculation results in the item set calculation is related to the next set of expansion methods, an output is required at the end of each calculation. The present invention designs a 3-bit data. The highest bit indicates whether the item set has expansion potential, and the middle bit indicates whether the item set has high utility. Both are represented by 1 for yes and 0 for no. The lowest bit indicates the validity of the current output result, and is always 1 when the result is output. The address of the calculation result BRAM remains unchanged each time it is written, and the address data is cleared to zero after each upper-level program read operation is completed. Such a design can ensure the correctness of the calculation result data transmitted each time. According to the utility value and MUTIL size in the current example, the re_en signal is pulled high after the calculation is completed and the result 0x7 is output.
[0090] If the itemset has expansion potential, the newly connected utility rows must be stored in the iterative itemset utility storage BRAM, which causes the IP core to transition from the R2_C state to the S state. Based on the example above, the IP core sequentially stores the matrix rows waiting in the register into the itr=1th group of BRAM locations for easy access during the next expansion. During storage, a 32-bit data structure consisting of the utility value and the remaining utility value for the transaction is written per clock cycle, depending on the BRAM bandwidth. This means that the value is 0x00020001 for the first 768 clocks and 0x00200001 for the last 256 clocks. After these 1024 cycles, the IP core returns to the IDLE state, awaiting the next computation or storage operation.
[0091] Furthermore, the beneficial effects of the method of the present invention are illustrated through specific simulation experiments.
[0092] 1. Simulation conditions Based on the Xilinx Zynq-7000 board (model: XC7Z100-2FFG900I), the ARM frequency is 666.67 MHz, the DDR frequency is 533.33 MHz, and the FPGA clock frequency is 200 MHz. The database uses the open source database SPMF.
[0093] The methods compared in the experiment are as follows: One is the first single-stage high-utility itemset mining algorithm based on utility list, which is referred to as algorithm 1 in the experiment, and the reference is Liu, Mengchi and Qu, Junfeng, “Mining high utility itemsets without candidate generation,” in Proceedings of the 21st ACM international conference on Information and knowledge management, 2012, pp. 55–64.
[0094] The other is a high-utility itemset mining method based on utility matrix and prefix-free expansion method, which is optimized based on algorithm 1, referred to as algorithm 2 in the experiment, and the reference is Li Gufeng, Fang Weiyi, Wang Jialong, Xiang Jiawei, Shang Tao, “High-utility itemset mining method and device”, patent application number: 202411207964.3, publication number: CN119179727A.
[0095] Simulation content The above algorithms have the following related settings in the experimental process: platform selection, running mode, etc. Table 11, wherein algorithm 3 is the method of the present application.
[0096] Table 11. Settings of comparison algorithms
[0097] The comparison test is based on the SPMF open source database. Since the number of transactions in the original database is indefinite, in order to compare the performance of each algorithm when facing a data stream structure with a window size of 1024, 1024 transactions are randomly extracted from each database as experimental objects for comparison. The comparison of the characteristics of each database is shown in Table 12.
[0098] Table 12. Parameters of each database
[0099] According to the specific embodiments of the present application, the running time of the three algorithms on each database is compared, and the results are as follows: Figure 12 、 Figure 13 and Figure 14As shown. Wherein, the acceleration ratio = (heterogeneous parallel architecture running time / ARM architecture running time). The degree of acceleration of the method of the present invention varies under databases with different sparsity. In the sparser test1 and test2 databases, the acceleration effect is only about 2 times. In contrast, in the densest Connect and Chess databases, the acceleration effect can reach 5-20 times compared with Algorithm 3, and the heterogeneous acceleration ratio exceeds 3 under some thresholds, showing extremely high computing efficiency. For the T20I20D10K and Mushroom databases with medium sparsity, although the limitation of database scanning time is broken, due to the influence of task configuration time during parallel computing, the acceleration ratio exceeds 2 in some cases, and still hovers below 2 in some cases.
[0100] The acceleration ratio reflects the window response speed of high-utility item set mining in the data stream environment. It can be seen that the present invention not only improves the real-time performance of the system, but also effectively reduces the storage pressure of the sliding window on past data and optimizes resource utilization efficiency.
[0101] According to the specific embodiment of the present invention, the number of iterations of the three algorithms on each database is compared. The number of iterations of the sequential execution algorithm refers to the number of item set expansions, while the number of iterations of the parallel execution algorithm refers to the number of times from the first parallel call of the four IP cores to the last group of IP cores being called. The results are as follows: Figure 15 、 Figure 16 and Figure 17 As shown in the figure, the iteration compression ratio = (number of iterations before parallelization / number of iterations after parallelization).
[0102] The number of iterations after parallel computing is significantly reduced. Using the iteration compression ratio as a reference, the algorithm's iteration compression ratio remains stable at around 3 across various databases, even exceeding 3 in the test2 database. Based on the characteristics of four-way parallel computing, the iteration compression ratio should ideally reach 4. This demonstrates the excellent effectiveness of the present invention.
[0103] The contribution of each module in the architecture to the overall resources of the heterogeneous accelerator is calculated, with a focus on the usage of lookup table resources and register resources. Figure 18 From the overall architecture perspective, the lookup table and register resources of the key parallel computing modules account for 87.91% and 95.78% of the total resources respectively. Figure 18Among key computing modules, the func_k module occupies a particularly high proportion of resources within parallel_computing, with its lookup table and register resources accounting for 97.82% and 99.19%, respectively. This indicates that the itemset matrix connection and utility judgment IP core corresponding to func_k is a computationally intensive core component, dominating overall resource consumption.
[0104] Table 13. Figure 18 Medium module function
[0105] In summary, the method presented here achieves significant acceleration gains, doubling or more the algorithm's runtime and achieving an iterative compression ratio of 75% of the theoretical maximum. Furthermore, the IP core responsible for calculating high-utility itemsets in parallel computing, as the module with the highest resource consumption and the highest runtime share, should be the core focus of overall algorithm optimization. Therefore, it is foreseeable that this architecture will achieve even better acceleration performance with further development and optimization of onboard resources.
[0106] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A high-utility pattern mining heterogeneous parallel acceleration method for data streams, characterized by: include: Step 1: Scan the original database, which includes multiple transactions, obtain the utility of each item in each transaction and construct a calculation auxiliary structure; Step 2: Calculate the residual utility of the single item in each transaction according to the calculation auxiliary structure, and store the utility and residual utility of the single item in each transaction in the DDR; Step 3: Prune the constructed TWU value hash table according to the TWU pruning strategy to obtain a pruned TWU value hash table; Step 4: Using four parallel IP cores, based on the utility and residual utility of the items stored in the DDR in each transaction, a polling method is used to calculate the high utility of each item in the pruned TWU value hash table, and the high-utility item is determined based on the calculation result of the high utility of the item. Step 5: Divide the four parallel IP cores into a parallel first IP core group and a second IP core group, each IP core group includes two parallel IP cores, and adopt a dynamic and static combined task allocation strategy to allocate all extended item sets in the pruned TWU value hash table to the first IP core group and the second IP core group. Calculate the high utility of the item set based on the utility and residual utility of the single item stored in the DDR in each transaction, and determine the high-utility item set based on the calculation result of the high utility of the item set.
2. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 1 is characterized in that: The computational auxiliary structure includes a transaction utility vector, the TWU value hash table and a position vector structure; wherein, The transaction utility vector is used to record the total transaction utility of each transaction in the original database; The TWU value hash table is used to record the TWU value of each single item in the original database, and the TWU value hash table is sorted in ascending order according to the TWU value of the single item; The position vector structure is used to record the base address in the DDR corresponding to each single item divided according to the order in which each single item appears in the original database, wherein 1024 consecutive addresses in the DDR are regarded as an address range corresponding to a single item.
3. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 2 is characterized in that: The TWU value of the single item is the accumulation of the total transaction utility of all transactions in the original database including the current single item.
4. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 2 is characterized in that: The step 2 includes: Step 2.1: For each transaction, based on the total transaction utility of the current transaction and the utility of the individual item in the current transaction, the residual utility of the individual item in the current transaction is calculated according to the sorting order of the individual items in the TWU value hash table; Step 2.2: For the address range in the DDR corresponding to each single item, one address in the address range corresponds to storing the utility and residual utility of the current single item in a transaction, wherein for a transaction that does not include the current single item, the corresponding address remains 0, and the 32 bits of each address are divided into upper 16 bits and lower 16 bits, the upper 16 bits are used to store the utility of the current single item in the current transaction, and the lower 16 bits are used to store the residual utility of the current single item in the current transaction.
5. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 2, characterized in that: The step 3 comprises: The TWU value of each single item in the TWU value hash table is compared with a preset threshold, and the single items in the TWU value hash table that are less than the threshold are removed to obtain the pruned TWU value hash table.
6. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 1, characterized in that: Each IP core is correspondingly provided with a first BRAM and a second BRAM. Accordingly, step 4 includes: Step 4.1: Calling four parallel IP cores to calculate the high-item utilities of four consecutive items in the pruned TWU value hash table in parallel using a round-robin approach. During each round of high-item utilities calculation, the IP cores are configured using the ARM side. If the number of items in the last calculation is less than four, the required number of IP cores is called. Step 4.2: After each round of calculation, the calculation result of the single high utility is stored in the result register corresponding to the IP core in a left-shift manner. When the result register is full, the data therein is written to the first BRAM. When the IP core receives the end calculation flag, after the last round of calculation, the data in the result register is written to the first BRAM, and the validity flag is stored in the second BRAM. For the IP cores that are not called in the last round of calculation, the data in the corresponding result register is written to the first BRAM, and the validity flag is stored in the second BRAM. Step 4.3: The ARM side determines whether the calculation of all items in the pruned TWU value hash table is completed based on the pruned TWU value hash table pointer. If the parallel calculation of the current round is completed but there are still items in the pruned TWU value hash table that have not been calculated, return to step 4.1; if the parallel calculation of the current round is completed and all items in the pruned TWU value hash table are calculated, execute step 4.4; Step 4.4: The ARM side reads the contents of the first BRAM and the second BRAM according to the validity flag, analyzes the calculation result of the single item high utility, and determines and records the single item with high utility.
7. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 6, characterized in that: In step 4.1, the configuration content of the IP core configured by the ARM side during the calculation process of each round of single high utility includes: The base address of the item read by the IP core from the DDR, the threshold preset by the user, the IP core call enable, and the end calculation flag if the current item is the last round of calculation of the IP core.
8. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 6, characterized in that: The step 5 comprises: Step 5.1: Call the four IP cores in parallel and use static task partitioning to assign the expanded item set of the first single item in the pruned TWU value hash table to the first IP core group to calculate the item set high utility, and assign the expanded item set of the remaining single items in the pruned TWU value hash table excluding the first single item to the second IP core group to calculate the item set high utility. Within the first IP core group and the second IP core group, a dynamic task allocation method is used. When any IP core completes the calculation of the item set under the current branch, it receives the calculation task of the item set of the next branch according to the order of the item sets within the IP core group until all item sets within the IP core group are iterated. The IP core is configured using the ARM side during each item set high utility calculation process. Step 5.2: After each calculation is completed, the calculation result of the high utility of the itemset is stored in the second BRAM. Based on the calculation result of the high utility of the itemset, the utility and residual utility of the itemset with high utility and expansion potential are stored in the first BRAM. Step 5.3: After each calculation is completed, the ARM side is used to read the contents of the second BRAM, and the calculation results of the high utility of the item set are analyzed to determine and record the high utility item set. If all single items in the pruned TWU value hash table are expanded and calculated, the algorithm ends. Otherwise, the expansion operation is performed on the item set with expansion potential and then return to step 5.
1.
9. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 8, characterized in that: In step 5.1, the configuration content of configuring the IP core using the ARM side during each item set high utility calculation process includes: The base address of the single item read by the IP core from the DDR, the threshold set by the user, the IP core call enable, and the loc signal and itr signal of the auxiliary calculation item set utility.
10. The high-utility pattern mining heterogeneous parallel acceleration method for data streams according to claim 7 or 9, characterized in that: After the IP core is configured using the ARM side, when the IP core receives an instruction to read data from the DDR, it performs the first burst according to the configured base address of the single item read from the DDR, and automatically enables a new burst read and updates the base address after the burst transmission is completed, until the data in the address range corresponding to a single item in the DDR is read four times in a burst.
Citation Information
Patent Citations
Efficient item set mining method and device
CN119179727A