A High Fuzzy Utility Itemset Mining Method Based on Fuzzy List Buffer

By introducing fuzzy list buffers and evaluating the co-occurrence structure of fuzzy utility items into the high fuzzy utility item set mining algorithm, the problem of existing algorithms spending too much time and memory in the fuzzy list connection operation is solved, and a more efficient mining process is achieved.

CN115470262BActive Publication Date: 2025-06-20GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211048967.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-06-20
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The existing high fuzzy utility item set mining algorithm consumes too much running time and memory in the connection operation of fuzzy lists.

Method used

The high fuzzy utility item set mining method based on fuzzy list buffer is adopted. By evaluating the fuzzy utility co-occurrence structure and fuzzy list buffer construction program, the recursive search subroutine is optimized, the inspection of invalid fuzzy item sets is reduced, and memory consumption is reduced through the memory multiplexing mechanism.

Benefits of technology

It effectively reduces the running time and memory consumption of the high fuzzy utility item set mining algorithm and improves mining efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470262B_ABST
    Figure CN115470262B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and computer-readable storage medium for mining high fuzzy utility item sets based on a fuzzy list buffer. The method includes: S1: Initializing the operation parameters of data mining; S2: Scanning the transaction database D and calculating the fuzzy utility upper bound FUUB of single items according to the membership function R, and creating an initial list I*; S3: Storing the single fuzzy items whose fuzzy utility upper bound values are not less than the minimum threshold minUtil into the initial list I*, and sorting them in ascending order of the fuzzy utility upper bound values; S4: Scanning the database D again, constructing an evaluation fuzzy utility co-occurrence structure EFuCS, a fuzzy list buffer FLBuf and its auxiliary summary list SL; S5: Invoking the recursive search subroutine Search and passing in parameters; S6: Outputting all high fuzzy utility item sets HFUIs whose fuzzy utilities are not lower than the minimum threshold, and completing data mining. The present invention reduces the running time of the high fuzzy utility item set mining algorithm and reduces the memory consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-fuzzy utility item set mining, and more specifically, to a high-fuzzy utility item set mining, system and computer-readable storage medium based on a fuzzy list buffer. Background Art

[0002] With the increasing progress of science and technology, the hot fields involved in the data collection samples are becoming wider and wider (such as economy, military, logistics, finance, telecommunications, etc.). In general, the data in reality is complex or mixed in structure, structured or unstructured, incomplete, and the features are inaccurate. For these fuzzy and complex data sets, without extracting their important typical features, it is impossible to complete the statistical mining analysis of the data. The current research dilemma shows that although many researchers have achieved fruitful results in deterministic data mining techniques, proposed many effective implementation algorithms, and met various different actual rapid applications, in fact, the research on fuzzy and complex data mining techniques is still in an immature important stage, and there are still a large number of specific problems to be solved objectively. "Very fuzzy" is an important structural feature for humans and animals to perceive all things, acquire various knowledge, cognitive reasoning, and decision-making implementation. "Completely fuzzy" has a higher information storage capacity than "clear", richer philosophical connotations, and is more in line with the objective world. In human thinking, there are many new fuzzy concepts, such as the size of the diameter, hot and cold, etc. These industry concepts have no clear philosophical connotations and extensions, so it is impossible to find and apply traditional precise mathematical descriptions.

[0003] Traditional association rule mining (ARM) and frequent pattern mining (FIM) algorithms may output frequent but low-profit results, which are unacceptable in some cases. Therefore, to solve this problem, some scholars have proposed a new mining framework based on utility theory, called high-utility item set mining (HUIM). Shen et al. first attempted to apply utility constraints in association rule mining. They pointed out that utility includes the quantity of items and the profit of a unit item, and it is neither monotonic nor anti-monotonic. In addition, since utility can be positive or negative, the optimization methods applied in FIM cannot be directly applied to HUIM. Subsequently, some scholars proposed a transaction weighted utilization rate model to solve the above problems. Although the HUIM algorithm evaluates the importance of different items through the measurement of numerical utility, the mined results cannot provide other more useful information, such as the purchase quantity of items, etc.

[0004] Therefore, in the prior art, Wang et al. combined utility theory and fuzzy theory to propose a new architecture called Fuzzy Utility Mining (FUM) to mine high fuzzy utility item sets (HFUIs) from quantitative transaction databases. In addition, Lan et al. adopted a user-defined membership function to evaluate the fuzzy utility of item sets. The highlight of their work is the implementation of the downward closure property in FUM. In FUM, this is an effective fuzzy utility upper bound (FUUB). Recently, Wan (Wan, Shicheng, et al. "FUIM: Fuzzy Utility Itemset Mining." arXiv preprint arXiv:2111.00307 (2021).) et al. proposed an FUM algorithm called FUIM. They proposed the concept of residual fuzzy utility and a method for mining high fuzzy utility using fuzzy lists. A large number of experiments show that FUIM performs better than previous algorithms. However, due to the structure of the fuzzy list, FUIM consumes too much running time and memory in the join operation of the fuzzy list. Therefore, there is an urgent need to propose a method for mining high fuzzy utility item sets based on a fuzzy list buffer. Summary of the Invention

[0005] To overcome the deficiencies in the prior art, such as large memory consumption and long running time in mining high fuzzy utility item sets, the present invention provides a method, a system, and a computer-readable storage medium for mining high fuzzy utility item sets based on a fuzzy list buffer, which reduce the running time and memory consumption of the high fuzzy utility item set mining algorithm.

[0006] The primary objective of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:

[0007] The first aspect of the present invention provides a method for mining high fuzzy utility item sets based on a fuzzy list buffer, including the following steps:

[0008] S1: Initialize the data mining operation parameters, where the data mining operation parameters include: the quantitative database D to be mined, a predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set;

[0009] S2: Scan the transaction database D and calculate the fuzzy utility upper bound FUUB of single items according to the membership function R, and create an initial list I*;

[0010] S3: Deposit single fuzzy items whose fuzzy utility upper bound values are not less than the minimum threshold minUtil into the initial list I*, and sort them in ascending order according to the fuzzy utility upper bound values;

[0011] S4: Rescan the database D again to construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf, and its auxiliary summary list SL;

[0012] S5: Call the recursive search subroutine Search and pass in parameters, where the parameters include: the initial prefix fuzzy item set Initialize the list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf, and its summary list SL;

[0013] S6: Output all high fuzzy utility item sets HFUIs with fuzzy utility not lower than the minimum threshold to complete data mining.

[0014] Furthermore, calling the recursive search subroutine Search in step S5 includes the following steps:

[0015] S501: In the recursive search subroutine Search, for an extended fuzzy item set X of the fuzzy item set P, if the sum of the fuzzy utilities of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of high fuzzy utility item sets;

[0016] S502: If the sum of the fuzzy utilities sumFu in the summary list SL(X) of the fuzzy item set X and the sum of the remaining fuzzy utilities sumRfu add up to not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high fuzzy utility item set;

[0017] S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y is after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound of the fuzzy utilities of the fuzzy item sets X and Y in the evaluation fuzzy utility co-occurrence structure EFuCS is not less than the minimum threshold minUtil;

[0018] S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, the summary list SL, the fuzzy item sets P, X, Y, and the minimum threshold minUtil as parameters, and return the construction result;

[0019] S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X;

[0020] S506: Combine the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed.

[0021] Further, the fuzzy list buffer construction program described in step S504 includes the following steps:

[0022] S5041: In the fuzzy list buffer construction program, set pointers PPnt, PxPnt, and PyPnt as the starting positions of SL(P), SL(Px), and SL(Py) in the summary list respectively, and the pointers point to the tuples in the fuzzy list buffer;

[0023] S5042: Let the variable EAMeasure be the sum of the fuzzy utilities of the summary lists SL(Px) and SL(Py) of the fuzzy item sets Px and Py plus the sum of the remaining fuzzy utilities. Let the variable insertPos be the starting position of the last fuzzy item set in the summary list SL;

[0024] S5043: If the Tids in the tuple pointed to by the pointer PxPnt are less than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PxPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by PxPnt from the variable EAMeasure;

[0025] S5044: If the Tids in the tuple pointed to by the pointer PxPnt are greater than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PyPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by PyPnt from the variable EAMeasure;

[0026] S5045: If the Tids in the tuple pointed to by the pointer PxPnt are equal to the Tids in the tuple pointed to by the pointer PyPnt, and the summary list SL(P) is not empty, then continuously move the pointer of PPnt to the right until PPnt moves to the end of SL(P) or the Tids in the tuple pointed to by PPnt are equal to the Tids in the tuple pointed to by PxPnt;

[0027] S5046: If the insert position insertPos exceeds the size of the fuzzy list buffer, then allocate new memory space, otherwise recycle and reuse the memory space, add a new tuple to the fuzzy list buffer, set Tids as the Tids of PxPnt, fus as the sum of fus of PxPnt and fus of PyPnt minus the fus of PPnt, and rfus as the rfus of PyPnt;

[0028] S5047: After inserting the data, move the pointers PxPnt and PyPnt one position to the right at the same time;

[0029] S5048: When the pointer PxPnt does not point to the end position EndPos of the summary list SL(Px), and the pointer PyPnt does not point to the end position EndPos of the summary list SL(Py), repeatedly execute the fuzzy list buffer program;

[0030] S5049: If the variable EAMeasure is less than the minimum threshold minUtil, return the result false;

[0031] S50410: Update the summary list SL(Pxy), return the result true, and end the fuzzy list buffer construction program.

[0032] Furthermore, the fuzzy list buffer FLBuf is composed of triples (Tids, fus, rfus), where Tid is the transaction identifier in the database, fu is the fuzzy utility of the transaction, and rfu is the remaining fuzzy utility of the transaction.

[0033] Furthermore, the summary list SL is composed of tuples (Itemsets, StartPoss, EndPoss, sumFus, sumRfus), where Itemset represents the fuzzy item set, StartPos and EndPos respectively represent the start and end positions of the corresponding fuzzy item set in the fuzzy list buffer FLBuf, sumFu represents the sum of the fuzzy utilities fus of the corresponding fuzzy item set in the fuzzy list buffer, and sumRfu represents the sum of the remaining fuzzy utilities rfus of the corresponding fuzzy item set in the fuzzy list buffer FLBuf.

[0034] Furthermore, after the recursive search subroutine Search has checked a node and all its descendant nodes, the program starts to backtrack. At this time, the nodes that have been checked are no longer used, and the memory space allocated in the fuzzy list buffer FLBuf to store this node will be recycled and reused. The data of the new potential fuzzy item set is directly overwritten and written into the recycled memory space, and at the same time, the information in the summary list SL is updated to achieve memory reuse and reduce the memory consumption of the program.

[0035] Furthermore, the evaluation of the fuzzy utility co-occurrence structure EFuCS is represented in matrix form, with the index being the fuzzy item set and the value representing the upper bound FUUB of the fuzzy utility after the combination of two fuzzy item sets.

[0036] The second aspect of the present invention provides a high fuzzy utility item set mining system based on a fuzzy list buffer. The system includes: a memory and a processor. The memory includes a high fuzzy utility item set mining method program based on a fuzzy list buffer. When the high fuzzy utility item set mining method program based on a fuzzy list buffer is executed by the processor, the following steps are implemented:

[0037] S1: Initialize the data mining operation parameters, where the data mining operation parameters include: the quantitative database D to be mined, the predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set;

[0038] S2: Scan the transaction database D and calculate the upper bound of the fuzzy utility FUUB of a single item according to the membership function R, and create the initialization list I*;

[0039] S3: Store the single fuzzy items whose fuzzy utility upper bound values are not less than the minimum threshold minUtil into the initialization list I*, and sort them in ascending order according to the fuzzy utility upper bound values;

[0040] S4: Scan the database D again to construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its auxiliary summary list SL;

[0041] S5: Call the recursive search subroutine Search and pass in the parameters, where the parameters include: the initial prefix fuzzy item set the initialization list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its summary list SL;

[0042] S6: Output all the high fuzzy utility item sets HFUIs whose fuzzy utilities are not less than the minimum threshold to complete the data mining.

[0043] Furthermore, calling the recursive search subroutine Search in step S5 includes the following steps:

[0044] S501: In the recursive search subroutine Search, for an extended fuzzy item set X of the fuzzy item set P, if the sum sumFu of the fuzzy utilities of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of the high fuzzy utility item sets;

[0045] S502: If the sum sumFu of the fuzzy utilities in the summary list SL(X) of the fuzzy item set X and the sum sumRfu of the remaining fuzzy utilities add up to a result not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high fuzzy utility item set;

[0046] S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y is after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound values of the fuzzy utilities of the fuzzy item sets X and Y in the evaluation fuzzy utility co-occurrence structure EFuCS are not less than the minimum threshold minUtil;

[0047] S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, summary list SL, fuzzy item sets P, X, Y, and minimum threshold minUtil as parameters, and return the construction result;

[0048] S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X;

[0049] S506: Merge the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed.

[0050] The third aspect of the present invention provides a computer-readable storage medium, which includes a program for mining high fuzzy utility item sets based on a fuzzy list buffer. When the program for mining high fuzzy utility item sets based on a fuzzy list buffer is executed by a processor, the steps of the method for mining high fuzzy utility item sets based on a fuzzy list buffer are implemented.

[0051] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0052] The present invention uses an evaluation fuzzy utility co-occurrence structure to store the upper bound value of the fuzzy utility of the fuzzy item set, and uses the corresponding evaluation fuzzy utility pruning strategy to avoid checking a large number of invalid fuzzy item sets during the mining process of high fuzzy utility item sets, thereby reducing the running time of the algorithm;

[0053] The fuzzy list buffer is used to uniformly allocate and recycle the memory space. For the fuzzy lists of the fuzzy item sets that are no longer used during the search process, the fuzzy list buffer reclaims the memory space allocated for storing the fuzzy list of the item set and reallocates it to other fuzzy item sets that may become high fuzzy utility item sets, reducing memory consumption through memory reuse;

[0054] The fuzzy list of the fuzzy item set is stored in the fuzzy list buffer, and the position information of the corresponding fuzzy item set is stored in the summary list. By reading the position information of the fuzzy item set in the summary list, the required fuzzy item set can be directly located in the fuzzy list buffer, optimizing the search process for the fuzzy list of the fuzzy item set and reducing the running time of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of a method for mining high fuzzy utility item sets based on a fuzzy list buffer of the present invention.

[0056] Figure 2 This is a schematic diagram of the Search process of the recursive search subroutine of the present invention.

[0057] Figure 3 This is a schematic diagram of the process of constructing the fuzzy list buffer program of the present invention.

[0058] Figure 4 This is a schematic diagram of the fuzzy list buffer FLBuf, summary list SL, and evaluation fuzzy utility co-occurrence structure EFuCS of the present invention.

[0059] Figure 5 This is a schematic diagram of the comparison of the running time (s) effects of the FLB-Miner method and the FUIM method in the embodiment of the present invention.

[0060] Figure 6 This is a schematic diagram of the comparison of the memory consumption (MB) effects of the FLB-Miner method and the FUIM method in the embodiment of the present invention. Detailed implementation manners

[0061] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0062] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0063] Embodiment 1

[0064] As Figure 1 shown, the first aspect of the present invention provides a method for mining high-fuzzy-utility item sets based on a fuzzy list buffer, including the following steps:

[0065] S1: Initialize the data mining operation parameters, where the data mining operation parameters include: the quantitative database D to be mined, the predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set;

[0066] S2: Scan the transaction database D and calculate the upper bound of the fuzzy utility FUUB of single items according to the membership function R, and create an initial list I*;

[0067] S3: Store the single fuzzy items whose fuzzy utility upper bound values are not less than the minimum threshold minUtil into the initial list I*, and sort them in ascending order according to the fuzzy utility upper bound values;

[0068] S4: Rescan the database D again to construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf, and its auxiliary summary list SL;

[0069] S5: Call the recursive search subroutine Search and pass in parameters, where the parameters include: the initial prefix fuzzy item set Initialize the list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf, and its summary list SL;

[0070] As Figure 2 shown, calling the recursive search subroutine Search in step S5 includes the following steps:

[0071] S501: In the recursive search subroutine Search, for an extended fuzzy item set X of the fuzzy item set P, if the sum of the fuzzy utilities sumFu of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of high fuzzy utility item sets;

[0072] S502: If the sum of the fuzzy utilities sumFu in the summary list SL(X) of the fuzzy item set X and the sum of the remaining fuzzy utilities sumRfu add up to not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high fuzzy utility item set;

[0073] S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y is after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound value of the fuzzy utilities of the fuzzy item sets X and Y in the evaluation fuzzy utility co-occurrence structure EFuCS is not less than the minimum threshold minUtil;

[0074] S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, the summary list SL, the fuzzy item sets P, X, Y, and the minimum threshold minUtil as parameters, and return the construction result;

[0075] As Figure 3 shown, the fuzzy list buffer construction program described in step S504 includes the following steps:

[0076] S5041: In the fuzzy list buffer construction program, set pointers PPnt, PxPnt, and PyPnt to the starting positions of SL(P), SL(Px), and SL(Py) in the summary list respectively, and the pointers point to the tuples in the fuzzy list buffer;

[0077] S5042: Let the variable EAMeasure be the sum of the fuzzy utilities of the summary lists SL(Px) and SL(Py) of the fuzzy item sets Px and Py plus the sum of the remaining fuzzy utilities. Let the variable insertPos be the starting position of the last fuzzy item set in the summary list SL;

[0078] S5043: If the Tids in the tuple pointed to by the pointer PxPnt is less than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PxPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by PxPnt from the variable EAMeasure;

[0079] S5044: If the Tids in the tuple pointed to by the pointer PxPnt is greater than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PyPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by PyPnt from the variable EAMeasure;

[0080] S5045: If the Tids in the tuple pointed to by the pointer PxPnt is equal to the Tids in the tuple pointed to by the pointer PyPnt, and the summary list SL(P) is not empty, then continuously move the pointer of PPnt to the right until PPnt moves to the end of SL(P) or the Tids in the tuple pointed to by PPnt is equal to the Tids in the tuple pointed to by PxPnt;

[0081] S5046: If the insertion position insertPos exceeds the size of the fuzzy list buffer, then allocate new memory space, otherwise recycle and reuse the memory space, add a new tuple to the fuzzy list buffer, set Tids to the Tids of PxPnt, fus to the sum of fus of PxPnt and fus of PyPnt minus fus of PPnt, and rfus to the rfus of PyPnt;

[0082] S5047: After inserting the data, move the pointers PxPnt and PyPnt one position to the right at the same time;

[0083] S5048: When the pointer PxPnt does not point to the end position EndPos of the summary list SL(Px), and the pointer PyPnt does not point to the end position EndPos of the summary list SL(Py), repeat the fuzzy list buffer program;

[0084] S5049: If the variable EAMeasure is less than the minimum threshold minUtil, return the result false;

[0085] S50410: Update the summary list SL(Pxy), return the result true, and end the fuzzy list buffer construction program.

[0086] S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X;

[0087] S506: Merge the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed.

[0088] S6: Output all high fuzzy utility item sets HFUIs whose fuzzy utilities are not lower than the minimum threshold, and complete the data mining.

[0089] It should be noted that the fuzzy list buffer FLBuf, summary list SL, and evaluation fuzzy utility co-occurrence structure EFuCS used in the present invention are as Figure 4 shown. The fuzzy list buffer FLBuf is composed of triples (Tids, fus, rfus). Tid is the transaction identifier in the database, fu is the fuzzy utility of the transaction, and rfu is the remaining fuzzy utility of the transaction.

[0090] The summary list SL is composed of tuples (Itemsets, StartPoss, EndPoss, sumFus, sumRfus). Among them, Itemset represents the fuzzy item set, StartPos and EndPos respectively represent the start and end positions of the corresponding fuzzy item set in the fuzzy list buffer FLBuf, sumFu represents the sum of the fuzzy utilities fus of the corresponding fuzzy item set in the fuzzy list buffer, and sumRfu represents the sum of the remaining fuzzy utilities rfus of the corresponding fuzzy item set in the fuzzy list buffer FLBuf.

[0091] It should be noted that after the recursive search subroutine Search checks a node and all its descendant nodes, the program starts to backtrack. At this time, the nodes that have been checked are no longer used, and the memory space allocated in the fuzzy list buffer FLBuf to store this node will be recycled and reused. The data of the new potential fuzzy item set is directly overwritten and written into the recycled memory space, and at the same time, the information in the summary list SL is updated to achieve memory reuse and reduce the memory consumption of the program.

[0092] The evaluation fuzzy utility co-occurrence structure EFuCS is represented in matrix form. The index is the fuzzy item set, and the value represents the upper bound FUUB of the fuzzy utility after the merger of two fuzzy item sets.

[0093] Embodiment 2

[0094] In the second aspect of the present invention, a high-fuzzy utility item set mining system based on a fuzzy list buffer is provided. The system includes: a memory and a processor. The memory includes a program for a high-fuzzy utility item set mining method based on a fuzzy list buffer. When the program for the high-fuzzy utility item set mining method based on a fuzzy list buffer is executed by the processor, the following steps are implemented:

[0095] S1: Initialize the data mining operation parameters. The data mining operation parameters include: the quantitative database D to be mined, the predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set.

[0096] S2: Scan the transaction database D and calculate the upper bound of the fuzzy utility FUUB of single items according to the membership function R, and create the initialization list I*.

[0097] S3: Store the single fuzzy items whose upper bound values of fuzzy utility are not less than the minimum threshold minUtil into the initialization list I*, and sort them in ascending order according to the upper bound values of fuzzy utility.

[0098] S4: Scan the database D again, construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its auxiliary summary list SL.

[0099] S5: Call the recursive search subroutine Search and pass in the parameters. The parameters include: the initial prefix fuzzy item set the initialization list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its summary list SL.

[0100] S6: Output all high-fuzzy utility item sets HFUIs whose fuzzy utilities are not lower than the minimum threshold, and complete the data mining.

[0101] Further, calling the recursive search subroutine Search in step S5 includes the following steps:

[0102] S501: In the recursive search subroutine Search, for an extended fuzzy item set X of a fuzzy item set P, if the sum of the fuzzy utilities of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of high-fuzzy utility item sets.

[0103] S502: If the sum of the fuzzy utilities sumFu in the summary list SL(X) of the fuzzy item set X and the sum of the remaining fuzzy utilities sumRfu add up to a result not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high-fuzzy utility item set.

[0104] S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y comes after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound value of the fuzzy utility of the fuzzy item sets X and Y in the evaluation of the fuzzy utility co-occurrence structure EFuCS is not less than the minimum threshold minUtil;

[0105] S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, the summary list SL, the fuzzy item sets P, X, Y, and the minimum threshold minUtil as parameters, and return the construction result;

[0106] S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X;

[0107] S506: Merge the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed.

[0108] The third aspect of the present invention provides a computer-readable storage medium, which includes a program for the mining method of high fuzzy utility item sets based on a fuzzy list buffer. When the program for the mining method of high fuzzy utility item sets based on a fuzzy list buffer is executed by a processor, the steps of the mining method of high fuzzy utility item sets based on a fuzzy list buffer are implemented.

[0109] Example 3

[0110] The data mining method FLB-Miner (Fuzzy-List Buffer for high fuzzy utility itemset Miner) in this example is tested with the FUIM algorithm on a large number of data sets. The data set types include sparse and dense, real data and synthetic data. By comparison, the data mining method in this example has the following advantages:

[0111] (1) The time for fuzzy item set mining is greatly shortened: By running the FLB-Miner algorithm and the FUIM algorithm on different data sets, the running results are as Figure 5As shown in the figure. On the sparse dataset Retail, when the minimum fuzzy utility rate δ is set to 0.0001, the FLB-Miner algorithm reduces the time by nearly three times compared to the FUIM algorithm. On the synthetic dataset T40I10D100K, when the minimum fuzzy utility rate δ varies from 0.005 to 0.01, the time span of the FUIM algorithm is close to 150s, while the time span of the FLB-Miner algorithm is only close to 30s. By comparing other datasets, it can also be concluded that when comparing the running time of the FLB-Miner algorithm with that of the FUIM algorithm, the FLB-Miner shows excellent performance, greatly shortening the running time. Moreover, as the minimum fuzzy utility rate decreases, the performance superiority of the FLB-Miner algorithm is better demonstrated.

[0112] (2) The memory consumption of fuzzy utility mining is greatly reduced: The maximum memory consumed by the FLB-Miner algorithm and the FUIM algorithm when running on different datasets is as Figure 6 shown. On the datasets Chess and Mushroom, the maximum memory consumed by the FLB-Miner algorithm is approximately one-third of that consumed by the FUIM algorithm. On all the tested datasets, the FLB-Miner algorithm can efficiently reduce the memory consumption compared to the FUIM algorithm. The main reasons are as follows: By evaluating the fuzzy utility co-occurrence structure and its pruning strategy, the FLB-Miner algorithm can avoid checking a large number of invalid fuzzy item sets, reduce the search space, generate fewer candidate fuzzy item sets, and thus reduce the memory usage of the fuzzy list buffer; Through the fuzzy list buffer and the summary list, as well as an effective memory reuse mechanism, the memory space is recycled and reused, making more full use of the memory space, thereby reducing the memory space consumption.

[0113] The mining method of the present invention introduces the evaluation of the fuzzy utility co-occurrence structure, optimizes the pruning process of the recursive search program, and reduces the program running time. And through the fuzzy list buffer structure and the associated summary list structure, the construction of the fuzzy list is managed, the fuzzy list in the fuzzy list buffer is retrieved and stored more efficiently, the allocated useless memory space is recycled and reused, the memory reuse rate is increased, the memory consumption of mining is reduced, and the efficiency of mining high-fuzzy utility item sets is improved.

[0114] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for mining high-fuzzy utility item sets based on a fuzzy list buffer, characterized in that, Including the following steps: S1: Initialize the data mining running parameters, where the data mining running parameters include: the quantitative database D to be mined, the predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set; S2: Scan the transaction database D and calculate the upper bound of the fuzzy utility FUUB of a single item according to the membership function R, and create an initialization list I*; S3: Store the single fuzzy items whose fuzzy utility upper bound values are not less than the minimum threshold minUtil into the initialization list I*, and sort them in ascending order according to the fuzzy utility upper bound values; S4: Scan the database D again to construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its auxiliary summary list SL; S5: Call the recursive search subroutine Search and pass in parameters, where the parameters include: the initial prefix fuzzy item set Initialize the list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf, and its summary list SL; S6: Output all high fuzzy utility item sets HFUIs whose fuzzy utilities are not lower than the minimum threshold to complete the data mining; Among them, the recursive search subroutine Search called in step S5 includes the following steps: S501: In the recursive search subroutine Search, for an extended fuzzy item set X of the fuzzy item set P, if the sum of the fuzzy utilities sumFu of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of high fuzzy utility item sets; S502: If the sum of the fuzzy utilities sumFu in the summary list SL(X) of the fuzzy item set X and the sum of the remaining fuzzy utilities sumRfu add up to not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high fuzzy utility item set; S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y is after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound values of the fuzzy utilities of the fuzzy item sets X and Y in the evaluation fuzzy utility co-occurrence structure EFuCS are not less than the minimum threshold minUtil; S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, the summary list SL, the fuzzy item sets P, X, Y, and the minimum threshold minUtil as parameters, and return the construction result; S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X; S506: Merge the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed; Among them, the fuzzy list buffer construction program described in step S504 includes the following steps: S5041: In the fuzzy list buffer construction program, set the pointers PPnt, PxPnt, and PyPnt to the starting positions of SL(P), SL(Px), and SL(Py) in the summary list respectively, and the pointers point to the tuples in the fuzzy list buffer; S5042: Let the variable EAMeasure be the addition result of the sum of the fuzzy utilities of the summary lists SL(Px) and SL(Py) of the fuzzy item sets Px and Py and the sum of the remaining fuzzy utilities. Let the variable insertPos be the starting position of the last fuzzy item set in the summary list SL; S5043: If the Tids in the tuple pointed to by the pointer PxPnt are less than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PxPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by the pointer PxPnt from the variable EAMeasure; S5044: If the Tids in the tuple pointed to by the pointer PxPnt are greater than the Tids in the tuple pointed to by the pointer PyPnt, then move the pointer PyPnt one position to the right, and subtract the sum of fus and rfus of the tuple pointed to by the pointer PyPnt from the variable EAMeasure; S5045: If the Tids in the tuple pointed to by the pointer PxPnt are equal to the Tids in the tuple pointed to by the pointer PyPnt, and the summary list SL(P) is not empty, then continuously move the pointer of PPnt to the right until PPnt moves to the end of SL(P) or the Tids in the tuple pointed to by PPnt are equal to the Tids in the tuple pointed to by the pointer PxPnt; S5046: If the insertion position insertPos exceeds the size of the fuzzy list buffer, then allocate new memory space. Otherwise, recycle and reuse the memory space, add a new tuple to the fuzzy list buffer, set Tids to the Tids of PxPnt, fus to the sum of fus of PxPnt and fus of PyPnt minus fus of PPnt, and rfus to the rfus of PyPnt; S5047: After inserting the data, move the pointers PxPnt and PyPnt one position to the right simultaneously; S5048: When the pointer PxPnt does not point to the end position EndPos of the summary list SL(Px), and the pointer PyPnt does not point to the end position EndPos of the summary list SL(Py), repeat the fuzzy list buffer program; S5049: If the variable EAMeasure is less than the minimum threshold minUtil, return the result false; S50410: Update the summary list SL(Pxy), return the result true, and end the fuzzy list buffer construction program.

2. The method for mining high-fuzzy utility item sets based on a fuzzy list buffer according to claim 1, characterized in that, The fuzzy list buffer FLBuf is composed of triples (Tids, fus, rfus). Tid is the transaction identifier in the database, fu is the fuzzy utility of the transaction, and rfu is the remaining fuzzy utility of the transaction.

3. The method for mining high-fuzzy utility item sets based on a fuzzy list buffer according to claim 1, characterized in that, The summary list SL is composed of tuples (Itemsets, StartPoss, EndPoss, sumFus, sumRfus), where Itemset represents a fuzzy item set, StartPos and EndPos respectively represent the start and end positions of the corresponding fuzzy item set in the fuzzy list buffer FLBuf, sumFu represents the sum of the fuzzy utilities fus of the corresponding fuzzy item set in the fuzzy list buffer, and sumRfu represents the sum of the remaining fuzzy utilities rfus of the corresponding fuzzy item set in the fuzzy list buffer FLBuf.

4. The method for mining high-fuzzy utility item sets based on a fuzzy list buffer according to claim 3, characterized in that, After the recursive search subroutine Search has checked a node and all its descendant nodes, the program starts to backtrack. At this time, the nodes that have been checked are no longer used, and the memory space allocated in the fuzzy list buffer FLBuf to store this node will be recycled and reused. The data of the new potential fuzzy item set is directly overwritten and written into the recycled memory space, and at the same time, the information in the summary list SL is updated to achieve memory reuse and reduce the memory consumption of the program.

5. The method for mining high-fuzzy utility item sets based on a fuzzy list buffer according to claim 1, characterized in that, The evaluation of the fuzzy utility co-occurrence structure EFuCS is represented in matrix form, with the fuzzy item set as the index and the value representing the upper bound of the fuzzy utility FUUB after the combination of two fuzzy item sets.

6. A system for mining high-fuzzy utility item sets based on a fuzzy list buffer, characterized in that, The system includes: a memory and a processor. The memory includes a program for a method of mining high fuzzy utility item sets based on a fuzzy list buffer. When the program for the method of mining high fuzzy utility item sets based on a fuzzy list buffer is executed by the processor, the following steps are implemented: S1: Initialize the data mining operation parameters. The data mining operation parameters include: the quantitative database D to be mined, the predefined membership function R, and the minimum fuzzy utility threshold minUtil of the result set. S2: Scan the transaction database D and calculate the upper bound of the fuzzy utility FUUB of single items according to the membership function R, and create an initial list I*. S3: Deposit the single fuzzy items whose upper bound values of fuzzy utility are not less than the minimum threshold minUtil into the initial list I*, and sort them in ascending order according to the upper bound values of fuzzy utility. S4: Scan the database D again to construct the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its auxiliary summary list SL. S5: Call the recursive search subroutine Search and pass in parameters, where the parameters include: the initial prefix fuzzy item set Initialize the list I*, the minimum fuzzy utility threshold minUtil, the evaluation fuzzy utility co-occurrence structure EFuCS, the fuzzy list buffer FLBuf and its summary list SL; S6: Output all high fuzzy utility item sets HFUIs whose fuzzy utilities are not less than the minimum threshold to complete the data mining.

7. A high-fuzziness utility item set mining system based on a fuzzy list buffer according to claim 6, characterized in that, The recursive search subroutine Search called in step S5 includes the following steps: S501: In the recursive search subroutine Search, for an extended fuzzy item set X of a fuzzy item set P, if the sum of the fuzzy utilities sumFu of the fuzzy item set X stored in the summary list SL(X) is not less than the minimum threshold minUtil, then add the fuzzy item set X to the set HFUIs of high fuzzy utility item sets. S502: If the sum of the fuzzy utility sumFu and the sum of the remaining fuzzy utilities sumRfu in the summary list SL(X) of the fuzzy item set X is not less than the minimum threshold minUtil, then the extended fuzzy item set of the fuzzy item set X may be a high fuzzy utility item set. S503: For another extended fuzzy item set Y of the fuzzy item set P, where Y comes after the fuzzy item set X, find that the fuzzy item set Y satisfies: the upper bound value of the fuzzy utility of the fuzzy item sets X and Y in the evaluation of the fuzzy utility co-occurrence structure EFuCS is not less than the minimum threshold minUtil; S504: Call the fuzzy list buffer construction program with the fuzzy list buffer FLBuf, the summary list SL, the fuzzy item sets P, X, Y, and the minimum threshold minUtil as parameters, and return the construction result; S505: If the construction result returns true, then merge the fuzzy item sets X and Y into Pxy. If the sum of the fuzzy utilities of the summary list SL(Pxy) of the fuzzy item set Pxy is greater than 0, then add the fuzzy item set Pxy to the set ExtensionsOfX of the extended fuzzy item sets of the fuzzy item set X; S506: Merge the fuzzy item sets P and X as the new prefix fuzzy item set Px, and recursively call the search subroutine Search until all extended fuzzy item sets are traversed.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a program for the high fuzzy utility item set mining method based on the fuzzy list buffer. When the program for the high fuzzy utility item set mining method based on the fuzzy list buffer is executed by a processor, the steps of a high fuzzy utility item set mining method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • High-utility item set mining method containing negative utility

    CN110471960A

  • Top-k efficient item set mining method based on data buffer pool

    CN111241136A