Database index processing method and device, electronic equipment, storage medium and computer program product
By classifying and pruning the indicator items in the database metric transaction, the corresponding relationship is established, and the problem of high data mining load in the existing technology is solved, and the effect is applicable to large-scale data mining scenarios is achieved.
Patent Information
- Application Number
- CN202510075669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
AI Technical Summary
Existing data mining technologies have high load when processing database metrics, making them difficult to be suitable for large-scale data mining scenarios.
By classifying and pruning the indicator items in the indicator transaction in the database, the correspondence between each indicator item and the corresponding item set is determined, and the amount of data calculation and complexity in the data mining process is reduced.
The load during data mining is reduced, making this method suitable for large-scale data mining scenarios.
Smart Images

Figure CN119938737A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data mining, and in particular to a database index processing method, device, electronic device, storage medium and computer program product. Background Art
[0002] In related technologies, the indicators in the database hide the characteristics of users' use of the database. Through the data mining method of association rules, the linkage relationship between the indicators in the user's database use process is analyzed to provide a reference for in-depth understanding of the usage of the entire database indicators. The data mining of association rules often determines the frequent indicator item sets in the database indicators through the Apriori algorithm and the FP-tree (Frequent Pattern tree) algorithm.
[0003] Among them, the Apriori algorithm calculates candidate item sets and frequent item sets to generate a large number of candidate item sets, which consumes a lot of computing and memory resources. The computational complexity is high. Each calculation of a frequent item set requires scanning the database once, and the database needs to be scanned multiple times, resulting in high load, which makes it difficult to apply to large-scale data mining scenarios. The FP-growth (Frequent Pattern-growth) algorithm performs a secondary scan on the transaction database to avoid the generation of a large number of candidate sets. However, due to the need to construct an FP-tree, the tree has many child nodes, the memory overhead is large and the load is high, and it can only be used to mine single-dimensional Boolean association rules, and the applicable scenario is single. Therefore, the technical problem of data mining in related technologies is that the load is high during the data mining process, and it is difficult to apply to large-scale data mining scenarios. Summary of the invention
[0004] The embodiments of the present application provide a database index processing method, device, electronic device, storage medium and computer program product, which can reduce the load in the data mining process.
[0005] The technical solution of this application is implemented as follows:
[0006] The present application embodiment provides a database indicator processing method, including:
[0007] For the indicator items included in the indicator transaction in the database, the indicator items are classified and pruned, and a first corresponding relationship between each of the indicator items and at least one corresponding item set is determined; wherein the item set is formed by the indicator items in the corresponding indicator transaction;
[0008] A target item set is determined based on the indicator-related data corresponding to the item set in each of the first corresponding relationships.
[0009] In the above solution, the index items included in the index transaction in the database are classified and pruned, and the first corresponding relationship between each of the index items and the corresponding at least one item set is determined, including:
[0010] For the indicator items included in the indicator transaction in the database, perform indicator classification processing to determine a second corresponding relationship between each indicator item and the corresponding at least one item set;
[0011] Based on the number of occurrences of the indicator transaction corresponding to the item set, the formed second corresponding relationship is pruned to determine the first corresponding relationship between each indicator item and the corresponding at least one item set.
[0012] In the above solution, the index items included in the index transaction in the database are classified and the second corresponding relationship between each index item and the corresponding at least one item set is determined, including:
[0013] Traversing the indicator transactions in the database based on a preset item set, determining each indicator item in the indicator transaction that matches a preset indicator in the preset item set;
[0014] Determine each of the item sets based on each of the indicator items included in each of the indicator transactions;
[0015] The second corresponding relationship is constructed based on each of the indicator items and the item set to which each of the indicator items belongs.
[0016] In the above scheme, the method further comprises:
[0017] The indicator-related data of each item set is formed based on the number of the indicator items in each item set and the corresponding number of occurrences, and the indicator-related data is marked to the corresponding item set.
[0018] In the above solution, the pruning process is performed on the formed second corresponding relationship based on the number of occurrences of the indicator transaction corresponding to the item set to determine the first corresponding relationship between each of the indicator items and the corresponding at least one item set, including:
[0019] Based on the number of occurrences of the item set corresponding to the indicator-related data, determining a frequent item set and an infrequent item set in at least one item set included in each second corresponding relationship; wherein the number of occurrences corresponding to the frequent item set is not less than a first threshold;
[0020] Determine the support of the index item in each of the second corresponding relationships based on the number of occurrences of the frequent item set in each of the second corresponding relationships;
[0021] A target support less than a second threshold is determined in each of the supports, and the second corresponding relationship to which the index item corresponding to the target support belongs is deleted, and the remaining second corresponding relationship is determined to be the first corresponding relationship.
[0022] In the above solution, the step of determining the target item set based on the indicator-related data corresponding to the item set in each of the first corresponding relationships includes:
[0023] Determine, based on the indicator-related data of the item set in each of the first corresponding relationships, frequent item sets in descending order;
[0024] The target item set is determined based on the frequent item sets arranged in descending order.
[0025] In the above solution, determining the frequent itemsets in descending order based on the indicator-related data of the itemsets in each of the first corresponding relationships includes:
[0026] Based on the sum of the number of occurrences corresponding to the frequent item sets included in each of the first corresponding relationships and the number of the index items in the frequent item sets, each of the first corresponding relationships is arranged in descending order to determine the first corresponding relationships with a certain order; wherein, if the sum of the number of occurrences corresponding to N first corresponding relationships is the same, the N first corresponding relationships are arranged in descending order according to the number of the index items in the frequent item sets; N is an integer greater than 1;
[0027] Based on the frequent item sets included in the first corresponding relationship in a certain order, the frequent item sets arranged in descending order are determined.
[0028] In the above solution, determining the target item set based on the frequent item sets arranged in descending order includes:
[0029] Traversing the frequent item sets in descending order according to sorting, and determining M sub-item sets by combining indicators based on the indicator items included in the first frequent item set; wherein M is an integer greater than 1;
[0030] Based on the M sub-itemsets and the frequent itemsets in descending order, a target sub-itemset corresponding to the first frequent itemset is determined until the traversal of the frequent itemsets in descending order is completed, and the target itemset is determined by combining the target sub-itemset corresponding to each frequent itemset in descending order and the frequent itemset.
[0031] In the above solution, determining the target sub-item set corresponding to the first frequent item set based on the M sub-item sets and the frequent item sets arranged in descending order includes:
[0032] After performing deduplication processing on the M sub-item sets, K sub-item sets that do not match the frequent item sets are determined from the sub-item sets; wherein K is an integer not greater than M;
[0033] Based on the number of occurrences corresponding to each of the sub-item sets, T sub-frequent item sets are determined from the K sub-item sets; wherein T is an integer not greater than K;
[0034] Determine T of the sub-frequent item sets and the frequent item set in the first corresponding relationship as the target sub-item set.
[0035] In the above scheme, the method further comprises:
[0036] For a group of association index items, based on the item set to which each of the association index items belongs in the first corresponding relationship, an association rule between the group of association index items is determined.
[0037] The present application also provides a database indicator processing device, including:
[0038] A transaction processing unit, configured to classify and prune indicators for indicator items included in indicator transactions in a database, and determine a first corresponding relationship between each indicator item and at least one corresponding item set; wherein the item set is formed by the indicator items in the corresponding indicator transaction;
[0039] a determining unit, configured to determine a maximum itemset based on the indicator-related data of the itemset in each of the first corresponding relationships;
[0040] A determining unit is configured to determine a target item set based on the maximum item set and each of the first corresponding relationships.
[0041] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the computer program.
[0042] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.
[0043] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps in the above method when executed by a processor.
[0044] In the embodiment of the present application, the index items included in the index transaction in the database are classified and pruned to determine the first corresponding relationship between each index item and at least one corresponding item set; wherein the item set is formed by the index items in the corresponding index transaction; based on the index-related data corresponding to the item set in each first corresponding relationship, the target item set is determined. In this way, the classification and pruning of the index items included in the index transaction can greatly reduce the amount and complexity of data calculation in the data mining process, thereby reducing the load in the target item set mining process, so that the present solution can be applied to large-scale data mining scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0046] Figure 2 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0047] Figure 3 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0048] Figure 4 An optional effect diagram of the database indicator processing method provided in the embodiment of the present application;
[0049] Figure 5 An optional effect diagram of the database indicator processing method provided in the embodiment of the present application;
[0050] Figure 6 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0051] Figure 7 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0052] Figure 8 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0053] Fig. 9 An optional effect diagram of the database indicator processing method provided in the embodiment of the present application;
[0054] Fig.10 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0055] Fig.11 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0056] Fig.12 An optional flowchart of a database indicator processing method provided in an embodiment of the present application;
[0057] Fig.13 A schematic diagram of the structure of a database indicator processing device provided in an embodiment of the present application;
[0058] Fig.14 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0060] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0061] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0063] In related technologies, database monitoring indicators are mainly collected in real time through various exporter components (kube-exporter, node-exporter, mysqld-exporter, volume-exporter, etc.). Database monitoring indicators are mainly divided into CPU, memory, disk usage, iops, network bandwidth, number of connections, number of transactions, etc. At present, related research on database monitoring indicator analysis mainly uses statistical, classification and prediction methods to achieve data statistics, abnormal monitoring and abnormal diagnosis of a single monitoring indicator.
[0064] The association rule learning method reflects the interdependence between an indicator and other indicators, and can effectively discover the potential useful information in the data. The monitoring data in the database server hides the user's use characteristics of the database. Through the data mining method of association rules, the linkage relationship between the various monitoring sub-items in the user's database use process is analyzed, providing a reference for in-depth understanding of the use of the entire database monitoring indicators. The most classic algorithm of the association rule learning method mainly determines the frequent item sets through the Apriori algorithm and the FP-tree algorithm.
[0065] The Apriori association rule algorithm uses an iterative method of layer-by-layer search, traversing the data set multiple times, calculating candidate item sets of each length combination, and finding all frequent item sets whose support is greater than or equal to the minimum support threshold. Management rules are generated through confidence calculation, which is suitable for association rule mining of transaction databases and sparse data sets. However, the Apriori association rule algorithm has a large memory overhead and is relatively limited in its applicable scenarios.
[0066] In order to reduce the number of I / O traversals, the FP-tree association rule algorithm introduces a tree data structure to temporarily store data. It only needs to scan the data set twice to construct the FP tree and mine frequent item sets from the FP tree. Due to the large number of child nodes in the tree and the need to recursively generate the conditional database and conditional FP-tree, the memory overhead is large and it can only be used to mine single-dimensional Boolean association rules.
[0067] In order to solve the above technical problems, the present application embodiment provides a database indicator processing method, please refer to Figure 1 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, which will be combined with Figure 1 The steps shown are explained:
[0068] S101. Classify and prune the indicator items included in the indicator transaction in the database, and determine a first corresponding relationship between each indicator item and at least one corresponding item set.
[0069] In an embodiment of the present application, the database indicator processing device can traverse each indicator included in the indicator transaction in the database, classify each indicator item into a corresponding item set, and establish a corresponding relationship between each indicator item and the corresponding item set. Wherein, the item set is formed by the indicator items in the corresponding indicator transaction. Then, based on the number of indicator items in the item set and the number of times the indicator transaction corresponding to the item set appears in the database, pruning processing is performed on the established corresponding relationship to determine the first corresponding relationship between each indicator item and the corresponding at least one item set.
[0070] In the embodiment of the present application, the indicator transaction may be information recorded for each indicator stored in the database. For example, it may be transaction information recorded for the central processing unit (CPU), memory, storage, backup, number of transactions per second (TPS), query rate per second (QPS), number of read / write operations per second (IOPS), number of connections and throughput.
[0071] Each indicator item may include an identifier of a corresponding indicator, and the item set corresponding to the indicator transaction may be a set of identifiers of all indicator items included in the indicator transaction.
[0072] S102: Determine a target item set based on the indicator-related data corresponding to the item set in each of the first corresponding relationships.
[0073] In an embodiment of the present application, after the classification and pruning process determines multiple first corresponding relationships, the corresponding indicator-related data can be determined for the item set in each first corresponding relationship, and the corresponding item set can be marked. The indicator-related data includes the number of indicator items in the corresponding item set, and the number of occurrences of the corresponding indicator transaction. The processing device can determine frequent item sets and infrequent item sets in the item set in each first corresponding relationship based on the indicator-related data marked by the item set in each first corresponding relationship. And based on the number of occurrences corresponding to the frequent item set in each first corresponding relationship, all the first corresponding relationships are arranged in descending order according to the sum of the number of occurrences, and the item set with the largest number of occurrences and the largest number of indicator items is determined as the largest item set in the sorted multiple first corresponding relationships. Finally, the target item set is determined based on the largest item set and the determined frequent item sets.
[0074] In an embodiment of the present application, after determining the target item set, the relationship between the database monitoring indicator items or the indicators of each subsystem can be mined through association rule technology to provide technical support for abnormal monitoring or user usage habit analysis. By providing database monitoring indicator association rule analysis services, a higher-level data analysis model is provided to users. Data mining technology is used to deeply analyze the internal connections of monitoring data and provide a reference for the layout of monitoring sub-item pages. From the perspective of operation and maintenance managers, it is possible to avoid making a single decision based solely on some raw data, and instead use the feedback from the data mining system to plan decision-making plans from multiple angles, making it convenient for operation and maintenance personnel to locate the range of abnormal indicators more quickly and reduce the time for troubleshooting database anomalies.
[0075] In the embodiment of the present application, the index items included in the index transaction in the database are classified and pruned to determine the first corresponding relationship between each index item and at least one corresponding item set; wherein the item set is formed by the index items in the corresponding index transaction; based on the index-related data corresponding to the item set in each first corresponding relationship, the target item set is determined. In this way, the classification and pruning of the index items included in the index transaction can greatly reduce the amount and complexity of data calculation in the data mining process, thereby reducing the load in the target item set mining process, so that the present solution can be applied to large-scale data mining scenarios.
[0076] See also Figure 2 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 1 The illustrated S101 can also be implemented through S201 to S202, which will be described in conjunction with the steps:
[0077] S201: Classify the indicator items included in the indicator transaction in the database, and determine a second corresponding relationship between each indicator item and the corresponding at least one item set.
[0078] In an embodiment of the present application, the processing device traverses each indicator item included in the indicator transaction in the database, extracts each indicator item, and determines at least one indicator transaction to which each indicator item belongs. A corresponding item set is determined based on all indicator items in the indicator transaction. A second corresponding relationship is established between each indicator item and the corresponding at least one item set.
[0079] S202: Perform pruning processing on the formed second corresponding relationship based on the number of occurrences of the indicator transaction corresponding to the item set, and determine the first corresponding relationship between each of the indicator items and the corresponding at least one item set.
[0080] In an embodiment of the present application, based on the number of occurrences corresponding to the item sets in each second corresponding relationship, the frequent item sets and infrequent item sets in each second corresponding relationship are determined, and based on the sum of the number of occurrences corresponding to the frequent item sets included in each second corresponding relationship, the formed second corresponding relationships are pruned to determine the remaining second corresponding relationships as the first corresponding relationships.
[0081] In the embodiment of the present application, the index items included in the index transaction in the database are classified and processed to determine the second corresponding relationship between each index item and the corresponding at least one item set. Based on the number of occurrences of the index transaction corresponding to the item set, the formed second corresponding relationship is pruned to determine the first corresponding relationship between each index item and the corresponding at least one item set. In this way, the second corresponding relationship is determined by classifying the index items included in the index transaction, and then the second corresponding relationship is pruned based on the number of occurrences corresponding to the item set. This can greatly reduce the amount and complexity of data calculation in the data mining process, thereby reducing the load in the target item set mining process, so that this solution can be applied to large-scale data mining scenarios.
[0082] See also Figure 3 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 2 S201 to S202 shown can also be implemented by S301 to S306, which will be described in combination with the steps:
[0083] S301. Traverse the indicator transactions in the database based on a preset item set, and determine each indicator item in the indicator transaction that matches a preset indicator in the preset item set.
[0084] In an embodiment of the present application, the processing device may obtain a preset item set, wherein the preset item set includes a plurality of preset indicators. The processing device may match the indicator items included in the indicator transaction of the database based on each preset indicator, and determine each of the indicator items in the indicator transaction that matches the preset indicator in the preset item set.
[0085] S302: Determine each of the item sets based on each of the indicator items included in each of the indicator transactions.
[0086] In the embodiment of the present application, after each indicator item included in each indicator transaction is determined, the indicator items included in each indicator transaction may be combined to determine an item set corresponding to each indicator transaction.
[0087] For example, first determine a preset item set I = {I1, I2, ..., I m}, m is the number of preset indicators, and I is the preset indicator. Scan the database D={t1,t2,...,t n}, n is the number of indicator transactions, t is the indicator transaction. Based on the preset item set, each transaction t is traversed, the indicator item corresponding to the preset indicator in each indicator transaction is determined, and then all the indicator items of each indicator transaction are combined to determine the corresponding item set.
[0088] S303: construct the second corresponding relationship based on each of the indicator items and the item set to which each of the indicator items belongs.
[0089] In the embodiment of the present application, the processing device may construct a second corresponding relationship based on each indicator item included in the multiple item sets and at least one item set to which each indicator item belongs.
[0090] In the embodiment of the present application, the processing device forms the indicator-related data of each item set based on the number of the indicator items in each item set and the corresponding number of occurrences, and marks the indicator-related data to the corresponding item set.
[0091] Combination Figure 4 The processing device performs traversal matching on the indicator transactions (Event1-Event3) to determine the item set corresponding to each indicator transaction that matches the preset indicator. Determine the number of indicator items in the item set and the number of item set occurrences, and record the number of indicator items and the number of item set occurrences and store them in the item set Map collection. According to at least one item set to which each indicator item (I1-I5) in the multiple item sets belongs, insert the item set (I1I3I5 and I2I3) into the inverted index list corresponding to each indicator item in sequence, the inverted index key is the name of each indicator item, and the value is each item set, forming a second corresponding relationship.
[0092] In the embodiment of the present application, a candidate item set index (second corresponding relationship) is constructed by traversing the item set set, thereby avoiding directly using the full transaction to construct an inverted index, reducing the complexity of the inverted index construction, and further reducing the data mining load, thereby improving the efficiency of determining the target item set.
[0093] S304: Determine a frequent item set and an infrequent item set in at least one item set included in each second corresponding relationship based on the number of occurrences of the item set corresponding to the indicator-related data.
[0094] In an embodiment of the present application, the processing device determines, based on the number of occurrences in the indicator-related data corresponding to each item set, a frequent item set whose number of occurrences is not less than a first threshold, and an infrequent item set whose number of occurrences is less than the first threshold in at least one item set included in each second corresponding relationship.
[0095] Exemplarily, the number of occurrences of each determined item set is compared with the first threshold minsup, and the support(I i )≥minsup determines the frequent item sets. Among them, support(I i ) is the support (number of occurrences), and minsup is the first threshold. The item sets whose number of occurrences is not less than the first threshold are frequent item sets, and the item sets whose number of occurrences is less than the first threshold are frequent item sets. Figure 5, frequent itemsets can be marked for itemsets I1I3I5: true. Infrequent itemsets can be marked for itemsets I2I3: false.
[0096] S305: Determine the support of the indicator item in each of the second corresponding relationships based on the number of occurrences of the frequent item sets in each of the second corresponding relationships.
[0097] In the embodiment of the present application, the processing device determines the support corresponding to the indicator item in each second corresponding relationship based on the sum of the number of occurrences corresponding to the frequent item set in each second relationship.
[0098] Exemplary, combined Figure 5 , for the second corresponding relationship I1-I1I3I5, the support 2 corresponding to the index item can be determined based on the number of occurrences of the frequent item set I1I3I5.
[0099] S306: Determine a target support less than a second threshold value in each of the supports, delete the second corresponding relationship to which the indicator item corresponding to the target support belongs, and determine the remaining second corresponding relationship as the first corresponding relationship.
[0100] In the embodiment of the present application, the processing device traverses the support of the index item in each second corresponding relationship, determines the target support less than the second threshold, deletes the second corresponding relationship to which the index item corresponding to the target support belongs, and determines the remaining second corresponding relationship as the first corresponding relationship.
[0101] Exemplary, combined Figure 5 , traverse the inverted index and use formula (1) to calculate the support of each index item, where I i represents the i-th indicator item, represents the jth frequent item set in index item i, Represents frequent itemsets Support (number of occurrences). If the item set list in the inverted index (second correspondence) is empty or the index item support does not meet the minimum support (second threshold), delete the inverted index item. If the minimum support is set to 2, delete Figure 5 The second correspondence marked by the thick box determines that the remaining inverted index items are the first correspondence.
[0102]
[0103] In an embodiment of the present application, for the determined second corresponding relationship, the itemsets therein are pruned based on the number of occurrences of the frequent itemsets, thereby avoiding the subsequent support calculation of redundant itemsets, improving the algorithm execution efficiency, and thereby reducing the data mining load and improving the efficiency of determining the target itemsets.
[0104] See also Figure 6 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 1 The illustrated S102 can also be implemented through S401 to S402, which will be described in conjunction with the steps:
[0105] S401: Determine frequent item sets in descending order based on the indicator-related data of the item sets in each of the first corresponding relationships.
[0106] In the embodiment of the present application, the processing device can determine the first corresponding relationships in a certain order by arranging the first corresponding relationships in descending order according to the sum of the number of occurrences of the corresponding frequent itemsets and the number of indicators of the corresponding frequent itemsets based on the index-related data of the itemsets in each first corresponding relationship. Then, the frequent itemsets in descending order are determined based on the frequent itemsets included in the first corresponding relationships in the certain order.
[0107] S402: Determine the target item set based on the frequent item sets arranged in descending order.
[0108] In the embodiment of the present application, the processing device uses the frequent item set law: all subsets of the maximum frequent item set are frequent item sets. Thus, it can be quickly determined that all subsets under the maximum frequent item set are frequent item sets. The subsets of each frequent item set in descending order are determined, and the determined subsets are deduplicated and combined with the frequent item sets in the first corresponding relationship to determine the target item set.
[0109] In the embodiment of the present application, based on the index-related data of the item set in each first corresponding relationship, the frequent item set in descending order is determined. The target item set is determined based on the frequent item set in descending order. In this way, the process of querying redundant item sets can be effectively avoided, the computational overhead is reduced, and the efficiency of determining the target item set is improved.
[0110] See also Figure 7 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 6 The illustrated S401 can also be implemented through S501 to S502, which will be described in conjunction with the steps:
[0111] S501: Based on the sum of the number of occurrences of the frequent itemsets included in each first corresponding relationship and the number of the index items in the frequent itemsets, each first corresponding relationship is sorted in descending order to determine the first corresponding relationships with a certain order.
[0112] In the embodiment of the present application, based on the sum of the number of occurrences corresponding to the frequent item sets included in each first corresponding relationship and the number of index items of the frequent item sets in each first corresponding relationship, each first corresponding relationship can be arranged in descending order according to the sum of the number of occurrences to determine the first corresponding relationships with a certain order. If the sum of the number of occurrences corresponding to N first corresponding relationships is the same, the N first corresponding relationships are arranged in descending order according to the number of index items in the frequent item sets; N is an integer greater than 1.
[0113] In the embodiment of the present application, the processing device may perform a plurality of first corresponding items according to the sum of the number of occurrences of the corresponding frequent itemsets. Arrange in descending order to obtain a descending index item tag sequence. i ) in descending order, and the support calculation method is the sum of the supports of each frequent item set of the indicator item, as shown in formula (1), and finally the first corresponding relationship arrangement in descending order of support is obtained. Among them, the frequent item set items can be double-arranged in descending order according to the number of indicator items and support in the frequent item set items to find the frequent item sets with decreasing item set length and support, effectively avoiding the query of other redundant item set combinations and the calculation overhead of support.
[0114] S502: Determine the frequent item sets in descending order based on the frequent item sets included in the first corresponding relationship in a certain order.
[0115] In the embodiment of the present application, after completing the descending order arrangement of multiple first corresponding relationships, the frequent item sets included in each first corresponding relationship can be extracted in order, and the frequent item sets can be arranged in the order of each first corresponding relationship to determine the frequent item sets arranged in descending order.
[0116] In the embodiment of the present application, each first corresponding relationship is arranged in descending order according to the sum of the number of occurrences of the frequent item sets included and the number of index items in the frequent item sets, and a first corresponding relationship with a certain order is determined. Based on the frequent item sets included in the first corresponding relationship in a certain order, the frequent item sets arranged in descending order are determined. In this way, by arranging the first corresponding relationship in descending order according to the number of occurrences of the frequent item sets and the number of index items included, each item set can be accurately sorted according to the number of occurrences and the corresponding number of index items, and then the frequent item sets arranged in descending order can be quickly determined, thereby improving the efficiency of determining the target item set.
[0117] See also Figure 8 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 6 The illustrated S402 can also be implemented through S601 to S602, which will be described in conjunction with the steps:
[0118] S601 , traverse the frequent item sets arranged in descending order according to sorting, and determine M sub-item sets by combining indicators based on the indicator items included in the first frequent item set.
[0119] In the embodiment of the present application, the processing device may traverse the frequent itemsets arranged in descending order according to the sorting, and determine M sub-itemsets corresponding to the first frequent itemset by combining indicators based on multiple indicator items included in the first frequent itemset, where M is an integer greater than 1.
[0120] Exemplary, combined Fig. 9 , based on the multiple index items included in the frequent itemset, the index combination is performed from left to right, and the item combination order is in the fixed item order, and the length is n1,...,n1-k,...,1,1≤k≤n1-1, where n1 is the number of index items in the maximum item set, and the M sub-item sets of the maximum item set are obtained by combination. Among them, the index items in the maximum item set may include: I1, I2, I3, I4 and I5. Each index item in I1, I2, and I3 can be traversed, and combined with I4 and I5 to determine the sub-item set.
[0121] S602: Determine a target sub-item set corresponding to the first frequent item set based on the M sub-item sets and the frequent item sets in descending order, until the traversal of the frequent item sets in descending order is completed, and determine the target item set by combining the target sub-item set corresponding to each frequent item set in descending order and the frequent item set.
[0122] In the embodiment of the present application, the processing device can deduplicate the determined M sub-item sets, and determine a portion of the sub-item sets in the M sub-item sets that meet the requirements of the frequent item sets, and determine the portion of the sub-item sets as the target sub-item sets corresponding to the first frequent item set, until the traversal of the frequent item sets in descending order is completed, and the target item set is determined by combining the target sub-item sets corresponding to each frequent item set in descending order and the frequent item sets.
[0123] In an embodiment of the present application, the frequent item sets arranged in descending order are traversed in a sorted manner, and based on the indicator items included in the first frequent item set, M sub-item sets are determined by combining indicators. . Based on the M sub-item sets and the frequent item sets arranged in descending order, a target sub-item set corresponding to the first frequent item set is determined, until the traversal of the frequent item sets arranged in descending order is completed, and the target item set is determined by combining the target sub-item set corresponding to each frequent item set in descending order and the frequent item set. In this way, the corresponding sub-item set is determined according to the frequent item set law, and then the target item set is determined in combination with the frequent item sets in the first corresponding relationship, thereby avoiding the process of calculating redundant item sets, reducing the calculation overhead, and improving the efficiency of determining the target item set.
[0124] See also Fig.10 , is an optional flow chart of the database index processing method provided in the embodiment of the present application, Figure 8 The illustrated S602 can also be implemented through S701 to S703, which will be described in conjunction with the steps:
[0125] S701 , after performing deduplication processing on the M sub-item sets, determine K sub-item sets that do not match the frequent item sets from the sub-item sets.
[0126] In the embodiment of the present application, after performing deduplication processing on the determined M sub-item sets, the processing device matches each sub-item set with all frequent item sets in the first corresponding relationship to determine K sub-item sets that do not match the frequent item sets.
[0127] The processing device may perform deduplication processing on the determined M sub-item sets, including: comparing and deduplicating each sub-item set with all determined sub-item sets, and comparing and deduplicating each sub-item set with each determined frequent item set.
[0128] S702: Determine T sub-frequent item sets from the K sub-item sets based on the number of occurrences corresponding to each of the sub-item sets.
[0129] In the embodiment of the present application, the processing device may determine T frequent sub-item sets among the K sub-item sets based on the number of occurrences of each sub-item set in the item set corresponding to the indicator transaction.
[0130] Among them, the sub-item set whose corresponding number of occurrences is greater than the first threshold can be determined as the sub-frequent item set.
[0131] S703: Determine T of the sub-frequent item sets and the frequent item set in the first corresponding relationship as the target sub-item set.
[0132] In the embodiment of the present application, the processing device may determine that the T sub-frequent item sets and the frequent item set in the first corresponding relationship are the target item set.
[0133] In the embodiment of the present application, if the combined sub-frequent item set is not in the frequent item set, the sub-item set support is calculated according to formula (1). Sub-set marked frequent item sets of each length item set combination are calculated in sequence, and the lengths are decreased in sequence until the last marked frequent item set combination is reached. The query is terminated, and all length frequent k-item sets and supports are obtained, and the algorithm execution is terminated. Specific algorithm flow:
[0134] 1. Based on the first corresponding relationship of the number of index items and support of frequent itemsets, the set of labeled frequent itemsets with decreasing length and support is determined.
[0135] 2. Traverse the descending marked frequent itemsets in sequence, and combine the sub-frequent itemsets in the order from left to right. The item combination order follows the fixed item order, and the lengths are n1,...,n1-k,...,1,1≤k≤n1-1.
[0136] 3. Create a set of labeled sub-frequent itemsets with an initial value of empty.
[0137] 4. Traverse the sub-frequent item set combination to determine whether the sub-frequent item set is in the marked frequent item set set and the sub-marked frequent item set set. If it is not in the two sets, calculate the inverted index support according to formula
[0002] to determine whether it is a frequent item set. If it is a sub-frequent item set, save it in the marked sub-frequent item set set; if it is in both sets, skip it.
[0138] 5. After all sub-frequent item set combinations are calculated, the marked frequent item set set and the marked sub-frequent item set set are merged. The merged result is the full frequent item set, and the algorithm is executed.
[0139] In the embodiment of the present application, after deduplication processing is performed on the M sub-item sets, K sub-item sets that do not match the frequent item set are determined in the sub-item sets; based on the number of occurrences corresponding to each sub-item set, T sub-frequent item sets are determined in the K sub-item sets; and the T sub-frequent item sets and the frequent item set in the first corresponding relationship are determined as the target sub-item set. In this way, deduplication is performed on the determined sub-item sets, and the sub-frequent item sets therein are determined, which can reduce the redundant sub-item sets in the process of determining the target sub-item set, reduce the calculation load, and improve the efficiency of determining the target item set.
[0140] See also Fig.11 , which is an optional flow chart of the database indicator processing method provided in the embodiment of the present application, will be described in combination with the steps:
[0141] S801: For a group of association index items, determine an association rule between the group of association index items based on the item set to which each of the association index items belongs in the first corresponding relationship.
[0142] In an embodiment of the present application, the processing device can determine the confidence and lift between a group of associated index items based on the item set to which each associated index item belongs in the first corresponding relationship, and determine the association rule between the group of associated index items based on the determined confidence and lift.
[0143] In the embodiment of the present application, the processing device can calculate the confidence and lift of each association index item according to its support (number of occurrences), and the association rule that meets the minimum confidence and minimum lift is called a strong association rule. As shown in formula (2) and formula (3):
[0144]
[0145] in, It is used to indicate the confidence between indicator item I1 and indicator item I2. support(I1∪I2) is used to indicate the support of the item set that includes both indicator item I1 and indicator item I2. support(I1) indicates the support of the item set that includes indicator item I1.
[0146]
[0147] in, It is used to indicate the lift between index item I1 and index item I2, and support(I2) indicates the support of the item set including index item I2.
[0148] In the embodiment of the present application, the processing device mines the strong association rules between database indicators based on the multi-dimensional indicators of support, confidence and lift. The simulation data and the real monitoring data are selected according to different orders of magnitude in the selection of data volume to ensure the accuracy and robustness of the proposed algorithm under data sets of different orders of magnitude. Provide a strong association scheme for indicators for database management decision makers to achieve the purpose of reasonable supervision of the database.
[0149] In an embodiment of the present application, a database monitoring indicator analysis method based on inverted index association rules is proposed, which aims to realize rapid association rule modeling of database indicators in a controllable and automated manner, and intelligently analyze database monitoring data. Association rule data mining is used to discover the relationship between different attributes in the monitoring sub-items, indicating the conditional probability that other attributes also occur under the premise that one or more attributes occur. Based on the analysis results of data mining, managers can understand the relationship between different associated attributes and the usage characteristics of different users to formulate more reasonable database management plans and promptly discover problems in the use of the database.
[0150] See also Fig.12 , which is an optional flow chart of the database indicator processing method provided in the embodiment of the present application, will be described in combination with the steps:
[0151] S11. Scan the transaction data set according to the preset index, match the transactions to the corresponding item sets one by one, and record the number of index items and the number of occurrences of the item sets.
[0152] In the embodiment of the present application, the processing device scans each indicator transaction in the transaction data set according to the preset indicator, determines each indicator item corresponding to each transaction, combines the indicator items corresponding to each transaction to determine the corresponding item set, and records the number and occurrence times of the indicator items in each item set.
[0153] S12. Associate the matched item sets to each inverted index item.
[0154] In the embodiment of the present application, the processing device may establish a corresponding relationship between each index item and at least one item set to which each index item belongs, that is, associate the item set with the corresponding inverted index item.
[0155] S13, perform pruning operations: mark whether the matching item set is a frequent item set and delete inverted index items with less than support.
[0156] In the embodiment of the present application, the processing device performs pruning processing on the determined inverted index, marks whether the matching item set is a frequent item set, and deletes the inverted index items with a value less than the support.
[0157] S14. Construct subsets with decreasing lengths.
[0158] In the embodiment of the present application, the processing device traverses each frequent item set and constructs corresponding subsets with decreasing lengths based on the frequent item sets.
[0159] S15. Traverse k-frequent itemsets.
[0160] S16. Calculate the sum of the support of the same matching item sets in the inverted index, which is the support of the candidate item set. Determine whether to mark it as a frequent item set based on the size of the support.
[0161] In the embodiment of the present application, the processing device calculates the sum of the support of the same matching item sets of the inverted index, which is the support of the candidate item set, and determines whether to mark it as a frequent item set according to the size of the support.
[0162] S17. Based on the length of frequent itemsets, mark the sub-frequent itemsets, and scan the inverted index to calculate the support of the sub-frequent itemsets.
[0163] In the embodiment of the present application, the processing device marks the sub-frequent itemsets based on the length frequent itemsets, and scans the inverted index to calculate the support of the sub-frequent itemsets.
[0164] S18. Based on the labeled frequent item sets and support, calculate the confidence and lift indicators, and analyze the correlation between various monitoring sub-items in the database.
[0165] In the embodiment of the present application, the processing device calculates the confidence and lift indexes based on the labeled frequent item sets and the support, and analyzes the correlation between various monitoring sub-items in the database.
[0166] In the embodiment of the present application, based on the proposed method for fast association rule analysis of inverted index database monitoring indicators, the database is scanned once to construct an inverted index relationship, and the transaction is converted into a candidate item set and classified into a marked candidate item set set. The length of the candidate item set is the number of items, and the number of candidate item set hits is the support size, so as to achieve effective compression of the full transaction; the candidate item set index is constructed by traversing the marked candidate item set set, avoiding directly constructing the inverted index with the full transaction, and reducing the complexity of inverted index construction; judging whether the marked candidate item set set meets the minimum support, and updating whether it is a frequent item set mark. Judging whether the inverted index is empty, determining whether to delete the inverted index item, and reducing the complexity of sub-frequent item set calculation through pruning operations; quickly determining the maximum frequent item set by double descending order of item set length and support, combining sub-frequent item sets, and quickly calculating the sub-frequent item set support according to the inverted index and the marked candidate item set set, effectively avoiding redundant candidate item set support calculation, and achieving fast search and support calculation of the full frequent item set.
[0167] In an embodiment of the present application, a method for quickly mining association rules of database indicators based on an inverted index is proposed. It only needs to scan the database once to construct a marked candidate item set and an inverted index, and prune the marked candidate item set and the inverted index according to the support threshold to reduce the processing of redundant item sets. The maximum frequent k item set is searched in order from large to small according to the item set length, and each subset frequent item set and support are quickly generated according to the maximum frequent item set, so as to achieve fast search of frequent item sets and improve the mining efficiency of monitoring sub-item association rules. The candidate item set index is constructed by traversing the marked candidate item set set, avoiding directly constructing the inverted index with the full transaction, and reducing the complexity of inverted index construction. In addition, the maximum frequent item set is quickly determined by double descending order of item set length and support, and the sub-frequent item sets are combined. The support of the sub-frequent item set is quickly calculated according to the inverted index and the marked candidate item set set, which effectively avoids the calculation of the support of redundant candidate item sets and achieves fast search and support calculation of the full frequent item set.
[0168] See also Fig.13 , which is a structural diagram of the database indicator processing device provided in an embodiment of the present application.
[0169] The embodiment of the present application provides a database indicator processing device 800, including: a transaction processing unit 801 and a determination unit 802.
[0170] The transaction processing unit 801 is used to classify and prune the indicator items included in the indicator transaction in the database, and determine a first corresponding relationship between each indicator item and at least one corresponding item set; wherein the item set is formed by the indicator items in the corresponding indicator transaction;
[0171] The determining unit 802 is configured to determine a target item set based on the indicator-related data corresponding to the item set in each of the first corresponding relationships.
[0172] In the embodiment of the present application, the transaction processing unit 801 in the database indicator processing device 800 is used to perform indicator classification processing on the indicator items included in the indicator transaction in the database, and determine the second corresponding relationship between each indicator item and the corresponding at least one item set;
[0173] Based on the number of occurrences of the indicator transaction corresponding to the item set, the formed second corresponding relationship is pruned to determine the first corresponding relationship between each indicator item and the corresponding at least one item set.
[0174] In the embodiment of the present application, the transaction processing unit 801 in the database indicator processing device 800 is used to traverse the indicator transaction in the database based on the preset item set, and determine each indicator item in the indicator transaction that matches the preset indicator in the preset item set;
[0175] Determine each of the item sets based on each of the indicator items included in each of the indicator transactions;
[0176] The second corresponding relationship is constructed based on each of the indicator items and the item set to which each of the indicator items belongs.
[0177] In an embodiment of the present application, the transaction processing unit 801 in the database indicator processing device 800 is used to form the indicator-related data of each item set based on the number of the indicator items in each item set and the corresponding number of occurrences, and mark the indicator-related data to the corresponding item set.
[0178] In the embodiment of the present application, the transaction processing unit 801 in the database indicator processing device 800 is used to determine a frequent item set and a non-frequent item set in at least one item set included in each second corresponding relationship based on the number of occurrences of the item set corresponding to the indicator-related data; wherein the number of occurrences corresponding to the frequent item set is not less than a first threshold;
[0179] Determine the support of the index item in each of the second corresponding relationships based on the number of occurrences of the frequent item set in each of the second corresponding relationships;
[0180] A target support less than a second threshold is determined in each of the supports, and the second corresponding relationship to which the index item corresponding to the target support belongs is deleted, and the remaining second corresponding relationship is determined to be the first corresponding relationship.
[0181] In the embodiment of the present application, the determination unit 802 in the database index processing device 800 is used to determine the frequent itemsets in descending order based on the index-related data of the itemsets in each of the first corresponding relationships;
[0182] The target item set is determined based on the frequent item sets arranged in descending order.
[0183] In the embodiment of the present application, the determination unit 802 in the database indicator processing device 800 is used to arrange each first corresponding relationship in descending order based on the sum of the number of occurrences corresponding to the frequent item sets included in each first corresponding relationship and the number of the indicator items in the frequent item sets, and determine the first corresponding relationships with a certain order; wherein, if the sum of the number of occurrences corresponding to N first corresponding relationships is the same, the N first corresponding relationships are arranged in descending order according to the number of the indicator items in the frequent item sets; N is an integer greater than 1;
[0184] Based on the frequent item sets included in the first corresponding relationship in a certain order, the frequent item sets arranged in descending order are determined.
[0185] In the embodiment of the present application, the determination unit 802 in the database indicator processing device 800 is used to traverse the frequent item sets arranged in descending order according to the sorting, and determine M sub-item sets by indicator combination based on the indicator items included in the first frequent item set; wherein M is an integer greater than 1;
[0186] Based on the M sub-itemsets and the frequent itemsets in descending order, a target sub-itemset corresponding to the first frequent itemset is determined until the traversal of the frequent itemsets in descending order is completed, and the target itemset is determined by combining the target sub-itemset corresponding to each frequent itemset in descending order and the frequent itemset.
[0187] In the embodiment of the present application, the determination unit 802 in the database indicator processing device 800 is used to perform deduplication processing on the M sub-item sets, and then determine K sub-item sets that do not match the frequent item set in the sub-item sets; wherein K is an integer not greater than M;
[0188] Based on the number of occurrences corresponding to each of the sub-item sets, T sub-frequent item sets are determined from the K sub-item sets; wherein T is an integer not greater than K;
[0189] Determine T of the sub-frequent item sets and the frequent item set in the first corresponding relationship as the target sub-item set.
[0190] In the embodiment of the present application, the determination unit 802 in the database indicator processing device 800 is used to determine, for a group of associated indicator items, an association rule between the associated indicator items based on the item set to which each associated indicator item belongs in the first corresponding relationship.
[0191] It should be noted that in the embodiment of the present application, if the above-mentioned database index processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a database index processing device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0192] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method on the data processing device side are implemented.
[0193] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0194] It should be noted that Fig.14 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.14 As shown, an embodiment of the present application provides an electronic device 900, including a memory 902 and a processor 901, wherein the memory 902 stores a computer program that can be run on the processor 901, and the processor 901 implements the steps in the above method when executing the program, wherein;
[0195] The processor 901 generally controls the overall operations of the electronic device 900 .
[0196] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0197] Correspondingly, an embodiment of the present application also provides a computer program product, including a computer program, which can be executed by the processor 901 of the electronic device 900 to complete the steps in the method on the side of the database indicator processing device 800.
[0198] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0199] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0200] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0201] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0202] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0203] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.
[0204] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.
[0205] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A database index processing method, characterized in that: include: For the indicator items included in the indicator transaction in the database, the indicator items are classified and pruned, and a first corresponding relationship between each of the indicator items and at least one corresponding item set is determined; wherein the item set is formed by the indicator items in the corresponding indicator transaction; A target item set is determined based on the indicator-related data corresponding to the item set in each of the first corresponding relationships.
2. The database indicator processing method according to claim 1, characterized in that: The step of classifying and pruning the indicator items included in the indicator transaction in the database, and determining a first corresponding relationship between each of the indicator items and the corresponding at least one item set, includes: For the indicator items included in the indicator transaction in the database, perform indicator classification processing to determine a second corresponding relationship between each indicator item and the corresponding at least one item set; Based on the number of occurrences of the indicator transaction corresponding to the item set, the formed second corresponding relationship is pruned to determine the first corresponding relationship between each indicator item and the corresponding at least one item set.
3. The database indicator processing method according to claim 2, characterized in that: The step of performing index classification processing on the index items included in the index transaction in the database and determining a second corresponding relationship between each index item and the corresponding at least one item set includes: Traversing the indicator transactions in the database based on a preset item set, determining each indicator item in the indicator transaction that matches a preset indicator in the preset item set; Determine each of the item sets based on each of the indicator items included in each of the indicator transactions; The second corresponding relationship is constructed based on each of the indicator items and the item set to which each of the indicator items belongs.
4. The database indicator processing method according to claim 3 is characterized in that: The method further comprises: The indicator-related data of each item set is formed based on the number of the indicator items in each item set and the corresponding number of occurrences, and the indicator-related data is marked to the corresponding item set.
5. The database indicator processing method according to claim 4 is characterized in that: The pruning process is performed on the formed second corresponding relationship based on the number of occurrences of the indicator transaction corresponding to the item set to determine the first corresponding relationship between each indicator item and the corresponding at least one item set, including: Based on the number of occurrences of the item set corresponding to the indicator-related data, determining a frequent item set and an infrequent item set in at least one item set included in each second corresponding relationship; wherein the number of occurrences corresponding to the frequent item set is not less than a first threshold; Determine the support of the index item in each of the second corresponding relationships based on the number of occurrences of the frequent item set in each of the second corresponding relationships; A target support less than a second threshold is determined in each of the supports, and the second corresponding relationship to which the index item corresponding to the target support belongs is deleted, and the remaining second corresponding relationship is determined to be the first corresponding relationship.
6. The database indicator processing method according to any one of claims 1 to 5, characterized in that: The determining of a target item set based on the indicator-related data corresponding to the item set in each of the first corresponding relationships includes: Determining frequent item sets in descending order based on the indicator-related data of the item sets in each of the first corresponding relationships; The target item set is determined based on the frequent item sets arranged in descending order.
7. The database index processing method according to claim 6, characterized in that: The step of determining the frequent itemsets in descending order based on the indicator-related data of the itemsets in each of the first corresponding relationships comprises: Based on the sum of the number of occurrences corresponding to the frequent item sets included in each of the first corresponding relationships and the number of the index items in the frequent item sets, each of the first corresponding relationships is arranged in descending order to determine the first corresponding relationships with a certain order; wherein, if the sum of the number of occurrences corresponding to N first corresponding relationships is the same, the N first corresponding relationships are arranged in descending order according to the number of the index items in the frequent item sets; N is an integer greater than 1; Based on the frequent item sets included in the first corresponding relationship in a certain order, the frequent item sets arranged in descending order are determined.
8. The database index processing method according to claim 6, characterized in that: The determining the target item set based on the frequent item sets arranged in descending order comprises: Traversing the frequent item sets in descending order according to sorting, and determining M sub-item sets by combining indicators based on the indicator items included in the first frequent item set; wherein M is an integer greater than 1; Based on the M sub-itemsets and the frequent itemsets in descending order, a target sub-itemset corresponding to the first frequent itemset is determined until the traversal of the frequent itemsets in descending order is completed, and the target itemset is determined by combining the target sub-itemset corresponding to each frequent itemset in descending order and the frequent itemset.
9. The database index processing method according to claim 8, characterized in that: The determining a target sub-item set corresponding to the first frequent item set based on the M sub-item sets and the frequent item sets arranged in descending order comprises: After performing deduplication processing on the M sub-item sets, K sub-item sets that do not match the frequent item sets are determined from the sub-item sets; wherein K is an integer not greater than M; Based on the number of occurrences corresponding to each of the sub-item sets, T sub-frequent item sets are determined from the K sub-item sets; wherein T is an integer not greater than K; Determine T of the sub-frequent item sets and the frequent item set in the first corresponding relationship as the target sub-item set.
10. The database indicator processing method according to any one of claims 1 to 5, characterized in that: The method further comprises: For a group of association index items, based on the item set to which each of the association index items belongs in the first corresponding relationship, an association rule between the group of association index items is determined.
11. A database index processing device, characterized in that: include: A transaction processing unit, configured to classify and prune indicators for indicator items included in indicator transactions in a database, and determine a first corresponding relationship between each indicator item and at least one corresponding item set; wherein the item set is formed by the indicator items in the corresponding indicator transaction; A determination unit is used to determine a target item set based on the indicator-related data corresponding to the item set in each of the first corresponding relationships.
12. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the method according to any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 10 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.