Sensitive frequent information hiding method, device and system and medium

By identifying and selecting appropriate target transactions and sensitive items, and using the protozoan optimization algorithm strategy to clean the database, the problems of privacy leakage and data loss in frequent item set mining are solved, and efficient privacy protection and accuracy of data mining results are achieved.

CN120744987AActive Publication Date: 2025-10-03NANJING UNIV OF INFORMATION SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511269877.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-03
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively hide sensitive information in frequent itemset mining, resulting in the risk of privacy leakage. At the same time, the hiding process may cause the destruction of the database structure and the loss of accuracy of data mining results.

Method used

By identifying and selecting appropriate target transactions and sensitive items, and adopting the protozoan optimization algorithm strategy, the original transaction database is cleaned, and only sensitive items are deleted without deleting the entire transaction, forming a target database to ensure the accuracy and privacy protection of data mining results.

Benefits of technology

It achieves the goal of hiding sensitive frequent patterns while maintaining the accuracy of data mining results and the integrity of the database structure to the greatest extent, reducing data loss and ensuring that 99.99% of the overall database structure is retained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744987A_ABST
    Figure CN120744987A_ABST
Patent Text Reader

Abstract

The invention provides a sensitive frequent information hiding method, device and system and a medium, and belongs to the technical field of privacy protection data mining. The sensitive frequent information hiding method comprises the steps that an original transaction database is combined, a frequent item set is mined according to a given support degree threshold value, and a sensitive frequent item set is generated by the frequent item set; according to the method, the target database is obtained through optimization algorithm data cleaning according to the data, and the hiding effect is achieved by reducing the support degree of the sensitive item set to be below a given threshold value, so that the target database balances data privacy and data utility under the condition of the same mining threshold value; in the cleaning process, data privacy leakage is reduced as much as possible, and normal effectiveness of data is reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, device, system and medium for hiding sensitive and frequent information, and belongs to the technical field of privacy protection data mining. Background Art

[0002] With the development of data science, data mining techniques have been widely applied in various fields to extract valuable information and patterns from massive amounts of data. In this context, frequent itemset mining (FMS), as a key task in data mining, has attracted widespread attention from researchers. Frequent itemset mining aims to discover itemsets that appear frequently in a dataset. These frequent itemsets often contain key patterns and sensitive associations implicit in the data, and their practical applications are dual-purpose: for example, in healthcare, they may reveal strong correlations between patient privacy and disease characteristics; in finance, they may reveal patterns in the association between customer identity information and transaction behavior; and in e-commerce, they may identify the unique consumption preferences of high-value users. However, with increasingly stringent data privacy regulations, how to effectively conceal sensitive information in frequent itemset mining has become a critical issue that needs to be addressed.

[0003] When processing large amounts of data, practitioners need to extract the necessary information from it, and data mining algorithms are constantly improving to meet this demand. However, the strong association rules involved in frequent itemset mining often contain sensitive information, and these highly supported association patterns are easily exploited for malicious purposes. In an era where data value is increasingly prominent, reverse engineering attacks targeting sensitive frequent patterns are becoming increasingly frequent, posing a serious risk of privacy breaches. Sensitive Frequent Itemset Hiding, as a core technology to address this challenge, has become a key research direction in privacy-preserving data mining in recent years.

[0004] Hiding sensitive frequent itemsets presents unique technical challenges. Numerous methods have been developed for hiding sensitive frequent itemsets in existing research, both domestically and internationally. Some of these approaches lean toward traditional rule-based processing, while others employ machine learning optimization algorithms to hide sensitive information. This research addresses the challenges of machine learning algorithms, such as significant database loss or insufficient sensitive information hiding.

[0005] When processing data sets, existing methods often require significant deletion of original data transactions to hide sensitive frequent itemsets, resulting in significant damage to the database structure. When the data distortion exceeds a critical value, a contradictory phenomenon of "hiding failure" and "true pattern loss" coexists. How to effectively hide sensitive frequent patterns while maximizing the accuracy of data mining results remains a technical bottleneck that urgently needs to be overcome. Summary of the Invention

[0006] The purpose of the present invention is to provide a sensitive frequent information hiding method, device, system and medium, which can achieve privacy protection by hiding sensitive frequent item sets by identifying and selecting appropriate target transactions (the set of transaction numbers to be deleted corresponding to the victim items) and sensitive items (victim items).

[0007] In order to achieve the above objectives / solve the above technical problems, the present invention is implemented by adopting the following technical solutions.

[0008] In one aspect, the present invention provides a method for hiding sensitive and frequent information, comprising:

[0009] Obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements;

[0010] According to whether the transactions in the original transaction database contain a victim item, the transactions that do not contain the victim item are directly input into the target database;

[0011] The transactions containing the victim items are cleansed and hidden, and the transactions after the victim items are hidden are input into the target database to form a final target database;

[0012] The method for obtaining the victim item specifically includes:

[0013] Calculate the candidate counts of all "1" itemsets in the sensitive frequent itemset;

[0014] Arrange the "1" item sets according to the candidate counts, and use the "1" item sets with the candidate counts from small to large as the victim items. Delete all sensitive frequent item sets containing the current victim items until the sensitive frequent item set is empty. Determine all victim items and sequence them in order.

[0015] The cleaning process specifically includes:

[0016] The serialized victim items are regarded as sub-particles, and the original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle.

[0017] Furthermore, obtaining the frequent itemsets in the original transaction database according to the support of the data information specifically includes:

[0018] The data information in the original transaction database is divided into frequent item sets and infrequent item sets according to the set support threshold.

[0019] Furthermore, the calculation of the candidate count of the "1" item set in the sensitive frequent item set specifically includes:

[0020] ;

[0021] in, and For custom parameters, is the candidate count, is the support degree of the “1” item set in the sensitive frequent item set, is the support of the “1” item set in the non-sensitive frequent item set.

[0022] The candidate count is used to quantify the degree of damage that may be caused to the database by modifying the sensitive item. The smaller the candidate count, the smaller the damage, and the larger the candidate count, the greater the damage.

[0023] Furthermore, the calculation method of the sub-particle size is expressed as follows:

[0024] ;

[0025] in, is the sub-particle size of the i-th victim term, is the scaling factor, represents the i-th victim item, is the pth sensitive frequent item set in S, S is the set of sensitive frequent item sets,

[0026] is the maximum support function of the frequent itemset containing the victim item,

[0027] represents the sensitive frequent itemset where the i-th victim item exists,

[0028] The minimum support threshold for frequent itemset mining,

[0029] |D| is the original transaction database size.

[0030] By designing the size of the sub-particles, we try to ensure that sensitive items are removed as much as possible, and also prevent data utility loss caused by excessive deletion. In general, n defaults to 1, ensuring that while the support of the sensitive item set is reduced to below the minimum threshold, it will not cause data utility loss due to excessive deletion. The larger n is, the better the confidentiality, but the greater the data utility loss. The smaller n is, the worse the confidentiality is, but the data utility is retained.

[0031] Furthermore, the original transaction database is cleaned according to the protozoan optimization algorithm strategy in combination with the size of each sub-particle to obtain a target database that hides sensitive and frequent information, specifically including:

[0032] The particle is composed of the size of each sub-particle and the pre-built sensitive transaction retrieval table:

[0033] The sensitive transaction retrieval table is a set of transaction numbers corresponding to the victim item in the original transaction database, the sub-particle is a victim item and a target transaction number, the target transaction number is a set of transaction numbers corresponding to the victim item in the original transaction database to be deleted, the number of transaction number sets is the size of the sub-particle corresponding to the victim item, and the particle is a set of all the sub-particles;

[0034] Calculate the fitness value of each particle and sort them. After sorting, the particle with the smallest fitness value is the victim item of the preliminary global optimal solution and the corresponding target transaction number;

[0035] A set proportion of particles are selected to enter the dormant or reproductive stage, and particles that do not enter the dormant or reproductive state will enter the foraging stage;

[0036] The expression for obtaining particles that have entered the dormant or reproductive stage is as follows:

[0037] ;

[0038] in, For particles that enter dormancy or reproduction, is the population size, To customize the preset maximum dormancy or reproduction ratio, is a random number between 0 and 1, is the population random processing function,

[0039] The fitness value of each particle that enters the dormant or reproduction stage is sorted to calculate the dormant parameter, and then the particle is determined to enter the dormant or reproduction stage based on the dormant parameter. The expression is:

[0040] ;

[0041] ;

[0042] in, is the sleep parameter, is the ranking of the fitness value of the particle entering the dormant or reproductive stage in the population;

[0043] Particles that enter the dormant stage will be deleted and replaced by new particles generated randomly. Particles that enter the reproduction stage will modify the initial sub-particles of the non-global optimal solution in the internal part.

[0044] The foraging parameters are calculated based on the random number and iteration number of the particles entering the foraging stage. The foraging parameters determine whether the particles enter the autotrophic or heterotrophic stage. The expression is:

[0045] ;

[0046] ;

[0047] in, is the foraging parameter, is the current iteration number, is the maximum number of iterations;

[0048] In the foraging phase, each particle replaces a set proportion of transactions of each of its sub-particles with new transactions; in the autotrophic phase, particles will select new transactions from the preliminary non-global optimal solution candidate set for replacement, and in the heterotrophic phase, particles will select new transactions from the preliminary global optimal solution for replacement;

[0049] Then calculate the fitness values ​​of all particles after replacement to obtain a new global optimal solution, update the global optimal solution, and replace transactions based on the new global optimal solution in the next iteration. Repeat the above replacement process until the maximum number of iterations is reached and the optimal solution is obtained to complete the cleaning of the original transaction database.

[0050] By classifying particles: foraging, dormancy, and selection at different stages of reproduction, iterative optimization is performed within the maximum number of iterations. Combined with the particle fitness value, the cleaning effect is improved, so that a particle with the optimal fitness value, that is, the best solution, is determined during the particle update process.

[0051] Furthermore, the fitness value is calculated as follows:

[0052] ;

[0053] in: is the fitness value of the i-th particle, , , is a custom parameter, and , is the total number of sensitive frequent itemsets that cannot be hidden during the hiding process of the i-th particle, is the total number of non-sensitive frequent itemsets lost by the i-th particle during the hiding process, is the total number of infrequent itemsets incorrectly mined by the i-th particle.

[0054] In a second aspect, the present invention provides a device for hiding sensitive and frequent information, comprising:

[0055] The acquisition module is used to obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements;

[0056] The processing module is used to directly input transactions that do not contain victim items into the target database based on whether the transactions in the original transaction database contain victim items; clean the transactions that contain victim items and input the transactions after hiding the victim items into the target database to form a final target database;

[0057] The method for obtaining the victim item specifically includes:

[0058] Calculate the candidate counts of all "1" itemsets in the sensitive frequent itemset;

[0059] Arrange the "1" item sets according to the candidate counts, and use the "1" item sets with the candidate counts from small to large as the victim items. Delete all sensitive frequent item sets containing the current victim items until the sensitive frequent item set is empty. Determine all victim items and sequence them in order.

[0060] The cleaning process specifically includes:

[0061] The serialized victim items are regarded as sub-particles, and the original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle.

[0062] In a third aspect, the present invention provides a system for hiding sensitive and frequent information, characterized by comprising:

[0063] Memory, used to store computer programs / instructions;

[0064] A processor is used to execute the computer program / instructions to implement the steps of the above-mentioned sensitive and frequent information hiding method.

[0065] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, characterized in that when the computer program / instruction is executed by a processor, the steps of the above-mentioned method for hiding sensitive and frequent information are implemented.

[0066] In a fifth aspect, the present invention provides a computer program product, comprising a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the above-mentioned method for hiding sensitive and frequent information are implemented.

[0067] Compared with the existing technology, the beneficial effects achieved by the present invention are: by cleaning the privacy database and inputting the transactions behind the hidden victim items into the target database, the present invention can achieve better privacy information protection of electronic health records; by reducing the support of sensitive item sets, when applying the same minimum frequent mining threshold, no false cost is generated.

[0068] By selecting a victim item with a smaller Icount value and combining it with an appropriate victim transaction obtained through iteration, the present invention causes less data loss (smaller loss cost value) when processing a data set. It also significantly reduces the problem of data loss in dense data sets and retains a certain degree of data utility.

[0069] Since the present invention does not delete the entire transaction, but chooses a cleaning strategy of deleting the victim item in the victim transaction, it achieves better results in preserving the item set and the overall structure of the database; it can ensure that the overall structure of the database is above 99.99%. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a schematic diagram of the process of the present invention;

[0071] Figure 2 Schematic diagram of the cleaning process of the present invention. DETAILED DESCRIPTION

[0072] It should be noted that:

[0073] Hiding failure: when the private itemset fails to be hidden under the same threshold parameter mining condition, that is, the frequent itemsets that should be hidden are still mined; Sensitive frequent item sets mined from the cleaned database are generally Represents the sensitive frequent itemsets that are not hidden during the hiding process ( ) total, the formula is as follows: ;

[0074] Loss cost: After cleaning, the non-sensitive private itemsets that should have been mined are not mined. NSRI' is the non-sensitive frequent itemsets mined from the cleaned database. Generally, let represents the non-sensitive frequent itemsets lost in the hiding process ( )total: ;

[0075] False cost: After cleaning, itemsets that should not be frequent itemsets are mined. This is partly due to false information generated by modifying the database. Generally, Indicates the total number of infrequent itemsets mined incorrectly: ;

[0076] Item set support: Item set I in the transaction database The support in It represents the item set Number of occurrences in the transaction database;

[0077] Itemset frequency: The frequency of item set I appearing in the transaction database is expressed as Indicates that Represents a database The total number of transactions, ;

[0078] Frequent itemsets: given parameters ( ), when an itemset The frequency in the database meets the conditions , then we can call the itemset Frequent itemsets ;

[0079] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0080] The term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.

[0081] Example 1

[0082] like Figure 1 In one embodiment shown, this embodiment provides a method for hiding sensitive and frequent information, including:

[0083] Step 1: Obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements;

[0084] Step 1.1: The original transaction database obtained in this embodiment is the electronic health record database.

[0085] , Represents the unique number of the transaction, where Representative transactions, each of which has . Represents the first itemsets; if an itemset contains , then this item set is called a “k” item set, as shown in Table 1.

[0086] Table 1 Electronic health record database

[0087] ;

[0088] As can be seen from the table above, it consists of 10 electronic health records;

[0089] Step 1.1: Set the frequent item set mining threshold (support threshold) to 0.4. Then, according to the frequent item set mining rules, the frequent item sets of this database are obtained. The results are shown in Table 2:

[0090] Table 2 Frequent itemsets

[0091] ;

[0092] According to the hidden demand principle, {a2}, {c1}, {c1, g1} are set as sensitive frequent itemsets.

[0093] Step 2: Based on whether the transactions in the original transaction database contain victim items, the transactions that do not contain victim items are directly input into the target database;

[0094] Step 3: Input the transaction containing the victim item into the privacy database, clean the privacy database, and input the transaction after hiding the victim item into the target database to form the final target database;

[0095] The method for obtaining the victim item specifically includes:

[0096] Step 3.1: Calculate the candidate count of the "1" item set in the sensitive frequent item set, specifically including:

[0097] ;

[0098] in, and For custom parameters, is the candidate count, is the support degree of the “1” item set in the sensitive frequent item set, is the support degree of the “1” item set in the non-sensitive frequent item set;

[0099] In this implementation ;

[0100] Step 3.2: Arrange the "1" item sets according to the candidate counts, and take the "1" item sets with the smallest candidate counts as the largest as the victim items. Delete all sensitive frequent item sets containing "1" item sets until the sensitive frequent item sets are empty, and determine the sizes of all victim items and their sub-particles: and , and number them in sequence according to the Icount value: c1, a2;

[0101] Step 3.3: Construct the sensitive item and victim item index table. The sensitive transaction retrieval table is the set of transaction numbers corresponding to the sensitive frequent item sets in the original transaction database, such as {a2}→{ , },{c1}→ .

[0102] like Figure 2 As shown, the cleaning process specifically includes:

[0103] Step 3.4: Use the sequenced victim items as sub-particles, and construct the particle by combining the size of each sub-particle;

[0104] The sub-particle size calculation method is expressed as:

[0105] ;

[0106] in, is the sub-particle size of the i-th victim term, is the scaling factor, Represents the current victim item, is the pth sensitive frequent item set in S, S is the set of sensitive frequent item sets,

[0107] is the maximum support function of the frequent itemset containing the victim item,

[0108] represents the sensitive frequent itemset where the i-th victim item exists,

[0109] The minimum support threshold for frequent itemset mining,

[0110] |D| is the original database size;

[0111] Wherein: the sub-particle is a victim item and a target transaction number, the target transaction number is the set of transaction numbers to be deleted corresponding to the victim item in the original transaction database, the number of the transaction number set is the size of the sub-particle corresponding to the victim item, and the particle is the set of all the sub-particles;

[0112] Step 3.5: Clean the original transaction database according to the protozoan optimization algorithm strategy (APO2DT algorithm strategy) based on the size of each sub-particle.

[0113] Step 3.51: The cleaning phase begins, and the initialized particles and their fitness values ​​are obtained according to step 3. Assume that the population size is 3, that is, it contains 3 particles P1, P2, and P3. The parameters are set to , , The fitness value is calculated according to the fitness function calculation formula:

[0114] ;

[0115] , ,

[0116] .

[0117] According to the fitness value of each particle, the transaction number of the particle with the smallest fitness value after sorting is the preliminary global optimal solution, as shown in Table 3:

[0118] Table 3 Initialized population representation

[0119] ;

[0120] Step 3.52:

[0121] A certain proportion of particles are selected to enter the dormant or reproductive stage, and the particles that do not enter the dormant or reproductive state will enter the foraging stage;

[0122] The expression for obtaining particles that have entered the dormant or reproductive stage is as follows:

[0123]

[0124] in, For particles that enter dormancy or reproduction, is the population size, To customize the preset maximum dormancy or reproduction ratio, is a random number between 0 and 1, is the population random processing function,

[0125] The fitness value of each particle that enters the dormant or reproduction stage is sorted to calculate the dormant parameter, and then the particle is determined to enter the dormant or reproduction stage based on the dormant parameter. The expression is:

[0126] ;

[0127] ;

[0128] in, is a random number between (0, 1), is the sleep parameter, is the ranking of the fitness value of particles entering the dormant or reproductive stage in the population, is the population size;

[0129] Particles that enter the dormant stage will be deleted and replaced by new particles generated randomly. Particles that enter the reproduction stage will modify the initial sub-particles of the non-global optimal solution in the internal part.

[0130] The foraging parameters are calculated based on the random number and iteration number of the particles entering the foraging stage. The foraging parameters determine whether the particles enter the autotrophic or heterotrophic stage. The expression is:

[0131] ;

[0132] ;

[0133] in, is a random number between (0, 1), is the foraging parameter, is the current iteration number, is the maximum number of iterations;

[0134] Select the global optimal solution. Since the fitness values ​​are the same, randomly select a particle P2 as the optimal solution. The three particles generate corresponding random parameters and proceed to the next stage. Here, according to the formula: Since the result is obtained randomly, it is assumed here that With P2. Generate random numbers with P2, assuming Random parameter is less than Parameters, so Entering the dormant stage. P2 random parameter is greater than Parameters, enter the reproduction stage. Other particles enter the autotrophic and heterotrophic selection stage. Similarly, generate random parameters for P3. Assume that the random parameters of P3 are less than Parameters, P3 enters heterotrophic operation.

[0135] According to the above, P1 performs dormancy operation, P2 performs reproduction operation, and P3 performs heterotrophic operation, so:

[0136] Table 4 Population representation after iteration

[0137] ;

[0138] It can be seen that the particles It has the lowest fitness value, so P3 is the optimal solution after iteration.

[0139] Step 3.53: Repeat the above process until the maximum iteration value is reached and the optimal result is obtained, such as the solution provided by particle P3: delete the sensitive item c1 from transactions T2, T7, and T8; delete the sensitive item a2 from transactions T3 and T10;

[0140] The original database was cleaned according to the optimal solution, and Table 5 was obtained:

[0141] Table 5 Target database after complete cleaning

[0142] ;

[0143] Step 3.54: Target database To verify, we perform frequent item set mining on the target database in Table 5, with the threshold also set to 0.4, and obtain Table 6:

[0144] Table 6 Frequent itemsets obtained from the target database

[0145] ;

[0146] As can be seen from the above table, the target database Sensitive information in the file has been hidden, while more non-sensitive information has been retained.

[0147] Since the present invention does not delete the entire transaction, but chooses a cleaning strategy of deleting the victim item in the victim transaction, it achieves better results in preserving the item set and the overall structure of the database; it can ensure that the overall structure of the database is above 99.99%.

[0148] Example 2

[0149] Based on the sensitive and frequent information hiding method described in Example 1, this embodiment provides a sensitive and frequent information hiding device, including:

[0150] The acquisition module is used to obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements;

[0151] The processing module is used to directly input transactions that do not contain victim items into the target database based on whether the transactions in the original transaction database contain victim items; clean the transactions that contain victim items and input the transactions after hiding the victim items into the target database to form a final target database;

[0152] The method for obtaining the victim item specifically includes:

[0153] Calculate the candidate counts of all "1" itemsets in the sensitive frequent itemset;

[0154] Arrange the "1" item sets according to the candidate counts, and use the "1" item sets with the candidate counts from small to large as the victim items. Delete all sensitive frequent item sets containing the current victim items until the sensitive frequent item set is empty. Determine all victim items and sequence them in order.

[0155] The cleaning process specifically includes:

[0156] The serialized victim items are regarded as sub-particles, and the original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle.

[0157] Example 3

[0158] Based on the sensitive and frequent information hiding method described in Example 1, this embodiment provides a sensitive and frequent information hiding system, including:

[0159] Memory, used to store computer programs / instructions;

[0160] A processor is configured to execute the computer program / instructions to implement the steps of the sensitive and frequent information hiding method described in Example 1.

[0161] Example 4

[0162] Based on the sensitive and frequent information hiding method described in Example 1, this embodiment provides a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the sensitive and frequent information hiding method described in Example 1 are implemented.

[0163] Example 5

[0164] Based on the sensitive and frequent information hiding method described in Example 1, this embodiment provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the sensitive and frequent information hiding method described in Example 1 are implemented.

[0165] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0166] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0167] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0169] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.

Claims

1. A method for hiding sensitive and frequent information, characterized in that: include: Obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements; According to whether the transactions in the original transaction database contain a victim item, the transactions that do not contain the victim item are directly input into the target database; The transactions containing the victim items are cleansed and hidden, and the transactions after the victim items are hidden are input into the target database to form a final target database; The method for obtaining the victim item specifically includes: Calculate the candidate counts of all "1" itemsets in the sensitive frequent itemset; Arrange the "1" item sets according to the candidate counts, and use the "1" item sets with the candidate counts from small to large as the victim items. Delete all sensitive frequent item sets containing the current victim item until the sensitive frequent item set is empty. Determine all the victim items and sequence them in order. The cleaning process specifically includes: The serialized victim items are regarded as sub-particles respectively. The original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle, and the victim items are hidden.

2. The sensitive and frequent information hiding method according to claim 1 is characterized in that: The step of obtaining the frequent itemsets in the original transaction database according to the support of the data information specifically includes: The data information in the original transaction database is divided into frequent item sets and infrequent item sets according to the set support threshold.

3. The sensitive and frequent information hiding method according to claim 1 is characterized in that: The calculation of the candidate count of the "1" item set in the sensitive frequent item set specifically includes: ; in, and For custom parameters, is the candidate count, is the support degree of the "1" item set in the sensitive frequent item set, is the support of the "1" item set in the non-sensitive frequent item set.

4. The sensitive and frequent information hiding method according to claim 1, characterized in that: The calculation method of the sub-particle size is expressed as follows: ; in, is the sub-particle size of the i-th victim term, is the scaling factor, represents the i-th victim item, is the pth sensitive frequent item set in S, S is the set of sensitive frequent item sets, is the maximum support function of the frequent itemset containing the victim item, represents the sensitive frequent itemset where the i-th victim item exists, The minimum support threshold for frequent itemset mining, |D| is the original transaction database size.

5. The sensitive and frequent information hiding method according to claim 4 is characterized in that: The serialized victim items are sub-particles, and the original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle to hide the victim items, specifically including: The particle is composed of the size of each sub-particle and the pre-built sensitive transaction retrieval table: The sensitive transaction retrieval table is a set of transaction numbers corresponding to the victim item in the original transaction database, the sub-particle is a victim item and a target transaction number, the target transaction number is a set of transaction numbers corresponding to the victim item in the original transaction database to be deleted, the number of transaction number sets is the size of the sub-particle corresponding to the victim item, and the particle is a set of all the sub-particles; Calculate the fitness value of each particle and sort them. After sorting, the particle with the smallest fitness value is the victim item of the preliminary global optimal solution and the corresponding target transaction number; A set proportion of particles are selected to enter the dormant or reproductive stage, and particles that do not enter the dormant or reproductive state will enter the foraging stage; The expression for obtaining particles that have entered the dormant or reproductive stage is as follows: ; in, For particles that enter dormancy or reproduction, is the population size, To customize the preset maximum dormancy or reproduction ratio, is a random number between 0 and 1, is the population random processing function; The fitness value of each particle that enters the dormant or reproduction stage is sorted to calculate the dormant parameter, and then the particle is determined to enter the dormant or reproduction stage based on the dormant parameter. The expression is: ; ; in, is the sleep parameter, is the ranking of the fitness value of the particle entering the dormant or reproductive stage in the population; Particles that enter the dormant stage will be deleted and replaced by new particles that are randomly generated. Particles that enter the reproduction stage will modify the initial sub-particles of the non-global optimal solution. The foraging parameters are calculated based on the random number and iteration number of the particles entering the foraging stage. The foraging parameters determine whether the particles enter the autotrophic or heterotrophic stage. The expression is: ; ; in, is the foraging parameter, is the current iteration number, is the maximum number of iterations; In the foraging phase, each particle replaces a set proportion of transactions of each of its sub-particles with new transactions; in the autotrophic phase, particles will select new transactions from the preliminary non-global optimal solution candidate set for replacement, and in the heterotrophic phase, particles will select new transactions from the preliminary global optimal solution for replacement; Then calculate the fitness values ​​of all particles after replacement to obtain a new global optimal solution. Update the global optimal solution, and the next iteration will replace transactions based on the new global optimal solution. Repeat the above replacement process until the maximum number of iterations is reached. The optimal solution is obtained, the victim item is hidden, and the original transaction database is cleaned.

6. The sensitive and frequent information hiding method according to claim 5, characterized in that: The fitness value calculation expression is: ; in: is the fitness value of the i-th particle, , , is a custom parameter, and , is the total number of sensitive frequent itemsets that cannot be hidden during the hiding process of the i-th particle, is the total number of non-sensitive frequent itemsets lost by the i-th particle during the hiding process, is the total number of infrequent itemsets incorrectly mined by the i-th particle.

7. A device for hiding sensitive and frequent information, characterized in that: include: The acquisition module is used to obtain the frequent itemsets of the transaction according to the support of the transaction in the original transaction database, and divide the frequent itemsets into sensitive frequent itemsets containing victim items and non-sensitive frequent itemsets according to the set requirements; The processing module is used to directly input transactions that do not contain victim items into the target database based on whether the transactions in the original transaction database contain victim items; clean the transactions that contain victim items and input the transactions after hiding the victim items into the target database to form a final target database; The method for obtaining the victim item specifically includes: Calculate the candidate counts of all "1" itemsets in the sensitive frequent itemset; Arrange the "1" item sets according to the candidate counts, and use the "1" item sets with the candidate counts from small to large as the victim items. Delete all sensitive frequent item sets containing the current victim item until the sensitive frequent item set is empty. Determine all the victim items and sequence them in order. The cleaning process specifically includes: The serialized victim items are regarded as sub-particles, and the original transaction database is cleaned according to the protozoan optimization algorithm strategy based on the size of each sub-particle.

8. A sensitive and frequent information hiding system, characterized in that: include: Memory, used to store computer programs / instructions; A processor is configured to execute the computer program / instructions to implement the steps of the sensitive and frequent information hiding method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the method for hiding sensitive and frequent information described in any one of claims 1-6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method for hiding sensitive and frequent information according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • High-utility hiding protection method of sensitive information data

    CN105138926A

  • Self-adaptive differential privacy method for asset positioning data privacy protection

    CN118940315A

  • E-commerce transaction data cleaning method and system

    CN119648230A

  • Mining method and system for privacy protection periodic high-utility item set

    CN120144638A

  • Distortion apparatus for hiding sensitive association rules and method thereof

    KR1020170049048A