A method for generating negative database

By combining binary XOR and random algorithms to generate negative databases, the existing technology has solved the shortcomings in security and accuracy, and achieved negative database generation with high security and high computing accuracy.

CN114547694BActive Publication Date: 2025-05-23WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210188158.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-23
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The existing negative database generation algorithm has insufficient security, is susceptible to attack or decreases in accuracy, and similar original data can be recovered through probability statistics.

Method used

A binary XOR combined with a random algorithm is used to generate n position pairs (p, q), calculate the XOR value and zeroList, and randomly combine zeroList and oneList to generate a negative database to avoid using probability parameters.

Benefits of technology

It significantly improves the security of negative databases, enhances difficulty, and has high calculation accuracy, and can effectively resist attacks from probability statistical models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547694B_ABST
    Figure CN114547694B_ABST
Patent Text Reader

Abstract

The present invention discloses a negative database generation method, comprising: converting original decimal data into binary data; generating n position pairs (p, q); calculating the XOR value of the corresponding position according to the position pair and the hidden string, and then determining the value range of the zeroList of the corresponding position according to the XOR result; calculating oneList according to zeroList, and then generating a negative database according to zeroList and oneList. The method effectively ensures the security of the negative database; the generation probability of different types of records is not affected by parameters, so that the records in the generated negative database are more evenly distributed; the negative database generated by the method has higher calculation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of privacy protection and security, and specifically relates to a negative database generation method. Background Art

[0002] With the rapid development of Internet technology, we enjoy the convenience brought by the Internet in work, study, life, etc., and deal with social software, shopping software, and food delivery software every day. While enjoying this convenience, our personal information and private data are leaked unknowingly. In recent years, privacy data leakage and data security issues have caused damage to the interests of individuals or groups. Therefore, how to ensure personal privacy and data security has become a focus issue.

[0003] Negative representation of information is a data representation method that can effectively ensure data security. It mainly achieves the purpose of protecting privacy data and information security by storing information that is the complement of the original data. Negative database is a negative information representation scheme. Negative database (NDB) can effectively reduce storage space and significantly improve data security by compressing the complement of the original data. Therefore, negative database technology has been used in privacy protection, biometric information recognition, information hiding and other aspects.

[0004] At present, the work of negative database is mainly aimed at binary data. How to convert decimal data into binary data and then into negative database data has become a key research direction. Many negative database generation schemes have been proposed, and these schemes have their own advantages and disadvantages. The negative databases generated by the common prefix algorithm and RNDB algorithm have low security and are easy to be attacked; the negative databases generated by the q-hidden algorithm and the p-hidden algorithm generate more difficult negative databases, but the accuracy decreases to a certain extent when the negative database participates in the calculation; the K-hidden algorithm and QK-hidden have higher application effects, but similar original data can still be obtained through probability statistics, and the security is reduced to a certain extent. Therefore, it is very necessary to propose a negative database generation algorithm with higher security. Summary of the invention

[0005] In order to overcome the defects of the above-mentioned background technology, the present invention provides a method for generating a negative database to improve the security of the negative database.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] A method for generating a negative database, comprising:

[0008] Step 1, convert the original decimal data into binary data;

[0009] Step 2, generate n position pairs (p, q);

[0010] Step 3, calculate the XOR value of the corresponding position according to the position pair and the hidden string, and then determine the range of the value of zeroList of the corresponding position according to the XOR result;

[0011] Step 4, calculate oneList based on zeroList, and then generate a negative database based on zeroList and oneList;

[0012] oneList is an integer linked list of the number of times each bit of the original data appears as '1' in the negative database, and zeroList is an integer linked list of the number of times each bit of the original data appears as '0' in the negative database.

[0013] Preferably, step 1 specifically includes: first converting the data in the data set into integer data, and then converting the integer data into binary data.

[0014] Preferably, step 2 specifically includes:

[0015] Manually set the initial value of n, generate position pairs (p, q) one by one, and for the i-th position pair (p i ,q i ), first, randomly generate the values ​​of p and q. If p=q, p=i or q=i appears, regenerate the values ​​of P and q until p≠q≠i is used as (p i ,q i ), repeat this process until n position pairs are generated;

[0016] Where i is the number of the position pair, 1≤i≤n; p i is the p-value of the ith position pair; q i is the q value of the i-th position pair, and the i-th position pair is denoted as (p i ,q i ), p and q are two random values, and p, q∈[0, L) represents the position subscript of the binary string s.

[0017] Preferably, step 3 specifically includes:

[0018] like s i = 0, then zeroList[i] = rnd, rnd∈([m / 2], m), if s i = 0, then zeroList[i] = rnd, rnd∈[1,[m / 2]]; if si=1, then zeroList[i]=rnd, rnd∈[1,[m / 2]], if si =1, then zeroList[i]=rnd, rnd∈([m / 2], m);

[0019] Where L is the length of the binary string s, where s is a piece of binary data generated in step 1; s i It represents the value of the i-th bit of s. Indicates that s is in p i The value of the position; zeroList is an integer linked list of length L; zeroList[i] is the i-th position of s, and the number of times it appears in the negative database is represented by '0'; rnd is a random number; m is the number of times each position of s appears in the negative database; [m / 2] is rounded down.

[0020] Preferably, step 4 specifically includes: oneList[i]=t-zeroList[i], wherein t is an initialization value of the total number of times each bit in the binary data s appears in the negative database, which is manually set; and zeroList and oneList are randomly combined into a negative database record.

[0021] Preferably, the method of randomly combining zeroList and oneList into a negative database includes: randomly selecting k different positions from zeroList and oneList in turn, and using '0' or '1' corresponding to the k positions as the confirmation position of a negative database record of k-NDB.

[0022] The present invention designs a negative database generation method. Experimental results show that the negative database generated by this method is complete and difficult to solve; under the premise of ensuring the experimental accuracy, compared with the previous negative database generation algorithm, the security of the negative database can be significantly improved, and the problem that the previous negative database generation algorithm may be broken by the probability statistical model can be effectively solved; good experimental results have been achieved in Kmeans clustering and KNN classification experiments. The previous negative database generation algorithm of the present invention usually adopts the adjustment of multiple different probability parameters to control the generation probability of different types of records to control their distribution. However, this method does not adopt probability parameters, which effectively improves the security of the negative database.

[0023] The present invention adopts a scheme combining binary XOR with a random algorithm, which effectively ensures the security of the negative database; the generation probability of different types of records is not affected by parameters, so the records in the generated negative database are more evenly distributed; the negative database generated by the method has higher calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 The figure is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0026] Step 1: First convert the original decimal data into binary data.

[0027] Most of the existing negative database research is based on binary data. Therefore, we must first convert the decimal data into binary data. If the decimal data is a floating point number, we must first convert the floating point number into an integer.

[0028] Here we take the commonly used Iris dataset as an example. The Iris dataset contains 150 data samples, each with 4 attributes. The data in the Iris dataset are all floating point numbers. For floating point numbers, you must first convert the floating point numbers into integer data. The data in the Iris dataset has only one decimal place, so you only need to multiply by 10 to convert it into integer data.

[0029] Next, the integer data is converted into binary data. The traditional base conversion is not used here. Instead, the number of times '1' appears in the binary data is used to represent the original integer data. For the Iris data set, each data is less than 100 after conversion to integer, so each data is converted into a binary string with a length of 100. For example, '6' is converted to "111111" and 94 '0's are added in front to ensure the consistency of the binary data length. Each sample data of the Iris data set has 4 attributes, so it is converted into 4 binary strings with a length of 100. The 4 binary strings are sequentially concatenated to obtain binary data with a length of 400.

[0030] After the above processing, we convert the original 150 data samples into 150 binary data with a length of 400. This processing effectively ensures the retention of valid data information.

[0031] Step 2: Randomly generate n position pairs (p, q).

[0032] The initial value of n is set to 400. Here, position pairs are generated one by one. For the i-th position pair (p i , qx), first, randomly generate the values ​​of p and q. If p=q, p=i or q=i appears, regenerate the values ​​of p and q until p≠q≠i is used as (p i ,q i ). This process is repeated until 400 position pairs are generated.

[0033] Where i is the number of the position pair, 1≤i≤400; p i is the p-value of the ith position pair; q i is the q value of the i-th position pair; the i-th position pair is denoted as (pi ,q i ).

[0034] The value of n affects the security of the generated negative data and the accuracy of the calculation on the negative database. Here, the value of n is the length of the binary data record, which can provide higher calculation accuracy while ensuring security. The randomly generated values ​​of p and q provide another layer of protection to ensure the security of the negative database generation algorithm.

[0035] Step 3: Calculate the XOR value of the corresponding position based on the position pair and the hidden string, and then determine the range of the value of zeroList at the corresponding position based on the XOR result.

[0036] Take a binary string s of length L as an example, and generate a zeroList of length L. The generation rule is: if s=0, then zeroList[i]=rnd, rnd∈([m / 2],m), if si=0, then zeroList[i]=rnd, rnd∈[1,[m / 2]]; if si=1, then zeroList[i]=rnd, rnd∈[1,[m / 2]], if s i =1, then zeroList[i]=rnd, rnd∈([m / 2], m).

[0037] Where s is a binary data of Iris generated in step 1; si represents the value of the i-th bit of s, Indicates that s is in p i The value of the position; zeroList is an integer linked list of length L; zeroList[i] is the i-th position of s, and the number of times it appears in the negative database is represented by '0'; rnd is a random number; m is the number of times each position of s appears in the negative database; [m / 2] specifies rounding down.

[0038] For attackers, Equivalent to It can be seen as adding four negative database records; Equivalent to , which can also be regarded as adding four negative database records. Therefore, it is very difficult to restore the original data from the negative database generated by this method. This step innovatively adopts binary XOR combined with random algorithm, which effectively ensures the difficulty of negative database relative to local search strategy.

[0039] Step 4: Based on zeroList, calculate oneList, and then generate a complete negative database based on zeroList and oneList.

[0040] The total number of times each bit of binary data s appears in the negative database is fixed, so oneList can be calculated through zeroList. Assuming that the number of times the i-th bit of s appears in the negative database is t, then oneList[i] = t-zeroList[i]. Next, zeroList and oneList are randomly combined into negative database records. First, a bit is randomly selected from zeroList or oneList as the first determined bit of the negative database record, and then a bit is randomly selected from zeroList or oneList as the second determined bit of the negative database record. And so on, finally k bits are selected as a negative database record of k-NDB.

[0041] It is stipulated that the number of times the i-th bit of s appears in the negative database is 15. If zeroList[i]=3, then oneList[i]=15-zeroList[i]=12. Generate 3-NDB, then each record in the negative database contains 3 definite bits. NDBs is first initialized to an empty set. The number of data that NDBs is expected to contain is the length of s multiplied by the parameter r. Each time, 3 different positions of data are randomly selected from zeroList and oneList as a negative database record until all data in zeroList and oneList appear in the negative database. In this way, a complete negative database of s is generated.

[0042] Previous negative database generation algorithms usually use multiple different probability parameters to control the generation probability of different types of records to control their distribution. However, this method does not use probability parameters, which effectively improves the security of the negative database.

[0043] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A method for generating a negative database, It is characterized in that include: Step 1, convert the original decimal data into binary data; Step 2, generate n position pairs (p, q); Step 3, calculate the XOR value of the corresponding position according to the position pair and the hidden string, and then determine the range of the value of zeroList of the corresponding position according to the XOR result; Step 4, calculate oneList based on zeroList, and then generate a negative database based on zeroList and oneList; oneList is an integer linked list of the number of times each bit of the original data appears as '1' in the negative database, and zeroList is an integer linked list of the number of times each bit of the original data appears as '0' in the negative database; The step 2 specifically includes: Manually set the initial value of n, generate position pairs (p, q) one by one, and for the i-th position pair (p i ,q i ), first, randomly generate the values ​​of p and q. If p=q, p=i or q=i appears, regenerate the values ​​of p and q until p≠q≠i is used as (p i ,q i ), repeat this process until n position pairs are generated; Where i is the number of the position pair, 1≤i≤n; p i is the p-value of the ith position pair; q i is the q value of the i-th position pair, and the i-th position pair is denoted as (p i ,q i ), p and q are two random values, p, q∈[0,L) represents the position subscript of the binary string s; The step 4 specifically includes: oneList[i]=t-zeroList[i], where t is an initialization value of the total number of times each bit in the binary data s appears in the negative database, which is manually set; and zeroList and oneList are randomly combined into a negative database record.

2. A method for generating a negative database according to claim 1, It is characterized in that The step 1 specifically includes: firstly converting the data in the data set into integer data, and then converting the integer data into binary data.

3. A negative database generation method according to claim 1, It is characterized in that The step 3 specifically includes: like s=0, then zeroList[i]=rnd,rnd∈([m / 2],m), if s i =0, then zeroList[i]=rnd,rnd∈[1,[m / 2]]; if s i =1, then zeroList[i]=rnd,rnd∈[1,[m / 2]], if s i =1, then zeroList[i]=rnd,rnd∈([m / 2],m); Where L is the length of the binary string s, where s is a piece of binary data of Iris generated in step 1; s i It represents the value of the i-th bit of s. Indicates that s is in p i The value of the position; zeroList is an integer linked list of length L; zeroList[i] is the i-th position of s, and the number of times it appears in the negative database is represented by '0'; rnd is a random number; m is the number of times each position of s appears in the negative database; [m / 2] is rounded down.

4. A method for generating a negative database according to claim 1, It is characterized in that The method of randomly combining zeroList and oneList into a negative database includes: randomly selecting k different positions from zeroList and oneList in turn, and using '0' or '1' corresponding to the k positions as the confirmation position of a negative database record of k-NDB.

Citation Information

Patent Citations

  • Privacy protection k-means clustering method

    CN108154185A

  • Secure search processing system and secure search processing method

    US20140172830A1