A method for judging the legitimacy of data transactions based on a counting Bloom filter
By using counting Bloom filters in data transactions to map the feature vectors of the data set into a feature matrix, the problems of inefficient retrieval efficiency and data privacy threats in the prior art are solved, and fast and effective judgment of the legality of data sets and privacy protection are achieved.
Patent Information
- Application Number
- CN202210902991.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The prior art has problems of inefficient retrieval efficiency and data privacy threats in data transactions, and it is difficult to quickly and effectively judge the legitimacy of data sets.
Using a method based on a counting Bloom filter, the feature matrix is formed by mapping the eigenvectors of the data set into a counting Bloom filter, and the legality of the data set is judged by using a maximum likelihood estimation method.
It improves the efficiency of judging the legality of data sets in data transactions, and protects the characteristic space privacy of data sets to a certain extent.
Smart Images

Figure QLYQS_29
Abstract
Description
Technical Field
[0001] The present invention relates to a method for calculating and judging the similarity degree of data sets, and belongs to the fields of associated data retrieval and data transaction tracking. Background Art
[0002] Data has huge economic value and is regarded as emerging oil resources. The breakthrough progress of machine learning algorithms and the wide application of artificial intelligence technology also rely on the supply of a large amount of high-quality data. Each organization hopes to obtain valuable data resources to optimize performance and assist decision-making. As an emerging business model, data sharing and trading have received high attention from the business community and academia. In data transactions, there are many illegal profit-making behaviors among transaction participants. The most common one is to repeatedly sell or sell data sets after infringing modifications to profit based on the characteristics of easy replication and modification of data. This illegal behavior will cause serious damage to the good ecosystem of data sharing and trading.
[0003] The data set similarity degree calculation method is based on many existing data individual similarity degree calculation methods. Many existing data individual calculation methods map data individuals (pictures, texts) to a feature metric space of a certain dimension, and judge the similarity degree of data individuals by calculating the distance of data individuals in the feature metric space. Based on the data individual similarity degree calculation method, many existing methods attempt to use the distance in the feature space to retrieve the distances of similar individuals one by one in a data set, such as tree-based methods. However, the retrieval efficiency of these methods is limited because they need to calculate the most similar / approximate most similar individuals. At the same time, the method of uploading and reviewing all data in the data set is an infringement of the privacy of the data set, and the retrieval method does not consider this privacy issue. Summary of the Invention
[0004] The present invention is to solve the deficiencies of the above-mentioned existing technology in terms of retrieval efficiency and data privacy threats, and proposes a method for judging the legality of data transactions based on a counting Bloom filter, in order to help reviewers on the data transaction platform side in data sharing and trading quickly judge the legality of data sets participating in data transactions, so as to maintain a good data sharing and trading ecosystem.
[0005] In order to achieve the above invention purpose, the present invention adopts the following technical solutions:
[0006] The method for judging the legality of data transactions based on a counting Bloom filter according to the present invention is characterized in that it is applied to a data set transaction environment composed of a data transaction platform and M data set providers. Among them, the kth data set provided by the kth data set provider is denoted as i k,j indicating the jth data in the kth data set, mk Denote the data volume of the k-th data set, and the judgment method is carried out according to the following steps:
[0007] Step 1: If k = 1, then the data trading platform defaults that the k-th data set of the k-th data provider is a legal data set; and execute Step 2; otherwise, execute Step 4;
[0008] Step 2: The k-th data provider judges the j-th data i in the k-th data set k,j If the j-th data i k,j is text data, then the j-th data i k,j is extracted as a measurable vector in the feature space through the N-Gram method and the minhash method; if the j-th data i k,j is image data, then the j-th data i k,j is extracted as a measurable vector in the feature space through a similarity neural network; thus, the feature vector set of the k-th data set I k is obtained;
[0009] The k-th data provider maps its feature vector set into l different counting Bloom filters according to l different mappings provided by the data trading platform, thereby forming the feature matrix of the k-th data set I k and uploads it to the data trading platform; and the width of the feature matrix is the width w of the counting Bloom filter, and the length of the feature matrix is l;
[0010] Step 3: After assigning k + 1 to k, return to Step 1 until k > M;
[0011] Step 4: The k-th data provider splits the k-th data set into m k sets, and each set has only one data;
[0012] The k-th data provider calculates the feature matrix of each set according to the process of Step 2, thereby obtaining m k feature matrices and uploading them to the data trading platform to judge whether the data set is legal;
[0013] Step 5: Let the data set to be queried be the legal data sets in the 1st to k - 1st data sets;
[0014] After the data trading platform receives the m k feature matrices uploaded by the k-th data set provider, let event X be the uploaded m kThe natural numbers in the same grid of the counting Bloom filter in a row vector of a feature matrix and the feature matrix of the dataset to be queried are s; then the occurrence probability Probability(X) of event X is obtained using Equation (1):
[0015]
[0016] In Equation (1), t is the number of data determined to be illegal for the j-th data i in the dataset to be queried k,j ; ∈1 and ∈2 are two thresholds set by the data trading platform; r represents a loop variable; is the r-th power of ∈1, is the (s - r)-th power of ∈2; is the combination number;
[0017] Step 6: The data trading platform estimates m at two positions of t = 0 and t = 1 using the estimated likelihood function in the maximum likelihood estimation method k for the occurrence probability of event X for each feature matrix in the m feature matrices and the feature matrix of the dataset to be queried; if the likelihood value at t = 1 is higher than the likelihood value at t = 0, it is determined that the data corresponding to the corresponding feature matrix is illegal, otherwise it means that the data corresponding to the corresponding feature matrix is legal;
[0018] Step 7: After the data trading platform estimates each data in the k-th data set I k , it counts the legal proportion of the k-th data set I k . If the legal proportion is within the allowable range, the data trading platform determines that the k-th data set I k is legal data and uploads its feature matrix according to the process in Step 2; otherwise, it means that the k-th data set I k is illegal data, and the data trading platform rejects the k-th data set I k from participating in data trading.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0020] The present invention fully utilizes the principle of judging illegal data in data sharing and trading, that is, for individual data, the legality of data individuals is judged with a certain similarity threshold. After expanding to the scale of data sets, compared with many existing methods that still perform distance calculations, the efficiency of the present invention is further improved. While the judgment speed is increased on the one hand, the present invention also protects the feature space privacy of data sets. This method converts the content of the feature space of data sets into the content of Bloom filter matrices, thereby protecting the privacy of data in data sets to a certain extent. Specific embodiments
[0021] In this embodiment, a method for judging the legality of data transactions based on a counting Bloom filter is applied to a data set trading environment composed of a data trading platform and M data set providers. Among them, the k-th data set provided by the k-th data set provider is denoted as i k,j representing the j-th data in the k-th data set, and m k representing the data volume of the k-th data set. This judgment method is carried out according to the following steps:
[0022] Step 1: If k = 1, the data trading platform defaults that the k-th data set of the k-th data provider is a legal data set; and execute Step 2; otherwise, execute Step 4;
[0023] Step 2: The k-th data provider judges the j-th data i in the k-th data set k,j . If the j-th data i k,j is text data, the j-th data i k,j is extracted into a measurable vector in the feature space by the N-Gram method and the minhash method; if the j-th data i k,j is image data, the j-th data i k,j is extracted into a measurable vector in the feature space by a similarity neural network; thus, the feature vector set of the k-th data set I k is obtained;
[0024] The k-th data provider maps its feature vector set into l counting Bloom filters according to l different mappings provided by the data trading platform, thereby forming the feature matrix of the k-th data set I k and uploading it to the data trading platform; and the width of the feature matrix is the width w of the counting Bloom filter. In this embodiment, the width w is determined by the locality-sensitive hashing function LSH(x) used by the Bloom filter, and it is the same cluster of locality-sensitive hashing functions provided by the data platform, so the width is consistent. The length of the feature matrix is l. The influence of the matrix length l value is that the larger the l value, the more accurate the estimated result of whether the data is "illegal" is, and at the same time, the computational efficiency complexity will be higher. The length l value is an empirical value determined according to the scale of the data and the required time of the system. Generally, the number of Bloom filters is set to 128 or 256; each value of the feature matrix is a natural number, and the sum of each row of the matrix is m k ;
[0025] Step 3: After assigning k + 1 to k, return to Step 1 until k > M;
[0026] Step 4: The k-th data provider splits the k-th data set into mk a set, and each set has only one piece of data;
[0027] The k-th data provider calculates the feature matrix of each set according to the process in Step 2, so as to obtain m k feature matrices and uploads them to the data trading platform to determine whether the data set is legal;
[0028] Step 5: Let the data set to be queried be the legal data sets in the 1st to (k - 1)-th data sets;
[0029] After the data trading platform receives the m k feature matrices uploaded by the k-th data set provider, let event X be that the natural numbers in the same grid of the counting Bloom filter in a certain row vector of the uploaded m k feature matrices and the feature matrix of the data set to be queried are s; then the occurrence probability Probability(X) of event X is obtained by using Equation (1):
[0030]
[0031] In Equation (1), t is the number of data determined to be illegal data for the j-th data i k,j in the data set to be queried, which is the variable to be estimated in the probability theory method; ∈1 and ∈2 are two thresholds set by the data trading platform, and the selection of the thresholds comes from the locality-sensitive hashing function LSH;
[0032] dis(x,y)≤t * , Probability(LSH(x)=LSH(y))≥∈1 (2)
[0033] dis(x,y)≥t * , Probability(LSH(x)=LSH(y))≤∈2 (3)
[0034] In Equations (2) and (3), for the feature vectors x and y of any two data in the M data sets, dist(x,y) represents the distance between the feature vectors x and y, and t * is the distance threshold that the data trading platform believes is definitely illegal data, and t * is the distance threshold that the data trading platform believes is definitely not illegal data; r represents the loop variable; is the r-th power of ∈1, is the (s - r)-th power of ∈2; is the combination number;
[0035] Step 6: The data trading platform estimates m at two positions of t = 0 and t = 1 by using the estimated likelihood function in the maximum likelihood estimation methodk The occurrence probability of each feature matrix in the feature matrix for event X with the feature matrix of the dataset to be queried; if the likelihood value when t = 1 is higher than the likelihood value at t = 0, it is determined that the data corresponding to the corresponding feature matrix is illegal, otherwise it means that the data corresponding to the corresponding feature matrix is legal; note that some calculations in step 6 can be pre-calculated here, and only the values of two points need to be estimated in this step, so the calculation scale is not large.
[0036] Step 7: After the data trading platform estimates each data in the k-th data set I k it counts the legal ratio of the k-th data set I k If the legal ratio is within the allowed range, the data trading platform determines that the k-th data set I k is legal data and uploads its feature matrix according to the process of step 2; otherwise, it means that the k-th data set I k is illegal data, and the data trading platform rejects the k-th data set I k from participating in data trading.
Claims
1. A method for judging the legality of data transactions based on a counting Bloom filter, characterized in that Applied to a data set trading environment composed of a data trading platform and M data set providers, where the th data set provider provides the th data set, denoted as , represents the th data in the th data set, represents the data volume of the th data set, and the judgment method is carried out according to the following steps: Step 1. If k = 1, the data trading platform defaults the k-th data set of the k-th data provider as the legal data set and proceeds to Step 2; otherwise, proceeds to Step 4; Step 2: The k-th data provider makes a judgment on the k-th data set in the th data . If the th data is text data, the th data is extracted as a measurable vector in the feature space through the N-Gram method and the minhash method; if the th data is image data, the th data is extracted as a measurable vector in the feature space through the graph similarity neural network; thus, the feature vector set of the k-th data set is obtained. The k-th data provider maps its set of feature vectors into l count Bloom filters according to l different mappings provided by the data trading platform, thereby forming the feature matrix of the k-th data set and uploading it to the data trading platform; and the width of the feature matrix is the width w of the count Bloom filter, and the length of the feature matrix is ; and ; Step 3. After assigning k + 1 to k, return to Step 1 until k > M; Step 4. The k-th data provider splits the k-th data set into sets, and each set has only one piece of data; The k-th data provider calculates the feature matrix of each set according to the process of step 2, so as to obtain feature matrices and upload them to the data trading platform to determine whether the data set is legal; Step 5. Let the data sets to be queried be the legal data sets in the 1st to the (k - 1)-th data sets; After the data trading platform receives the feature matrices uploaded by the k-th dataset provider, let event X be that the natural numbers in the same cell of the counting Bloom filter in a certain row vector of the feature matrices uploaded and the feature matrix of the dataset to be queried are s; then the occurrence probability of event X is obtained using Equation (1) After the data trading platform receives the feature matrices uploaded by the k-th dataset provider, let event X be that the natural numbers in the same cell of the counting Bloom filter in a certain row vector of the feature matrices uploaded and the feature matrix of the dataset to be queried are s; then the occurrence probability of event X is obtained using Equation (1) After the data trading platform receives the feature matrices uploaded by the k-th dataset provider, let event X be that the natural numbers in the same cell of the counting Bloom filter in a certain row vector of the feature matrices uploaded and the feature matrix of the dataset to be queried are s; then the occurrence probability of event X is obtained using Equation (1) : (1) In formula (1), t is the number of data determined to be illegal data among the data in the queried dataset; are two thresholds set by the data trading platform; r represents the loop variable; is to the r-th power of is to the (s - r)-th power of is the combination number; Step 6. The data trading platform uses the estimated likelihood function in the maximum likelihood estimation method to estimate the probability of the occurrence of event X for each of the feature matrices at two positions and the feature matrix of the dataset to be queried; if the likelihood value of is higher than the likelihood value at then it is determined that the data corresponding to the corresponding feature matrix is illegal, otherwise it means that the data corresponding to the corresponding feature matrix is legal; Step 7: After the data trading platform estimates each data in the k-th data set , it counts the legal ratio of the k-th data set . If the legal ratio is within the allowed range, the data trading platform determines that the k-th data set is legal data and uploads its feature matrix according to the process in Step 2; otherwise, it indicates that the k-th data set is illegal data, and the data trading platform rejects the k-th data set from participating in data trading.
Citation Information
Patent Citations
Data encryption and search method capable of protecting file privacy in cloud environment
CN108768951A
Attribute hiding and cancelling method based on counting Bloom filter
CN112632187A