Data Retrieval Method Based on Neural Network and Private Set Intersection
Through the neural network-based feature extraction and privacy set interception methods, the problems of data confidentiality and similarity search in cloud computing services are solved, and efficient and secure multi-type data retrieval is achieved.
Patent Information
- Application Number
- CN202311689836.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-12-11
AI Technical Summary
In cloud computing services, traditional searchable encryption technology is difficult to ensure the confidentiality of private data and data characteristics, and at the same time realizes similarity retrieval of multiple types and unstructured data.
The data retrieval method based on neural network and privacy set intersecting is adopted to extract data and query the feature vectors of keywords through neural network models, and the similarity of feature vectors is calculated using privacy set intersecting to ensure the confidentiality of data features.
It realizes the similarity retrieval of multiple types and unstructured data without revealing data characteristics, which improves the reliability of data security sharing and retrieval.
Smart Images

Figure CN117540431B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data security, and particularly relates to a data retrieval method based on neural network and private set intersection. Background Art
[0002] In recent years, with the wide application of big data and cloud computing technologies, providing big data retrieval services in the cloud has become the norm. Data providers outsource query data for storage on cloud servers to reduce equipment costs and operation and maintenance costs by leasing cloud services. Users input query conditions through cloud service interfaces to retrieve query data. Different from traditional retrieval that only performs exact matching for data such as text and numerical values, big data retrieval aims to perform similarity retrieval on multi-type and unstructured data such as images and texts through technologies such as machine learning.
[0003] However, cloud computing services are often provided by third parties. The open nature of the cloud platform itself and the third party determine that the cloud platform is semi-trusted. That is, while the cloud platform faithfully provides services, it may also cause data security problems such as data assets being stolen or user query privacy being leaked. Once an information security incident occurs, it will cause unnecessary losses to data providers, cloud service providers, and users, especially in the case of sensitive data.
[0004] Although some scholars have proposed searchable encryption schemes for the semi-trusted problem in the cloud environment, by encrypting each keyword of the query data and performing ciphertext keyword retrieval, traditional searchable encryption technologies can only provide exact matching of text keywords or numerical data. Traditional retrieval is difficult to ensure strong confidentiality of the private data and data characteristics of each participating party while realizing data security sharing and retrieval, and providing the technical problem of similarity retrieval of multi-type and unstructured data. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a data retrieval method based on neural network and private set intersection to solve the technical problem of realizing data security sharing and retrieval while ensuring strong confidentiality of the private data and data characteristics of each participating party in cloud computing services, and providing similarity retrieval of multi-type and unstructured data.
[0006] The present invention discloses a data retrieval method based on neural network and private set intersection, including the following steps:
[0007] Multiple participating parties preprocess their respective private original data sets to obtain their respective sample sets;
[0008] Each participating party initializes its own neural network model and inputs its respective sample set for model training. After the model converges, each obtains its own trained neural network model. The respective sample sets are input into the respective trained neural network models, and feature vectors of the respective sample sets are extracted from the respective hidden layers. The feature vectors of the sample sets of each participating party are used to perform a private set intersection to obtain a common data set, and the common data set is used as a common database.
[0009] Any participating party inputs a query keyword, and based on the hidden layer of the trained neural network model of this participating party, a feature vector of the query keyword is extracted. The feature vector of the query keyword and the feature vectors of the common data set are used to perform the private set intersection to obtain the results in the common data set that match the feature vector of the query keyword.
[0010] Further, the multiple participating parties preprocess their respective private original data sets to obtain their respective sample sets, including:
[0011] Each of the respective private original data sets contains multiple data records, and each data record is cleaned and denoised.
[0012] Each of the cleaned and denoised data records is tokenized to separate the keywords therein.
[0013] The keywords in each data record are vectorized using a word vector model, and the keywords in each data record are mapped to corresponding word vectors to obtain the word vector sequences of each data record.
[0014] The word vector sequences of each data record are used as samples of each participating party, thereby forming the sample set.
[0015] Among them, the data of the respective private original data sets are in plaintext and are multi-type, unstructured data.
[0016] Further, during the training process of the neural network model, multiple cycles need to be iterated. In each cycle, each participating party batches the samples of its respective training sample set and inputs them into the input layer of the initialized neural network model.
[0017] The hidden layer of the initialized neural network model sequentially calculates the word vectors of each data record in the training sample set, and the output layer outputs classification labels.
[0018] The respective test sample sets are used to input into the neural network model to evaluate the model performance and accuracy.
[0019] After the accuracy meets the requirements, a trained neural network model is obtained. Each sample set is input into the trained neural network model, and the state vector at the last time step is extracted from the respective hidden layer as the feature vector of each sample set.
[0020] Among them, each participant stores the mapping relationship between the keywords in the original dataset and the feature vectors of their respective sample sets.
[0021] Further, in the hidden layer, assuming the time series \(t = 1, 2, 3, \cdots, T\), at \(t = 0\), a hidden state vector \(h\) is initialized 0 as a zero vector;
[0022] For each time step \(t\), by using the input word vector \(x\) at the current time step t and the hidden state vector \(h\) at the previous time step t-1 a new hidden state \(h\) is calculated t , and it becomes the input for the next time step for subsequent loop calculations. The calculation formula is as follows:
[0023] \(h\) t = Activation(W × [x t , h t-1 +b)
[0024] where \(h\) t is the hidden state vector at the current time step, Activation is the activation function, [x t , h t-1 represents the combination and superposition of the input word vector \(x\) t and the hidden state vector \(h\) at the previous time step t-1 , and \(W\) and \(b\) are the weights and biases learned by the neural network model;
[0025] After completing the loop calculation, the hidden state vector \(h\) at the last time step corresponding to the input word vector \(x\) t is obtained. The hidden state vectors at the last time step of the hidden layer corresponding to all word vectors are used as the feature vectors of each data record in the sample; t The feature vectors of all data records of each participant form the feature vector of their respective sample set.
[0026] The feature vectors of all data records of each participant form the feature vector of their respective sample set.
[0027] Further, using private set intersection on the feature vectors of the respective sample sets to obtain a common dataset includes:
[0028] Regarding the feature vectors of each participant's respective sample set as a set;
[0029] Perform an oblivious pseudo-random function mapping on the feature vectors in the set to irreversible pseudo-random values;
[0030] Take the intersection of the pseudo-random values of each party to obtain the common data set of each party;
[0031] Among them, each party saves the mapping relationship between the feature vector and the pseudo-random value.
[0032] Furthermore, the word vector sequence of the query keyword of any party is input into the trained neural network model, and the feature vector corresponding to the query keyword is extracted from the hidden layer as set A, and the common data set is regarded as set B;
[0033] Perform an oblivious pseudo-random function on each feature vector in sets A and B to obtain the corresponding sets of pseudo-random values HA and HB;
[0034] Take the intersection of HA and HB to obtain the intersection set, which is used as the result that matches the feature vector of the input query keyword in the common data set.
[0035] Furthermore, the number of elements in the intersection set of sets HA and HB represents the similarity between sets A and B. The more the number of identical elements in the intersection set, the higher the similarity between the feature vector of the query keyword and the feature vector of the common data set;
[0036] Sort the intersection set in descending order according to the similarity.
[0037] Furthermore, the intersection set is a set of pseudo-random values corresponding to feature vectors;
[0038] Based on the mapping relationship between the pseudo-random value and the feature vector, and the mapping relationship between the feature vector and the original data, obtain the original data corresponding to the pseudo-random value in the intersection set, and further obtain the original data corresponding to the intersection set queried by the party.
[0039] Furthermore, the initialization of the neural network model by each party includes:
[0040] Each party uses the same neural network model on the basis of using the same operating system and hardware environment. The model includes an input layer, a hidden layer, and an output layer, and at least one hidden layer is included;
[0041] Among them, the neural network model is an RNN model.
[0042] Furthermore, the word vector model uses the GloVe model;
[0043] 70% of the said sample set is used as the training sample set, and 30% is used as the test sample set.
[0044] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0045] 1. The present invention extracts the feature vectors of the original data and query keywords through neural network encoding, retains the similarity of the original plaintext data in the N-dimensional feature space, and calculates the similarity between the query keywords and the possible results by performing private set intersection on the feature vectors. While ensuring the confidentiality of data features, the similarity of the feature vectors in the N-dimensional space is not changed, and similarity retrieval of multiple types of unstructured data can be provided;
[0046] 2. Through private set intersection, the confidentiality of the private original plaintext data and data features of each participating party is ensured;
[0047] 3. It is applicable to the similarity retrieval of multiple types of unstructured data, improving the practicality of the retrieval;
[0048] 4. An oblivious pseudorandom function is adopted, and the process is irreversible, ensuring that sensitive information is not leaked during the data sharing process;
[0049] 5. Through the neural network model and similarity calculation, an efficient retrieval method for multiple types of unstructured data is provided, providing a more reliable data retrieval solution for cloud computing services.
[0050] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0052] Figure 1 is a flowchart of a data retrieval method based on neural network and private set intersection;
[0053] Figure 2 is a schematic diagram of each participating party using their respective original plaintext data to train their respective neural network models using the same neural network to obtain a common data set;
[0054] Figure 3 is a schematic diagram of the training process of neural network models of multiple participating parties;
[0055] Figure 4 Schematic diagram of the data processing flow when the participating parties conduct retrieval;
[0056] Figure 5 Schematic diagram of the query application scenarios for multiple participating parties. Specific implementation manners
[0057] The preferred embodiments of the present invention will be specifically described below with reference to the accompanying drawings, where the accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.
[0058] A specific embodiment of the present invention, as Figure 1 shown, discloses a data retrieval method based on neural network and private set intersection, including the following steps:
[0059] Step S1, multiple participating parties preprocess their respective private original data sets to obtain their respective sample sets;
[0060] Step S2, each participating party initializes its own neural network model and inputs its respective sample set for model training. After the model converges, each obtains its own trained neural network model. Input the respective sample set into its own trained neural network model, and extract the feature vectors of the respective sample set from the respective hidden layers; Use private set intersection for the feature vectors of the sample sets of each participating party to obtain a common data set, and use the common data set as a common database;
[0061] Step S3, any one of the participating parties inputs a query keyword, and extracts the feature vector of the query keyword based on the hidden layer of the trained neural network model of this participating party; Use private set intersection for the feature vector of the query keyword and the feature vectors of the common data set to obtain the results in the common data set that match the feature vector of the query keyword.
[0062] Step S1 includes steps S11 - S13, specifically.
[0063] Each of the respective private original data sets contains multiple data records, and each data record is cleaned and denoised;
[0064] Each data record after cleaning and denoising is segmented to separate the keywords therein;
[0065] The keywords in each data record are vectorized using a word vector model, and the keywords in each data record are mapped to corresponding word vectors to obtain the word vector sequences of each data record;
[0066] Take the word vector sequences of each data record as the samples of each participating party, thereby forming the sample set;
[0067] Among them, the data in the respective private original data sets are plaintext, which are multi-type and unstructured data.
[0068] In step S11, each participating party preprocesses its respective private original data set to obtain vectorized original data.
[0069] The original data sets privately owned by each participating party are plaintext data, which must be of the same type. For example, multiple hospitals may have similar or different medical records of the same patient.
[0070] The original plaintext data privately owned by the participating parties are multi-type and unstructured data. The multi-type in the present invention refers to contents such as words, symbols, formulas, etc.; unstructured data usually refers to data information such as names, addresses, diseases, occupations, organizations, etc. embedded in texts, documents, logs, images; in the present invention, unstructured data mainly refers to information such as names and addresses embedded in text documents.
[0071] (1) Clean and denoise the original data set.
[0072] The original data set contains multiple data records; before each participating party trains its respective private neural network model, it is necessary to preprocess the original data set it owns.
[0073] Segment each data record in the original data set to separate the keywords therein;
[0074] Remove words, punctuation marks, and special characters that have no practical meaning;
[0075] If there are English letters, convert them all to lowercase to eliminate the inconsistency between uppercase and lowercase.
[0076] (2) Vectorize the original data set after cleaning and denoising.
[0077] Use a word vector model to map each word in the original data set after preprocessing into a word vector;
[0078] Map all the words in each data record into corresponding word vectors to obtain the word vector sequence of each data record.
[0079] Preferably, a pre-trained word vector model GloVe is selected. The GloVe model pays more attention to capturing global vocabulary statistics, which is very helpful for calculating the semantic similarity between words.
[0080] In step S12, the vectorized original data set is used as a sample set, and the sample set is divided into a training sample set and a test sample set.
[0081] Exemplarily, 70% of the vectorized original dataset is used as the training sample set, and 30% is used as the test sample set.
[0082] The training sample set is used to train the neural network model so that the neural network model can capture the patterns and features in the data;
[0083] The test sample set is used to evaluate the performance of the neural network model on unseen data and check the generalization ability of the data. After the model training is completed, the test set is used for evaluation. By inputting the test set into the model, the performance, effect, and accuracy of the model are evaluated. The use of the test set helps to determine whether the model overfits the training data, prevent overfitting, and improve the generalization performance of the model.
[0084] Step S2 is divided into steps S21 - S23, as Figure 2 shown, specifically.
[0085] Step S21, each party initializes its own neural network model;
[0086] Based on using the same operating system and hardware environment, each party uses the same neural network model, which includes an input layer, a hidden layer, and an output layer, with at least one hidden layer;
[0087] Among them, the neural network model is an RNN model.
[0088] Each party selects a suitable neural network structure. Preferably, for the present invention, RNN (Recurrent Neural Network) is selected. The RNN model performs well in processing sequence data and is very suitable for text tasks considering context relationships.
[0089] When selecting the model deep learning framework, preferably, TensorFlow (a deep learning platform) which is relatively more mature in deployment in the production environment is selected.
[0090] The neural network model includes an input layer, a hidden layer, and an output layer, with at least one hidden layer.
[0091] Each party's neural network model selects the same parameters, the same method and framework, and the same configuration environment.
[0092] The neural network model has the same loss function (Contrastive Loss), pseudo - random function, and optimization algorithm (Adam algorithm); the same framework is TensorFlow.
[0093] The same parameters, configuration information, method frameworks, etc. are selected to reduce the impact on the final result caused by the differences in the neural network model and the methods used in the remaining stages; if different parameters are selected, it may lead to the same feature vectors being obtained even when the original information is different.
[0094] Among them, the number and type of network layers of the neural network model, and the number of iterations are the same, and the learning rate is the adaptive learning rate of each parameter calculated by the optimization algorithm Adam.
[0095] Calculations are performed for multiple time steps in the hidden layer, and during the calculations, the calculation results are corrected backward, which includes information such as weights, correction amounts, and learning rates. At the same time, the weights need to be initialized; meanwhile, during the process of training the neural network, a loss function is used to measure the similarity or distance between samples, so as to represent the association between samples; during this period, in order to make the effect of the finally obtained model better, some optimization algorithms are used to minimize the loss function. Using optimization algorithms can make the similarity between related keywords greater and the association between unrelated keywords smaller, so that the performance of the model is better and the keyword matching is more accurate.
[0096] Step S22, each participating party uses the training set to input their respective private neural network models for training. When the accuracy meets the requirements, a trained neural network model is obtained, and the feature vectors corresponding to their respective private original plaintext data sets are extracted from the hidden layer of the trained neural network model.
[0097] As Figure 3 shown, during the training process of the neural network model, multiple cycles need to be iterated. In each cycle, each participating party batches the samples of their respective training sample sets and inputs them into the input layer of the initialized neural network model;
[0098] The hidden layer of the initialized neural network model sequentially calculates the word vectors of each data record in the training sample set, and the output layer outputs classification labels;
[0099] Use their respective test sample sets to input into the neural network model to evaluate the model performance and accuracy;
[0100] When the accuracy meets the requirements, a trained neural network model is obtained, and their respective sample sets are input into the trained neural network model, and the state vector of the last time step is extracted from their respective hidden layers as the feature vector of their respective sample sets;
[0101] Among them, each participating party saves the mapping relationship between the keywords in the original data set and the feature vectors of their respective sample sets.
[0102] Each participating party uses the word vectors of the training sample set as input and the classification label vector as output to train their respective neural network models. After the accuracy meets the requirements, a trained neural network model is obtained, and the feature vectors of their respective private original plaintext data are extracted from the hidden layer of the trained neural network model;
[0103] Use the neural network model to convert their respective sample sets into corresponding feature vectors.
[0104] Since the result of the output layer is only the prediction result for a specific task and does not contain sufficient information vectors to represent the original plaintext data, the state vector of the last time step in the hidden layer is selected, which can contain the information of the entire input sequence.
[0105] Specifically, taking a data record in the original plaintext data as an example, X = [x 1 , x 2 ,...., x m , where x 1 to x m are the keywords that may be included in a piece of data. The number of keywords included in each data record may be different, so m is not a definite number. After preprocessing the data of X and inputting it into the neural network model, when calculating in the hidden layer, the word vectors of x 1 to x m are input in sequence, and finally the feature vector lable = [l 1 , l 2 ,...., l n belonging to the entire data record output by the hidden layer is obtained. This process is performed for each data record held by the participating party.
[0106] The obtained word vectors corresponding to each data record are input into the input layer of the neural network. For the multiple word vectors corresponding to multiple keywords in a data record, they need to be calculated sequentially when calculating in the hidden layer. The word vectors corresponding to each data record are input into the input layer of the neural network at the same time, and the feature vectors are calculated sequentially according to the order before and after the words in the hidden layer. Because the input information contains multiple keywords, and each keyword has its own characteristics, so they are calculated sequentially.
[0107] When the calculation in the hidden layer is completed, the feature vector corresponding to the entire data record can be obtained from the state vector of the last time step in the hidden layer.
[0108] In the hidden layer, assume that the time series t = 1, 2, 3,..., T. At t = 0, a hidden state vector h 0 is initialized as a zero vector;
[0109] For each time step t, by using the input word vector x of the current time stept and the hidden state vector h of the previous time step t-1 compute the new hidden state h t , and it becomes the input for the next time step for subsequent loop calculations. The calculation formula is as follows:
[0110] h t = Activation(W × [x t , h t-1 + b)
[0111] where h t is the hidden state vector of the current time step, Activation is the activation function, [x t , h t-1 means combining and stacking the input word vector x t and the hidden state vector h of the previous time step t-1 ; W and b are the weights and biases learned by the neural network model;
[0112] After completing the loop calculation, obtain the hidden state vector h of the last time step corresponding to the input word vector x t , and the hidden state vectors of the last time step of the hidden layers corresponding to all word vectors are used as the feature vectors of each data record in the sample; t The feature vectors of all data records of each party form the feature vectors of their respective sample sets.
[0113] Each party's feature vectors of all data records form the feature vectors of their respective sample sets.
[0114] Define the loss function: When selecting the loss function, it needs to be selected according to the task type. Preferably, the Contrastive Loss is selected. The Contrastive Loss is a loss function specifically used to measure the similarity of sample pairs and is used to measure the similarity between two samples.
[0115] The Contrastive Loss can usually be expressed by the following formula:
[0116] L = (1 - Y) × D 2 + Y × max(margin - D, 0) 2
[0117] where L is the loss value; Y is the binary label, 0 represents a negative example text pair, and 1 represents a positive example text pair; D is the distance or similarity score between the feature vectors; margin is a hyperparameter representing the interval threshold between positive and negative example text pairs; for positive example text pairs (Y = 1), it is desired that their distance is as small as possible, so the loss is penalized by the distance between the two embedding vectors. For negative example text pairs (Y = 0), it is desired that their distance is greater than the margin, otherwise they will be penalized.
[0118] Positive example texts refer to vector feature samples with labels; negative example texts refer to vector feature samples without labels.
[0119] Use gradient descent or other optimization algorithms to minimize the loss function. By continuously adjusting the model parameters, which are the weight W and bias b, the vector features of positive example text pairs become more similar and closer, while the vector features of negative example text pairs become more distant.
[0120] This process is repeated continuously until the accuracy meets the requirements. Exemplarily, the number of training iterations is selected as 20 times. The actual number of training iterations can be adjusted according to the performance of the neural network model on the training set and the test set.
[0121] The neural network model will learn how to bring the vector features of similar text pairs closer and pull the vector features of dissimilar text pairs farther apart to achieve the task of text similarity matching.
[0122] Select an optimizer: The optimization algorithm can choose Adam. Adam is a widely used adaptive learning rate algorithm for each parameter and is used to update the parameters during the training process of the neural network model. The Adam algorithm dynamically adjusts the parameter weights W and bias b of the neural network model during the training process. This algorithm usually can achieve good results for the similarity calculation task.
[0123] When training the model, multiple epochs need to be iterated. In each epoch, the training sample set data is input into the model in batches, the loss is calculated through forward propagation, and then the model parameters are updated through backpropagation.
[0124] The neurons in the hidden layer perform forward calculations and at the same time perform backward weight corrections. They receive the original query data output by the input layer and also have information such as weights, weight correction amounts, and learning rates. The weight W is initialized randomly, and the value range is [0, 1];
[0125] Each participating party sequentially executes the above training process for each piece of data it holds. After the neural network model converges, finally, the feature vectors obtained by calculating the hidden layer of the datasets each party holds are obtained. The intersection of the feature vectors held by all participating parties is calculated using private set intersection to obtain a common dataset.
[0126] Finally, store this common dataset as a database. At this time, the data type in the database is the pseudo-random values of the feature vectors.
[0127] In step S23, use private set intersection to find the intersection of the feature vectors of the respective private original plaintext datasets to obtain the feature vectors of the common dataset, and store the feature vectors of the common dataset as a common database.
[0128] Performing private set intersection on the feature vectors of the respective sample sets to obtain a common data set includes:
[0129] Regarding the feature vectors of each participant's respective sample set as a set;
[0130] Performing an oblivious pseudorandom function on the feature vectors in the set to map them to irreversible pseudorandom values;
[0131] Taking the intersection of the pseudorandom values of each participant to obtain the common data set of each participant;
[0132] Wherein, each participant stores the mapping relationship between the feature vector and the pseudorandom value.
[0133] After using private set intersection to obtain the common part of the data sets of the participating parties, that is, after obtaining the common part of the feature vectors corresponding to the keywords in the data they each hold, taking the pseudofeature values of this common feature vector set as the data set and storing it in the database;
[0134] Step S3 includes steps S31 - S32, specifically.
[0135] In step S31, any participating party inputs a query keyword, and extracts the feature vector of the query keyword based on the hidden layer of the neural network model trained by this participating party;
[0136] As Figure 4 shown, the word vector sequence of the query keyword of any participating party is input into the trained neural network model, and the feature vector corresponding to the query keyword extracted from the hidden layer is regarded as set A, and the common data set is regarded as set B;
[0137] Performing an oblivious pseudorandom function on each feature vector in sets A and B to obtain corresponding sets of pseudorandom values HA and HB;
[0138] Taking the intersection of HA and HB to obtain an intersection set, which is used as the result in the common data set that matches the feature vector of the input query keyword.
[0139] In step S32, using the private set intersection on the feature vector of the query keyword and the feature vectors of the common data set to obtain the result in the common data set that matches the feature vector of the query keyword.
[0140] The number of elements in the intersection set of sets HA and HB represents the similarity degree of sets A and B. The more the number of identical elements in the intersection set, the higher the similarity between the feature vector of the query keyword and the feature vectors of the common data set;
[0141] Arrange the intersection set in descending order according to the similarity.
[0142] The intersection set is a set of pseudo-random values corresponding to the feature vectors;
[0143] Based on the mapping relationship between the pseudo-random values and the feature vectors, and the mapping relationship between the feature vectors and the original data, obtain the original data corresponding to the pseudo-random values in the intersection set, and further obtain the original data corresponding to the intersection set queried by the participating parties.
[0144] Among them, the query keyword of any participating party is one or more, and the logical AND relationship is default between multiple query keywords.
[0145] The word vector of the query keyword of any participating party is input into the input layer of the neural network model, and the query keyword label (i.e., the query keyword feature vector) is obtained through the hidden layer, and the classification label vector representation is performed through the output layer;
[0146] The query condition feature vector output by the hidden layer goes to the output layer of the neural network, and the classification label vector is the output result of the output layer.
[0147] According to the labels of the public data set, set the classification label set and represent it with a one-dimensional vector. Define the classification label vector lable = [l 1 ,l 2 ,....,l n , where n is the dimension of the label vector, and the i-th dimension l i represents a query keyword label. If the keyword of a certain piece of data conforms to the query keyword label l i , then set the value corresponding to the query keyword label in the label vector to l i = 1, otherwise l i = 0.
[0148] The input layer receives the word vector of the query keyword; the hidden layer calculates the feature vector of the keyword corresponding to the input word vector of the query keyword. This feature vector captures all the information of the input query keywords, and the output layer outputs the classification label vector corresponding to the query result. The feature vector obtained from the hidden layer contains the information of all the keywords included in the entire data record.
[0149] The input word vector of the query keyword X' = [x' 1 ,x' 2 ,....,x' n . The trained neural network calculates the newly input query keyword, and obtains the feature vector output by the hidden layer as the feature vector of the query keyword. The query keyword refers to the keyword input by any participating party when retrieving data.
[0150] The described method for private set intersection includes: Any participating party hopes to obtain query results in the public dataset through query keywords. After inputting the query keywords, the feature vectors of the input keywords are obtained through a pre-trained neural network. The set of these feature vectors is regarded as a set A. At the same time, all the feature vectors in the public database are also regarded as a set B. Through a given oblivious pseudorandom function (the mapping relationship of this function is one-to-one), the two sets are respectively mapped into HA and HB to obtain the intersection of the sets HA and HB and the corresponding intersection elements. Since the mapping relationship of the mapping function is one-to-one, that is, when using this mapping function, given a certain input, the output result is also uniquely determined. Therefore, the intersection of A and B and the intersection elements can be obtained from the intersection of HA and HB. The more elements there are in the intersection, the more similar the two sets A and B are, that is, the more similar the feature vectors of the query keywords are to some feature vectors in the public database.
[0151] Suppose the set converted from the feature vectors corresponding to the query keywords is A, and the set converted from the feature vectors corresponding to the public dataset is B. At this time, the number of elements in the two sets is the same, set as n;
[0152] The pseudorandom function is used to map the set A converted from the feature vectors corresponding to the query keywords and the set B corresponding to the public dataset to new values. The purpose of the mapping is to perform similarity calculation while protecting privacy.
[0153] By mapping the feature vectors to irreversible pseudorandom values, due to their irreversibility, it is impossible to infer the feature vectors of set A or B from the pseudorandom values.
[0154] For each element in B, a corresponding oblivious pseudorandom function is executed, that is, when a specific element in the set can only select the corresponding pseudorandom function to execute. After executing the pseudorandom function, it is mapped to another different pseudorandom value, and the set of pseudorandom values of set B thus obtains set HB;
[0155] For each element in A, the corresponding pseudorandom function is also executed and mapped to a pseudorandom value, and the set of pseudorandom values of A obtains set HA;
[0156] Since the pseudorandom functions executed by specific elements are determined, the results obtained are also the same. That is, when an element in set HA is the same as an element in set HB, the corresponding set elements in the original sets A and B are also the same;
[0157] After obtaining HA, the intersection of set HA and HB is calculated to obtain the number of elements in the intersection and the corresponding set B;
[0158] The number of intersection elements of the two new sets after being mapped by the pseudo-random function represents the similarity. The more the number of intersection elements, the greater the similarity. The fewer the number of elements in the intersection, the smaller the similarity.
[0159] Since the mapping relationship of the pseudo-random function is one-to-one, when there is an intersection between set HA and set HB and there are identical elements, there is also an intersection between the original sets A and B, and the elements corresponding before being converted by the pseudo-random function are also the same. That is, a part of the feature vectors corresponding to the keywords to be searched is the same as the feature vectors corresponding to the public data set. That is, a part of the information of the original data in the public data set is the same as the keywords to be searched.
[0160] Finally, the query results obtained are displayed and sorted in descending order according to the number of elements in the intersection set.
[0161] The present invention extracts the feature vectors of the query keywords and query conditions through a neural network, retains the similarity of the original data of the participating parties in the N-dimensional feature space, obtains relatively qualified results through private set intersection of the sets after the feature vectors are converted, and while ensuring the confidentiality of data features, does not change the similarity of the feature vectors in the N-dimensional space, and can provide similarity retrieval of multiple types of unstructured data.
[0162] In the present invention, unstructured data usually refers to names, addresses, diseases, occupations, organizations, etc., and these information are usually embedded in texts, documents, logs, images, etc.; in this document, it mainly refers to information such as names and addresses embedded in text documents.
[0163] Exemplarily, multiple participating parties are multiple medical institutions.
[0164] Such as Figure 5 As shown, assume there are three medical institutions A, B, and C, and they hope to perform secure retrieval and sharing of their respective medical data through cloud computing services. The following is an application example of the present invention in the field of medical data:
[0165] (1) Data preprocessing:
[0166] Each medical institution first extracts multiple patient records from its original medical data set.
[0167] Preprocess each record, including cleaning and denoising, and extracting keywords by word segmentation to ensure the availability and consistency of medical data.
[0168] (2) Training of respective neural network models:
[0169] Each medical institution uses the same neural network model, initializes and trains it. The model learns the feature data of the patient records to ensure that each medical institution uses similar feature representations.
[0170] (3) Private Set Intersection:
[0171] Using the method of private set intersection, the feature vectors of each patient in each medical institution are mapped to irreversible pseudo-random values to obtain a common data set. This step ensures the privacy of medical data.
[0172] (4) Query Processing:
[0173] When medical institution A inputs a query keyword, such as the medical insurance card number or ID card number of a patient, the neural network model extracts the feature vector of the query keyword.
[0174] Through private set intersection, a result that matches the query keyword is obtained, indicating that similar patient records have been found in the common data set.
[0175] (5) Similarity Calculation and Sorting:
[0176] The number of elements in the intersection set represents the similarity between the query keyword and the common data set. The higher the similarity, the greater the matching degree.
[0177] The results are sorted in descending order of similarity to provide the most relevant retrieval results.
[0178] (6) Restoration of Original Data:
[0179] Through the mapping relationship, the pseudo-random values in the intersection set are restored to the original medical data, enabling medical institutions to understand the specific patient records corresponding to the query results.
[0180] A patient is being treated for pneumonia in medical institution A. It is found that the patient was transferred to A after treatment in medical institution B. Based on the treatment plan of B, a corresponding treatment plan is formulated to accelerate the treatment process of the patient's disease.
[0181] The data privacy between medical institutions is ensured, and the data characteristics of different institutions remain confidential. Allowing similarity retrieval of multiple types of unstructured medical data improves the applicability of the retrieval. Through the neural network model and similarity calculation, an efficient medical data retrieval method is provided, which is helpful for medical research and collaborative work.
[0182] This example illustrates how the inventive method is applied to the secure retrieval of medical data, enabling different medical institutions to share query results without disclosing patient privacy information.
[0183] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0184] 1. The present invention extracts feature vectors of the original data and query keywords through neural network encoding, preserves the similarity of the original plaintext data in the N-dimensional feature space, calculates the similarity between the query keywords and possible results by performing private set intersection on the feature vectors, ensures the confidentiality of data features while not changing the similarity of the feature vectors in the N-dimensional space, and can provide similarity retrieval for various types of unstructured data.
[0185] 2. Through private set intersection, the confidentiality of the private original plaintext data and data features of each participating party is ensured.
[0186] 3. It is applicable to the similarity retrieval of various types of unstructured data, improving the practicality of the retrieval.
[0187] 4. By using oblivious pseudorandom functions, the process is irreversible, ensuring that sensitive information is not leaked during the data sharing process.
[0188] 5. Through the neural network model and similarity calculation, an efficient retrieval method for various types of unstructured data is provided, providing a more reliable data retrieval solution for cloud computing services.
[0189] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0190] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A data retrieval method based on neural network and private set intersection, characterized in that, it includes the following steps: Multiple parties preprocess their respective private original data sets to obtain their respective sample sets; Each party initializes its own neural network model and inputs its respective sample set for model training. After the model converges, each party obtains its own trained neural network model. Input its respective sample set into its own trained neural network model, and extract the feature vectors of its respective sample set from the respective hidden layers; Use private set intersection for the feature vectors of the sample sets of each party to obtain a common data set, and use the common data set as a common database; Any one of the parties inputs a query keyword, and extracts the feature vector of the query keyword based on the hidden layer of the trained neural network model of this party; Use the private set intersection of the feature vector of the query keyword and the feature vectors of the common data set to obtain the results in the common data set that match the feature vector of the query keyword; Using private set intersection for the feature vectors of the respective sample sets to obtain a common data set includes: Regarding the feature vectors of the respective sample sets of each party as a set; Performing an oblivious pseudo-random function on the feature vectors in the set to map them to irreversible pseudo-random values; Take the intersection of the pseudo-random values of each party to obtain the common data set of each party; Among them, each party saves the mapping relationship between the feature vector and the pseudo-random value; The word vector sequence of the query keyword of any one of the parties is input into the trained neural network model, and the feature vector corresponding to the query keyword is extracted from the hidden layer and regarded as set A, and the common data set is regarded as set B; Perform an oblivious pseudo-random function on each feature vector in sets A and B to obtain corresponding sets of pseudo-random values HA and HB; Take the intersection of HA and HB to obtain an intersection set, which is used as the result in the common data set that matches the feature vector of the input query keyword.
2. The method according to claim 1, characterized in that, The multiple parties preprocess their respective private original data sets to obtain their respective sample sets, including: Each of the respective private original data sets contains multiple data records, and each data record is cleaned and denoised; Perform word segmentation on each data record after cleaning and denoising to separate the keywords therein; Vectorize the keywords in each data record using a word vector model, map the keywords in each data record to corresponding word vectors, and obtain the word vector sequence of each data record; Use the word vector sequences of each data record as samples of each party, thereby forming the sample set; Among them, the data of the respective private original data sets is in plaintext and is multi-type, unstructured data.
3. The method according to claim 2, characterized in that, During the training process of the neural network model, multiple cycles need to be iterated. In each cycle, each party batches the samples of its respective training sample set and inputs them into the input layer of the initialized neural network model; The hidden layer of the initialized neural network model sequentially calculates the word vectors of each data record in the training sample set, and the output layer outputs classification labels; Use their respective test sample sets to input into the neural network model to evaluate the model performance and accuracy; When the accuracy meets the requirements, obtain the trained neural network model, input their respective sample sets into the trained neural network model, and extract the state vector of the last time step from their respective hidden layers as the feature vector of their respective sample sets; Among them, each party saves the mapping relationship between the keywords in the original data set and the feature vectors of their respective sample sets.
4. The method according to claim 3, wherein, In the hidden layer, assume that the time series is \(t = 1, 2, 3,\cdots, T\). At \(t = 0\), initialize a hidden state vector \(h\) 0 as a zero vector; For each time step t, by using the input word vector x at the current time step t and the hidden state vector h at the previous time step t-1 a new hidden state h is calculated t , and it becomes the input for the next time step for subsequent recurrent calculations. The calculation formula is as follows: h t = Activation(W × [x t , h t-1 + b) Among them, h t is the hidden state vector at the current time step, Activation is the activation function, [x t , h t-1 represents combining and superimposing the input word vector x t and the hidden state vector h t-1 from the previous time step, and W and b are the weights and biases learned by the neural network model; After completing the loop calculation, the input word vector x is obtained t The hidden state vector h at the last time step corresponding to it t , and the hidden state vectors at the last time step of the hidden layers corresponding to all word vectors are used as the feature vectors of each data record in the sample set; The feature vectors of all data records of each party form the feature vector of their respective sample sets.
5. The method according to claim 1, wherein, The number of elements in the intersection set of the sets HA and HB represents the similarity between the sets A and B. The more the number of identical elements in the intersection set, the higher the similarity between the feature vector of the query keyword and the feature vector of the common data set; Sort the intersection set in descending order according to the similarity.
6. The method according to claim 5, wherein, The intersection set is a set of pseudo-random values corresponding to the feature vectors; Based on the mapping relationship between the pseudo-random value and the feature vector, and the mapping relationship between the feature vector and the original data, obtain the original data corresponding to the pseudo-random value in the intersection set, and further obtain the original data corresponding to the intersection set queried by the party.
7. The method according to claim 6, wherein, The initialization of the respective neural network models by each party includes: Each party uses the same neural network model on the basis of using the same operating system and hardware environment. The model includes an input layer, a hidden layer, and an output layer, and at least one hidden layer is included; Among them, the neural network model is an RNN model.
8. The method according to any one of claims 1-7, wherein, The word vector model adopts the GloVe model; 70% of the sample set is used as the training sample set, and 30% is used as the test sample set.
Citation Information
Patent Citations
Efficient searchable proxy privacy set intersection method and device
CN114491613A
Training method of longitudinal federated neural network model
CN116992955A