Image retrieval method and system supporting sharing of multiple data sources
By using key distribution and ciphertext recryption technologies in the cloud computing environment, a layered index is built, and the secure and efficient retrieval of encrypted images shared by multiple data sources is achieved, solving the privacy and efficiency of image retrieval in the cloud environment.
Patent Information
- Application Number
- CN202510268711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In the cloud computing environment, how to achieve secure and efficient retrieval of encrypted images, especially in multi-data source sharing scenarios, it is necessary to protect data privacy, and ensure the accuracy and efficiency of retrieval.
Through a trusted third party, keys are generated for each data owner and query user and key distribution is performed. The data owner extracts image feature vectors and encrypts them, forming ciphertext clusters and ciphertext image sets, and uploads them to the cloud server. The cloud server group reencrypts the ciphertext clusters through reencrypts key pairs, builds a hierarchical index, query users reencrypt the query trap gate through reencrypts key pairs, and searches, and finally query users to use image keys to decrypt the plaintext image.
It realizes secure and efficient retrieval of encrypted images shared by multiple data sources in the cloud environment, protects data privacy, reduces retrieval computing overhead, and improves user query response speed.
Smart Images

Figure CN119788424B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image retrieval technology, and in particular to an image retrieval method and system supporting sharing of multiple data sources. Background Art
[0002] The rapid popularization of smart devices has led to a sharp increase in image data, which has brought an increasing storage burden to data owners with limited resources. The emergence of cloud servers has met their needs for large-scale image storage and can achieve more convenient data sharing. Therefore, outsourcing local large-scale images to third-party cloud servers has become a very popular measure. However, since plaintext images contain a large amount of sensitive information (such as remote sensing images, medical images, etc.) and cloud servers are semi-trusted, privacy and security have become the main issues that data owners need to pay attention to. In order to protect the information in the image set from being leaked, data owners usually encrypt the image set before outsourcing, but the encrypted images also bring the problem of difficult retrieval. Therefore, how to achieve secure and efficient retrieval of encrypted images in a cloud computing environment has become a hot issue in the current field. Summary of the invention
[0003] In order to solve the above problems, the present invention provides an image retrieval method and system supporting sharing of multiple data sources.
[0004] The present invention provides an image retrieval method supporting sharing of multiple data sources, comprising:
[0005] The trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, generates a re-encryption key pair for a cloud server group, and distributes the keys, wherein the cloud server group includes a first cloud server and a second cloud server;
[0006] Each of the data owners extracts a feature vector of each image in the respective image set, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; encrypts each of the images in the image set based on the image key to obtain a ciphertext image set, and sends each of the ciphertext image sets and the ciphertext cluster to the first cloud server;
[0007] The second cloud server assists the first cloud server in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and constructs a hierarchical index based on each of the new ciphertext clusters;
[0008] The query user extracts a feature vector of the query image and generates a query trapdoor, and sends the query trapdoor to the first cloud server;
[0009] The second cloud server assists the first cloud server in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; and searches the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image;
[0010] The first cloud server sends the at least one ciphertext image and the identifier of the corresponding data owner to the query user;
[0011] The querying user decrypts each ciphertext image using the corresponding image key based on the identifier to obtain the queried image.
[0012] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the step of constructing a hierarchical index based on each of the new ciphertext clusters includes:
[0013] Step 1: For each new ciphertext cluster, take the cluster center of the new ciphertext cluster as the representative vector, take the feature vector in the new ciphertext cluster as the stored data, and construct the initial node ,in, is the representative vector of the initial node, is the level number of the initial node in the hierarchical index, The data stored in the initial node is set to a first set whose content is empty;
[0014] Step 2: Add all the initial nodes to the first set;
[0015] Step 3: If the first set is not empty, randomly select an initial node from the first set and record it as the current node, and add the current node to the preset list; if the first set is empty, execute step 8;
[0016] Step 4: Calculate the distance between the current node and each other node, and record the other node with the smallest distance to the current node as the first node, where the other node is any initial node except the current node;
[0017] Step 5: If the first node already exists in the list, calculate the average vector of the representative vector of the current node and the representative vector of the first node, and use the average vector as the second node; record the current node as a child node of the second node, and record the first node as an auxiliary child node of the second node, and obtain a retrieval tree composed of all the initial nodes in the list, wherein the number of layers of the root node of the retrieval tree is the sum of the number of layers of the current node and one; return to step 3 to continue execution;
[0018] Step 6: If the first node already exists in the search tree, record the current node as a child node of the first node, and merge all the initial nodes in the list as subtrees into the search tree; return to step 3 to continue execution;
[0019] Step 7: If the first node exists in the first set, the first node is taken out from the first set, and the current node is recorded as a child node of the first node, and the first node is recorded as a new current node, and the process returns to step 4 to continue.
[0020] Step 8: If the number of all current root nodes is greater than 1, each root node is recorded as a new initial node, and the process returns to step 2 to continue execution; if the number of all current root nodes is 1, the hierarchical index construction is completed.
[0021] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the calculating of the distance between the current node and each other node includes:
[0022] For any of the other nodes, perform the following steps:
[0023] The first cloud server calculates the difference of ciphertext data at the same position in each dimension between the representative vector of the current node and the representative vector of the other nodes to obtain a ciphertext difference;
[0024] Based on the blinding method, a first random number is introduced into the ciphertext difference to obtain a first ciphertext difference, and a second random number is introduced into the ciphertext difference to obtain a second ciphertext difference, wherein the first random number is different from the second random number; the first ciphertext difference and the second ciphertext difference are partially decrypted to obtain a first intermediate ciphertext difference and a second intermediate ciphertext difference;
[0025] The second cloud server processes the first ciphertext difference and the second ciphertext difference respectively based on a partial decryption algorithm to obtain a third intermediate ciphertext difference and a fourth intermediate ciphertext difference; processes the first intermediate ciphertext difference and the third intermediate ciphertext difference based on a decryption sharing algorithm to obtain first blinded data, and processes the second intermediate ciphertext difference and the fourth intermediate ciphertext difference to obtain second blinded data;
[0026] Calculating a first ciphertext product of the first blinded data and the second blinded data, and encrypting the first ciphertext product according to an encryption algorithm based on a public key in the re-encryption key pair to obtain a second ciphertext product;
[0027] The first cloud server encrypts the product of the first random number and the second random number based on the encryption algorithm to obtain a third ciphertext product;
[0028] Based on the second ciphertext product, the third ciphertext product and the ciphertext difference, the ciphertext of the square of the Euclidean distance between the current node and the other nodes is obtained; and the ciphertext of the square of the Euclidean distance between the current node and the other nodes is used as the distance between the current node and each other node.
[0029] According to an image retrieval method supporting multi-data source sharing provided by the present invention, searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image includes:
[0030] Step I: Create an empty second set and initialize the root node of the hierarchical index as the entry node, the second set is used to record the minimum distance between the new query trapdoor and each layer node of the hierarchical index under the current path during the search process;
[0031] Step II: Obtain all descendant nodes of the next layer of the entry node, calculate the distance between the new query trapdoor and each descendant node, and determine the minimum distance between the new query trapdoor and each descendant node;
[0032] Step III: If the second set is empty, all the descendant nodes are recorded as target nodes, and the number of layers and the minimum distance of each target node are added to the second set, and step VI is executed;
[0033] Step IV: If the second set is not empty, filter out the descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, and record them as target nodes; if there are no descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, record the auxiliary child nodes of the entry node and the child nodes of the next layer as target nodes;
[0034] Step V: Based on a size comparison algorithm, select the smaller of the minimum distance and the minimum distance corresponding to the number of layers of the entry node in the second set, and record it as the minimum distance corresponding to the number of layers of the target node;
[0035] Step VI: If the target node has a node with a layer number greater than 2, the target node is used as a new entry node and the process returns to step II to continue; otherwise, the target node is added to the intermediate node list;
[0036] Step VII: When the intermediate node list is no longer updated, obtain the descendant nodes of all nodes in the intermediate node list, calculate the distance between the new query trapdoor and each descendant node, and select a set number of descendant nodes with the smallest distance to the new query trapdoor, and record them as the final nodes obtained by searching the hierarchical index, where the set number is the number of data owners;
[0037] Step VIII: searching for the feature vector corresponding to the new query trapdoor from the feature vectors stored in the final node, and using the ciphertext image corresponding to the feature vector corresponding to the new query trapdoor as the ciphertext image corresponding to the query image.
[0038] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the distance between the new query trapdoor and each of the descendant nodes is a ciphertext distance;
[0039] The determining of the minimum distance between the new query trapdoor and each of the descendant nodes comprises:
[0040] Comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes, and determining the minimum plaintext distance among the plaintext distances;
[0041] Taking the ciphertext distance corresponding to the minimum plaintext distance as the minimum distance;
[0042] The step of comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes includes:
[0043] For any two ciphertext distances between the new query trapdoor and each of the descendant nodes, the following steps are performed, where the two ciphertext distances include a first ciphertext distance and a second ciphertext distance:
[0044] The first cloud server encrypts one based on the encryption algorithm and the public key in the re-encryption key pair to obtain a ciphertext of one;
[0045] Multiply the ciphertext of the first ciphertext by the square of the first ciphertext distance and the second ciphertext distance respectively to obtain first ciphertext data and second ciphertext data;
[0046] Generate a third random number and a fourth random number, and encrypt the fourth random number to obtain a ciphertext random number; the third random number is greater than the fourth random number, and the binary length of the third random number is less than a first set length, the first set length is one quarter of the binary length of a public parameter, and the public parameter is a part of a public key in a re-encryption key pair;
[0047] Calculate based on the ciphertext random number, the third random number, the first ciphertext data and the second ciphertext data to obtain third ciphertext data; and partially decrypt the third ciphertext data to obtain fourth ciphertext data;
[0048] The second cloud server partially decrypts the third ciphertext data based on a partial decryption algorithm to obtain fifth ciphertext data, and decrypts the fourth ciphertext data and the fifth ciphertext data based on a decryption sharing algorithm to obtain plaintext data;
[0049] If the binary length of the plaintext data is greater than the second set length, the plaintext distance corresponding to the first ciphertext distance is less than the plaintext distance corresponding to the second ciphertext distance; if the binary length of the plaintext data is less than or equal to the second set length, the plaintext distance corresponding to the first ciphertext distance is greater than or equal to the plaintext distance corresponding to the second ciphertext distance, and the second set length is half of the binary length of the public parameter.
[0050] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, and generates a re-encryption key pair for a cloud server group, including:
[0051] For each of the data owners , each of the query users And any one of the cloud server groups performs the following steps:
[0052] The trusted third party Generate four different odd prime numbers , , , ,in, , , .
[0053] Calculate common parameters and private key , choose the order Generators of ,and ,in, , Representation model The multiplicative group under ; Generate a key pair .
[0054] The private key in the key pair Randomly split into two parts and get the split partial private key pair , and at the same time satisfy the following formula:
[0055]
[0056] Wherein, m is 1 or 2;
[0057] At the same time, for the data owner Generate Image Key ,in Indicates Data owners;
[0058] Traverse each of the data owners , each of the query users and cloud server group, obtain the data owner Image key With index key pair , the query user Retrieval key pair , and the re-encryption key pair of the cloud server ,in, Indicates Query users.
[0059] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the key distribution includes:
[0060] The trusted third party sends the image key to with the index key pair Sent to the data owner , the image key Retrieve the key pair with Sent to the query user , the re-encryption key pair In and , the index key pair In and , and the retrieval key pair In and Send the re-encryption key pair to the first cloud server In and , the index key pair In and , and the retrieval key pair In and Send it to the second cloud server.
[0061] According to an image retrieval method supporting multi-data source sharing provided by the present invention, each of the data owners extracts a feature vector of each image in their respective image sets, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster, including:
[0062] The data owner Extract the feature vectors of each image in the image set, and use a clustering algorithm to cluster the feature vectors of each image in the image set to obtain Cluster ;
[0063] Each of the clusters Each dimension of the data in each cluster center and its characteristic vector is processed in a unified format, and based on the public key in the index key pair , for the cluster The cluster centers and their characteristic vectors are encrypted after the unified format to obtain the ciphertext clusters. .
[0064] According to an image retrieval method supporting multi-data source sharing provided by the present invention, the second cloud server assists the first cloud server to re-encrypt each ciphertext cluster based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, including:
[0065] The first cloud server determines the index key pair based on the fifth random number and the public key of the index key pair. , for the ciphertext cluster Perform blinding processing to obtain a new ciphertext cluster , and based on the partial private key in the index key pair For the new ciphertext cluster Perform partial decryption to obtain the first intermediate ciphertext cluster ;
[0066] The second cloud server uses a partial private key in the index key pair , for the new ciphertext cluster Perform partial decryption to obtain the second intermediate ciphertext cluster ; Based on the decryption sharing algorithm, the first intermediate ciphertext cluster and the second intermediate ciphertext cluster Processing is performed to obtain blind clusters ; Based on the encryption algorithm and the public key in the re-encryption key pair For the blinded cluster Encrypt and get the third intermediate ciphertext cluster ;
[0067] The first cloud server uses the public key in the re-encryption key pair and the fifth random number, to the third intermediate ciphertext cluster Perform deblinding processing to obtain the re-encrypted new ciphertext cluster.
[0068] The present invention also provides an image retrieval system supporting sharing of multiple data sources, comprising:
[0069] A trusted third party, at least one data owner, at least one query user, and a cloud server group, wherein the cloud server group includes a first cloud server and a second cloud server;
[0070] The trusted third party is used to generate an index key pair and an image key for each of the data owners, generate a retrieval key pair for each of the query users, generate a re-encryption key pair for the cloud server group, and distribute the keys;
[0071] Each of the data owners is configured to extract a feature vector of each image in the respective image set, and process each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; encrypt each of the images in the image set based on the image key to obtain a ciphertext image set, and send each of the ciphertext image set and the ciphertext cluster to the first cloud server;
[0072] The second cloud server is used to assist the first cloud server in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and construct a hierarchical index based on each of the new ciphertext clusters;
[0073] The query user is used to extract a feature vector of the query image and generate a query trapdoor, and send the query trapdoor to the first cloud server;
[0074] The second cloud server is further used to assist the first cloud server in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; and searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image;
[0075] The first cloud server is used to send the at least one ciphertext image and the identifier of the corresponding data owner to the query user;
[0076] The query user is further configured to decrypt each ciphertext image using the corresponding image key based on the identifier to obtain the queried image.
[0077] The image retrieval method and system supporting multi-data source sharing provided by the present invention generates and distributes keys for each retrieval member through a trusted third party. The data owner constructs a ciphertext cluster and a ciphertext image set based on the public key, image key and image set in the index key pair, and uploads them to the first cloud server; the cloud server group re-encrypts the ciphertext cluster and constructs a hierarchical index; the query user processes the query image using the public key in the retrieval key pair to obtain a query trapdoor, and uploads it to the first cloud server; the cloud server group re-encrypts the query trapdoor to obtain a new query trapdoor, and searches the hierarchical index based on the new query trapdoor to obtain search results; the query user obtains the image in plaintext based on the search results. The method supports multi-data source sharing, adopts the improved Paillier cryptographic system as the feature vector encryption method, ensures the security of privacy information during retrieval, and realizes complete encryption of calculations during the retrieval process. By adopting the ciphertext re-encryption protocol, the balance problem between rich image resources and different retrieval requests in the cloud environment is solved. By assigning different keys to different data owners and query users, any query user can retrieve image sets from multiple data owners stored on the cloud server. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0079] Figure 1 This is one of the flow charts of the image retrieval method supporting multi-data source sharing provided by the present invention.
[0080] Figure 2 This is the second flow chart of the image retrieval method supporting multi-data source sharing provided by the present invention.
[0081] Figure 3 It is a schematic diagram of the structure of the image retrieval system supporting multi-data source sharing provided by the present invention. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0083] Combine the following Figure 1-Figure 3The present invention describes an image retrieval method and system supporting sharing of multiple data sources.
[0084] In order to facilitate a clearer understanding of the technical solutions of the embodiments of the present application, some technical contents related to the embodiments of the present application are first introduced.
[0085] There are many image retrieval schemes, but they still have various deficiencies in specific application environments. First, some schemes select traditional manual features specified by the Multimedia Content Description Interface (MPEG-7) standard in the image preprocessing stage, resulting in low retrieval accuracy. In addition, although there are many encrypted image retrieval schemes that take into account the needs of multiple users and can well protect the privacy of query users, in actual applications, these schemes lack security designs for multiple data owners, and there is still room for further improvement in terms of privacy protection of data owners and comprehensiveness of retrieval results. Finally, the secure K-Nearest Neighbor (KNN) algorithm has been adopted by most retrieval schemes because of its advantages such as lower computational and communication overhead required for encrypted image feature vectors, but the algorithm has been proven to be unsafe and cannot resist security threats in outsourced environments.
[0086] Figure 1 This is one of the flowcharts of the image retrieval method supporting multiple data source sharing provided by the present invention. Figure 1 As shown, the method includes steps 101 to 107:
[0087] Step 101: A trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, generates a re-encryption key pair for a cloud server group, and distributes keys. The cloud server group includes a first cloud server and a second cloud server.
[0088] In practical applications, an image retrieval system that supports sharing of multiple data sources includes at least one data owner (Image Owners, IO), at least one search user (Search User, SU), a first cloud server (Cloud Sever A, CSA), a second cloud server (Cloud Sever B, CSB) and a trusted third party (TrustedAgent, TA). TA generates keys for other members in the image retrieval system and distributes the keys to the corresponding members.
[0089] Specifically, TA has a right to the data owner. Generate index key pair and the image key , for query users Generate a retrieval key pair , generate a re-encryption key pair for the cloud server ,in, Indicates Data owners, Indicates Furthermore, TA distributes the generated key through a secure channel.
[0090] Step 102: Each of the data owners extracts a feature vector of each image in their respective image set, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; encrypts each of the images in the image set based on the image key to obtain a ciphertext image set, and sends each of the ciphertext image sets and the ciphertext clusters to the first cloud server.
[0091] Specifically, the data owner Extract image set The feature vector of each image in the clustering algorithm is combined with its own public key Constructing ciphertext clusters ; Use the block cipher algorithm (SM4) to encrypt each image in the image set, that is, , get the ciphertext image set ,Will and Send to CSA.
[0092] Step 103: The second cloud server assists the first cloud server in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and constructs a hierarchical index based on each of the new ciphertext clusters.
[0093] Specifically, CSA, with the assistance of CSB, performs Re-encrypt and obtain the public key of the cloud server (Re-encrypt the public key in the key pair ) The encrypted new ciphertext cluster , and use each data owner The corresponding new ciphertext cluster Build a hierarchical index.
[0094] Step 104: The query user extracts a feature vector of the query image and generates a query trapdoor, and sends the query trapdoor to the first cloud server.
[0095] Specifically, query user Extract query image The feature vector of Constructing a query trapdoor ;Will Send to CSA.
[0096] Step 105: The second cloud server assists the first cloud server in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; and searches the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image.
[0097] Specifically, CSA, with the assistance of CSB, searches for trapdoors Re-encrypt to obtain the cloud server public key New query trap under , and then use this new query trap The hierarchical index is searched to obtain the k ciphertext images that are most similar to the query image, where k is a positive integer.
[0098] Step 106: The first cloud server sends the at least one ciphertext image and the identifier of its corresponding data owner to the query user.
[0099] Specifically, CSA finds k ciphertext images according to their corresponding image identifiers (Identity Document, ID), and returns the k ciphertext images and their corresponding data owner IDs as the retrieval result R to .
[0100] Step 107: The querying user decrypts each ciphertext image using the corresponding image key based on the identifier to obtain the queried image.
[0101] Specifically, According to the data owner ID in the search result R, find the corresponding image key ; Further, the block cipher algorithm is used to decrypt the ciphertext image in R, that is , and finally obtain the k plaintext images that are most similar to the query image.
[0102] The image retrieval method and system supporting multi-data source sharing provided by the present invention generates and distributes keys for each retrieval member through a trusted third party. The data owner constructs a ciphertext cluster and a ciphertext image set based on the public key, image key and image set in the index key pair, and uploads them to the first cloud server; the cloud server group re-encrypts the ciphertext cluster and constructs a hierarchical index; the query user processes the query image using the public key in the retrieval key pair to obtain a query trapdoor, and uploads it to the first cloud server; the cloud server group re-encrypts the query trapdoor to obtain a new query trapdoor, and searches the hierarchical index based on the new query trapdoor to obtain search results; the query user obtains the image in plaintext based on the search results. The method supports multi-data source sharing, adopts the improved Paillier cryptographic system as the feature vector encryption method, ensures the security of privacy information during retrieval, and realizes complete encryption of calculations during the retrieval process. By adopting the ciphertext re-encryption protocol, the balance problem between rich image resources and different retrieval requests in the cloud environment is solved. By assigning different keys to different data owners and query users, any query user can retrieve image sets from multiple data owners stored on the cloud server.
[0103] In one or more optional embodiments of the present invention, the trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, and generates a re-encryption key pair for the cloud server group, including:
[0104] For each of the data owners , each of the query users And any one of the cloud server groups performs the following steps:
[0105] The trusted third party Generate four different odd prime numbers , , , ,in, , , .
[0106] Calculate common parameters and private key , choose the order Generators of ,and ,in, , Representation model The multiplicative group under ; Generate a key pair .
[0107] The private key in the key pair Randomly split into two parts and get the split partial private key pair , and at the same time satisfy the following formula:
[0108]
[0109] At the same time, for the data owner Generate Image Key ,in Indicates Data owners;
[0110] Traverse each of the data owners , each of the query users and cloud server group, obtain the data owner Image key With index key pair , the query user Retrieval key pair , and the re-encryption key pair of the cloud server ,in, Indicates Query users.
[0111] Specifically, the process of generating a key pair includes the following steps:
[0112] S1.1, for each data owner , Query users And the cloud server group, TA based on security parameters Generate different odd prime numbers , , , .
[0113] in, , , .
[0114] S1.2, TA calculation , , and choose an order Generators of , so that satisfaction ,in, , Representation model The multiplication group under Finally, the key pair of each entity is obtained. .
[0115] S1.3, TA sends the private key of each entity's key pair Randomly split into two parts, we get , and the following conditions are met at the same time: .
[0116] It should be noted that for data owners , you also need to generate the corresponding image key .
[0117] Through this embodiment, a corresponding key can be generated for each image retrieval member, so that different data owners and query users hold different index key pairs and retrieval key pairs respectively, thereby ensuring the security of the data of each image retrieval member.
[0118] In one or more optional embodiments of the present invention, the key distribution includes:
[0119] The trusted third party sends the image key to with the index key pair Sent to the data owner , the image key Retrieve the key pair with Sent to the query user , the re-encryption key pair In and , the index key pair In and , and the retrieval key pair In and Send the re-encryption key pair to the first cloud server In and , the index key pair In and , and the retrieval key pair In and Send it to the second cloud server.
[0120] In actual applications, TA sends the image key to With index key pair Send to , the image key Retrieve key pair Send to , and re-encrypt the key pair With index key pair , retrieve the key pair Sent to the cloud server group, where , , , , and Assigned to CSA, , , , , and Assigned to CSB. In this way, not only can each member have its own key, but each cloud server can also have part of the key of the data owner and the query user, preventing the key of the data owner and the query user from being leaked on the cloud server side, while ensuring the security of image retrieval.
[0121] In one or more optional embodiments of the present invention, each of the data owners extracts a feature vector of each image in the respective image set, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster, including:
[0122] The data owner Extract the feature vectors of each image in the image set, and use a clustering algorithm to cluster the feature vectors of each image in the image set to obtain Cluster ;
[0123] Each of the clusters Each dimension of the data in each cluster center and its characteristic vector is processed in a unified format, and based on the public key in the index key pair , for the cluster The cluster centers and their characteristic vectors are encrypted after the unified format to obtain the ciphertext clusters. .
[0124] Specifically, the data owner The process of obtaining a ciphertext cluster includes the following steps:
[0125] S2.1, The Convolutional Neural Network (CNN) model is used to extract the feature vectors of each image in the image set, and then the mini batch K-means clustering algorithm is used to divide the feature vectors into Class, get Cluster ,in, Is a positive integer.
[0126] S2.2, Will Each dimension of data in each cluster center and its eigenvector is processed in a unified format: Will Multiply each dimension of each cluster center and its eigenvector by 1000, and round each dimension. right Each cluster center and its characteristic vector are encrypted to obtain the ciphertext cluster And send it to CSA. The encryption algorithm is as follows:
[0127]
[0128] in, Indicated by public key Each dimension of the encrypted vector is a random number , plaintext data , Indicates less than The set of positive integers.
[0129] Accordingly, encrypting each image in the image set based on the image key to obtain a ciphertext image set, and sending each ciphertext image set and the ciphertext cluster to the first cloud server may be: Adopt SM4 encryption algorithm and use image key Image set for itself Encrypt and generate a ciphertext image set , that is, for each image Encrypt to get the ciphertext image ,in, For image sets The images.
[0130] This embodiment adopts the improved Paillier cryptographic system as the feature vector encryption method, which can make the image features more secure.
[0131] In one or more optional embodiments of the present invention, the second cloud server assists the first cloud server to re-encrypt each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, including:
[0132] The first cloud server determines the index key pair based on the fifth random number and the public key of the index key pair. , for the ciphertext cluster Perform blinding processing to obtain a new ciphertext cluster , and based on the partial private key in the index key pair For the new ciphertext cluster Perform partial decryption to obtain the first intermediate ciphertext cluster ;
[0133] The second cloud server uses a partial private key in the index key pair , for the new ciphertext cluster Perform partial decryption to obtain the second intermediate ciphertext cluster ; Based on the decryption sharing algorithm, the first intermediate ciphertext cluster and the second intermediate ciphertext cluster Processing is performed to obtain blind clusters ; Based on the encryption algorithm and the public key in the re-encryption key pair For the blinded cluster Encrypt and get the third intermediate ciphertext cluster ;
[0134] The first cloud server uses the public key in the re-encryption key pair and the fifth random number, to the third intermediate ciphertext cluster Perform deblinding processing to obtain the re-encrypted new ciphertext cluster.
[0135] Specifically, the process of re-encrypting each ciphertext cluster includes the following steps:
[0136] S3.1, CSA for ciphertext clusters To blind, use The public key Encrypt a (fifth) random number ,calculate , thus obtaining a new ciphertext cluster ,in, Represents a ciphertext cluster Each dimension of each vector in Represents a ciphertext cluster Each dimension of each vector in . Then use Part of the private key right Perform partial decryption to obtain the first intermediate ciphertext cluster ,Will and Sent to CSB. Part of the decryption algorithm is as follows:
[0137]
[0138] in, Represents a ciphertext cluster Each dimension of each vector in Power, Indicates the use of part of the private key The intermediate ciphertext data obtained by decryption.
[0139] S3.2, CSB is used first Part of the private key right Perform partial decryption to obtain the second intermediate ciphertext cluster , the algorithm is as follows:
[0140]
[0141] in, Represents a ciphertext cluster Each dimension of each vector in Power, Indicates the use of part of the private key The intermediate ciphertext data obtained by decryption.
[0142] CSB , And the function , according to the decryption sharing algorithm, we get The blinded cluster The decryption sharing algorithm is as follows:
[0143]
[0144] Finally, CSB uses the public key according to the encryption algorithm in S2.2 right Encrypt to get the third intermediate ciphertext cluster , and Send to CSA.
[0145] S3.3, CSA uses the public key according to the above encryption algorithm Encrypted random numbers get , and removed by the deblinding algorithm The random number contained in , get the new ciphertext cluster , and finally complete the ciphertext cluster The de-blinding algorithm is as follows:
[0146] .
[0147] In this way, the re-encryption of the ciphertext cluster is completed without exposing the plaintext data, thereby improving the security of image retrieval.
[0148] In one or more optional embodiments of the present invention, constructing a hierarchical index based on each of the new ciphertext clusters includes:
[0149] Step 1: For each new ciphertext cluster, take the cluster center of the new ciphertext cluster as the representative vector, take the feature vector in the new ciphertext cluster as the stored data, and construct the initial node ,in, is the representative vector of the initial node, is the level number of the initial node in the hierarchical index, The data stored in the initial node is set to a first set whose content is empty;
[0150] Step 2: Add all the initial nodes to the first set;
[0151] Step 3: If the first set is not empty, randomly select an initial node from the first set and record it as the current node, and add the current node to the preset list; if the first set is empty, execute step 8;
[0152] Step 4: Calculate the distance between the current node and each other node, and record the other node with the smallest distance to the current node as the first node, where the other node is any initial node except the current node;
[0153] Step 5: If the first node already exists in the list, calculate the average vector of the representative vector of the current node and the representative vector of the first node, and use the average vector as the second node; record the current node as a child node of the second node, and record the first node as an auxiliary child node of the second node, and obtain a retrieval tree composed of all the initial nodes in the list, wherein the number of layers of the root node of the retrieval tree is the sum of the number of layers of the current node and one; return to step 3 to continue execution;
[0154] Step 6: If the first node already exists in the search tree, record the current node as a child node of the first node, and merge all the initial nodes in the list as subtrees into the search tree; return to step 3 to continue execution;
[0155] Step 7: If the first node exists in the first set, the first node is taken out from the first set, and the current node is recorded as a child node of the first node, and the first node is recorded as a new current node, and the process returns to step 4 to continue.
[0156] Step 8: If the number of all current root nodes is greater than 1, each root node is recorded as a new initial node, and the process returns to step 2 to continue execution; if the number of all current root nodes is 1, the hierarchical index construction is completed.
[0157] Specifically, the process of building a hierarchical index includes the following steps:
[0158] S4.1, CSA will The cluster center in is taken as the representative vector, and each The feature vector in is used as the stored data to obtain the initial node of the hierarchical index (with The ciphertext cluster in the hierarchical index is the initial node. Each initial node in the hierarchical index is recorded as ,in, represents the representative vector of the node, Indicates the level of the node in the hierarchical index, Indicates the data stored in the node (when When greater than 1, is NULL, that is, empty).
[0159] S4.2, CSA, with the assistance of CSB, constructs a hierarchical index according to the following rules:
[0160] (1) Add all initial nodes to a first set S1, which is initially empty;
[0161] (2) If S1 is not empty, randomly select a node from it and record it as the current node C, and add it to a list L1. Otherwise, execute step (7);
[0162] (3) Find the node with the shortest distance to C among the initial nodes, denoted as P.
[0163] (4) If P already exists in the list L1, calculate the average vector of the representative vectors in the two nodes C and P, and use this as the new representative vector to generate a new node R, and record C as the child node of R, and record P as the auxiliary child node of R. At this point, the nodes in the list L1 form a tree, and the root node of the tree is R, where , return to step (2);
[0164] (5) If the node P already exists in a tree T, add C to the child nodes of P, thereby merging the nodes in L1 into T as a subtree, and return to step (2);
[0165] (6) If node P is an active node, that is, it exists in S1, then P is taken out of S1, and C is recorded as the child node of P, and P is recorded as the new current node C, and return to step (3);
[0166] (7) If the number of root nodes obtained is 1, it means that the hierarchical index is constructed, and the node is the root node of the hierarchical index. Otherwise, all the root nodes obtained are recorded as new initial nodes and return to step (1).
[0167] The hierarchical index provided in this embodiment aggregates the re-encrypted clusters from different data owners according to certain rules, thereby achieving lower retrieval computing overhead and timely user query response.
[0168] In one or more optional embodiments of the present invention, the calculating the distance between the current node and each other node includes:
[0169] For any of the other nodes, perform the following steps:
[0170] The first cloud server calculates the difference of ciphertext data at the same position in each dimension between the representative vector of the current node and the representative vector of the other nodes to obtain a ciphertext difference;
[0171] Based on the blinding method, a first random number is introduced into the ciphertext difference to obtain a first ciphertext difference, and a second random number is introduced into the ciphertext difference to obtain a second ciphertext difference, wherein the first random number is different from the second random number; the first ciphertext difference and the second ciphertext difference are partially decrypted to obtain a first intermediate ciphertext difference and a second intermediate ciphertext difference;
[0172] The second cloud server processes the first ciphertext difference and the second ciphertext difference respectively based on a partial decryption algorithm to obtain a third intermediate ciphertext difference and a fourth intermediate ciphertext difference; processes the first intermediate ciphertext difference and the third intermediate ciphertext difference based on a decryption sharing algorithm to obtain first blinded data, and processes the second intermediate ciphertext difference and the fourth intermediate ciphertext difference to obtain second blinded data;
[0173] Calculating a first ciphertext product of the first blinded data and the second blinded data, and encrypting the first ciphertext product according to an encryption algorithm based on a public key in the re-encryption key pair to obtain a second ciphertext product;
[0174] The first cloud server encrypts the product of the first random number and the second random number based on the encryption algorithm to obtain a third ciphertext product;
[0175] Based on the second ciphertext product, the third ciphertext product and the ciphertext difference, the ciphertext of the square of the Euclidean distance between the current node and the other nodes is obtained; and the ciphertext of the square of the Euclidean distance between the current node and the other nodes is used as the distance between the current node and each other node.
[0176] Specifically, S4.3 calculates the distance between two nodes, including the following steps:
[0177] (1) CSA for two node representative vectors , The ciphertext data at the same position in each dimension , ,according to Calculate ciphertext difference ;
[0178] (2) CSA first uses the blinding method in S3.1 to Introduce two different random numbers , get two new ciphertext data , ,in is the first random number, is the second random number; then the partial decryption algorithm in S3.1 is used to decrypt it, and the first intermediate ciphertext difference is obtained. and the second intermediate ciphertext difference ,Will , , and Send to CSB;
[0179] (3) CSB uses the partial decryption algorithm and the decryption sharing algorithm in S3.2. , Get the third intermediate ciphertext difference and the fourth intermediate ciphertext ,according to and , and Get the first blinded data after blinding and the second blinded data ;
[0180] (4) CSB calculates the first blind data and the second blinded data The product of , and then use the public key according to the encryption algorithm in S2.2 encryption get , recorded as the second ciphertext product ,Will Send to CSA;
[0181] (5) CSA uses the same encryption method to encrypt and The product of , and then calculate separately , , , thus obtaining ;
[0182] (6) CSA calculation ,get , The square of the Euclidean distance of the corresponding plaintext vector in the public key The following ciphertext, where Denotes the dimension of the representation vector.
[0183] It should be noted that since the distance between the current node and each other node is the distance under ciphertext, and when comparing the distances, what needs to be compared is the distance under plaintext, therefore, in the process of executing the step of "calculating the distance between the current node and each other node, and recording the other node with the smallest distance to the current node as the first node", it is necessary to adopt the method of plaintext comparison under ciphertext to compare the distance between the current node and each other node to determine the other node with the smallest distance to the current node.
[0184] Specifically, the method for comparing plaintext under ciphertext is usually to compare two data under ciphertext. , Corresponding to the first plaintext and the second plaintext The size comparison process includes:
[0185] (1) CSA uses the public key according to the encryption algorithm in S2.2 Encrypt the integer 1 to get the ciphertext of 1 , and then compare them with the ciphertext data and ciphertext data Multiply the square of to get the first ciphertext data and the second ciphertext data ;
[0186] (2) CSA selects two random numbers (Must meet and ), use the same encryption algorithm to get the ciphertext random number ,in, is the third random number, is the fourth random number, is the first set length, is a public parameter, , Respectively , The length corresponding to binary;
[0187] (3) CSA is calculated first , recorded as the third ciphertext data , and then use the partial decryption algorithm in S3.1 to decrypt it and get the fourth ciphertext data ,Will , Send to CSB;
[0188] (4) CSB uses the partial decryption algorithm and the decryption sharing algorithm in S3.2. Get the fifth ciphertext data ,according to , Decrypt to get the plaintext data ;
[0189] (5) CSB is based on the binary length of the plaintext data and the second set length For comparison, if Get the result , otherwise we get ,Will Send to CSA;
[0190] (6) CSA based on Make a judgment, if It indicates ,like It indicates .
[0191] In this embodiment, the two ciphertext data , They are the two distances between the current node and each other node; the first plaintext and the second plaintext are the plaintext distances corresponding to these two distances respectively.
[0192] In this way, the size of the squares of two Euclidean distances can be compared without revealing the Euclidean distance of the plaintext.
[0193] In one or more optional embodiments of the present invention, the process of calculating the average vector of the representative vector of the current node and the representative vector of the first node, that is, calculating the average vector of the representative vectors in the two nodes C and P, may be as follows:
[0194] (1) CSA first uses the blinding method in S3.1 to add the representative vector of the current node and the representative vector of the first node The ciphertext data at the same position in each dimension , Introducing the same random number , get two new ciphertext data , , and then use the partial decryption algorithm in S3.1 to decrypt it and obtain the intermediate ciphertext data , ,Will , , , Send to CSB;
[0195] (2) CSB uses the partial decryption algorithm and the decryption sharing algorithm in S3.2. , Get the intermediate ciphertext data , ,according to and , and Obtaining blinded data , ;
[0196] (3) CSB calculation ,in, Express Round down and use the public key according to the encryption algorithm in S2.2 encryption get ,Will Send to CSA;
[0197] (4) CSA based on and the blinding process used in step (1) , calculated , The mean of the public key The following ciphertext , thus obtaining the vector , The average vector in the public key The ciphertext vector .
[0198] In one or more optional embodiments of the present invention, the query user extracts a feature vector of the query image and generates a query trapdoor, and sends the query trapdoor to the first cloud server, which specifically includes the following steps:
[0199] S5.1, Extract query images using the CNN model in S2.1 The feature vector of , and multiply each dimension of the data by 1000, and then round each dimension of the data to the integer, recorded as the query vector ;
[0200] S5.2, According to the encryption algorithm in S2.2, use the assigned public key Encrypt the query vector to obtain the query trapdoor And send it to CSA.
[0201] In one or more optional embodiments of the present invention, searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image includes:
[0202] Step I: Create an empty second set and initialize the root node of the hierarchical index as the entry node, the second set is used to record the minimum distance between the new query trapdoor and each layer node of the hierarchical index under the current path during the search process;
[0203] Step II: Obtain all descendant nodes of the next layer of the entry node, calculate the distance between the new query trapdoor and each descendant node, and determine the minimum distance between the new query trapdoor and each descendant node;
[0204] Step III: If the second set is empty, all the descendant nodes are recorded as target nodes, and the number of layers and the minimum distance of each target node are added to the second set, and step VI is executed;
[0205] Step IV: If the second set is not empty, filter out the descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, and record them as target nodes; if there are no descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, record the auxiliary child nodes of the entry node and the child nodes of the next layer as target nodes;
[0206] Step V: Based on a size comparison algorithm, select the smaller of the minimum distance and the minimum distance corresponding to the number of layers of the entry node in the second set, and record it as the minimum distance corresponding to the number of layers of the target node;
[0207] Step VI: If the target node has a node with a layer number greater than 2, the target node is used as a new entry node and the process returns to step II to continue; otherwise, the target node is added to the intermediate node list;
[0208] Step VII: When the intermediate node list is no longer updated, obtain the descendant nodes of all nodes in the intermediate node list, calculate the distance between the new query trapdoor and each descendant node, and select a set number of descendant nodes with the smallest distance to the new query trapdoor, and record them as the final nodes obtained by searching the hierarchical index, where the set number is the number of data owners;
[0209] Step VIII: searching for the feature vector corresponding to the new query trapdoor from the feature vectors stored in the final node, and using the ciphertext image corresponding to the feature vector corresponding to the new query trapdoor as the ciphertext image corresponding to the query image.
[0210] Specifically, the process of querying a hierarchical index includes the following steps:
[0211] S6.1, CSA uses the blinding method in S3.1 to first query the trapdoor Blind it, and then collaborate with CSB to re-encrypt the trapdoor to get the new query trapdoor ;
[0212] S6.2, CSA starts from the root node of the hierarchical index and searches in a top-down, depth-first manner. The specific rules are as follows:
[0213] (1) Create a second set S2 to record the query traps in the current path during the search process The minimum distance to each layer node, where the initial S2 is empty; and the root node of the hierarchical index is initialized as the entry node;
[0214] (2) Obtain all descendant nodes of the next layer of the entry node and calculate according to the algorithm in S4.3 The distances to each descendant node, and compare the distances to find the minimum distance;
[0215] (3) If S2 is an empty set, all descendant nodes are recorded as target nodes, and the number of layers and the minimum distance of the target node are added to S2, and step (6) is executed;
[0216] (4) If S2 is not empty, use the size comparison algorithm in S4.4 to filter out descendant nodes whose distance is less than the minimum distance of the layer number of the entry node recorded in S2, and record the descendant nodes that meet the conditions as target nodes. If there are no descendant nodes that meet the conditions, the auxiliary child nodes of the entry node and the child nodes of the next layer are recorded as target nodes;
[0217] (5) According to the size comparison algorithm in S4.4, select the smaller one between the above minimum distance and the minimum distance corresponding to the number of layers of the entry node recorded in S2, and record it as the minimum distance corresponding to the number of layers of the target node;
[0218] (6) If the number of layers of the target node is greater than 2, use the node as the new entry node and return to step (2); otherwise, add the node to the intermediate node list L2, i.e., the second list;
[0219] (7) When L2 is no longer updated, obtain the lower-level descendant nodes of all nodes in L2 and calculate according to the algorithm in S4.3 The distance between each descendant node and the node with the smallest distance is selected using the size comparison algorithm in S4.4. descendant nodes, i.e., the final nodes obtained by searching the hierarchical index, where Indicates the number of data owners.
[0220] 6.3, CSA is calculated according to the algorithm in S4.3 With each final node The distances of all stored ciphertext feature vectors (i.e., the feature vectors stored in the final node) are calculated, and the size comparison algorithm in S4.4 is used to screen out the k ciphertext feature vectors with the smallest distances, and k ciphertext images are found according to their corresponding image IDs.
[0221] In this way, retrieval is convenient and fast, and the accuracy and completeness of the retrieval results can be guaranteed.
[0222] In one or more optional embodiments of the present invention, the distances between the new query trapdoor and each of the descendant nodes are all ciphertext distances; and determining the minimum distance between the new query trapdoor and each of the descendant nodes includes:
[0223] Comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes, and determining the minimum plaintext distance among the plaintext distances;
[0224] Taking the ciphertext distance corresponding to the minimum plaintext distance as the minimum distance;
[0225] The step of comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes includes:
[0226] For any two ciphertext distances between the new query trapdoor and each of the descendant nodes, the following steps are performed, where the two ciphertext distances include a first ciphertext distance and a second ciphertext distance:
[0227] The first cloud server encrypts one based on the encryption algorithm and the public key in the re-encryption key pair to obtain a ciphertext of one;
[0228] Multiply the ciphertext of the first ciphertext by the square of the first ciphertext distance and the second ciphertext distance respectively to obtain first ciphertext data and second ciphertext data;
[0229] Generate a third random number and a fourth random number, and encrypt the fourth random number to obtain a ciphertext random number; the third random number is greater than the fourth random number, and the binary length of the third random number is less than a first set length, the first set length is one quarter of the binary length of a public parameter, and the public parameter is a part of a public key in a re-encryption key pair;
[0230] Calculate based on the ciphertext random number, the third random number, the first ciphertext data and the second ciphertext data to obtain third ciphertext data; and partially decrypt the third ciphertext data to obtain fourth ciphertext data;
[0231] The second cloud server partially decrypts the third ciphertext data based on a partial decryption algorithm to obtain fifth ciphertext data, and decrypts the fourth ciphertext data and the fifth ciphertext data based on a decryption sharing algorithm to obtain plaintext data;
[0232] If the binary length of the plaintext data is greater than the second set length, the plaintext distance corresponding to the first ciphertext distance is less than the plaintext distance corresponding to the second ciphertext distance; if the binary length of the plaintext data is less than or equal to the second set length, the plaintext distance corresponding to the first ciphertext distance is greater than or equal to the plaintext distance corresponding to the second ciphertext distance, and the second set length is half of the binary length of the public parameter.
[0233] In practical applications, since the distances between the new query trapdoor and each descendant node are all ciphertext distances, when comparing the sizes, it is necessary to use the method of plaintext comparison under ciphertext to compare the sizes of the plaintext distances corresponding to each ciphertext distance, so as to determine the minimum plaintext distance among the plaintext distances. Furthermore, the ciphertext distance corresponding to the minimum plaintext distance is taken as the minimum distance, and the descendant node corresponding to the minimum distance is determined as the current descendant node.
[0234] It should be noted that the method of comparing plaintext under ciphertext is used to compare the distances of each ciphertext. , the second ciphertext distance Corresponding plaintext distance Distance from plaintext The size comparison process includes:
[0235] (1) CSA uses the public key according to the encryption algorithm in S2.2 Encrypt the integer 1 to get the ciphertext of 1 , and then separate them from the first ciphertext and the second ciphertext distance Multiply the square of to get the first ciphertext data and the second ciphertext data ;
[0236] (2) CSA selects two random numbers (Must meet and ), use the same encryption algorithm to get the ciphertext random number ,in, is the third random number, is the fourth random number, is the first set length, is a public parameter, , Respectively , The length corresponding to binary;
[0237] (3) CSA is calculated first , recorded as the third ciphertext data , and then use the partial decryption algorithm in S3.1 to decrypt it and get the fourth ciphertext data ,Will , Send to CSB;
[0238] (4) CSB uses the partial decryption algorithm and the decryption sharing algorithm in S3.2. Get the fifth ciphertext data ,according to , Decrypt to get the plaintext data ;
[0239] (5) CSB is based on the binary length of the plaintext data and the second set length For comparison, if Get the result , otherwise we get ,Will Send to CSA;
[0240] (6) CSA based on Make a judgment, if It indicates ,like It indicates .
[0241] S4.4 is applied in this embodiment, the two data in the ciphertext are , are the first ciphertext distance and the second ciphertext distance respectively; the first plaintext and the second plaintext They are the plaintext corresponding to the first ciphertext distance and the plaintext corresponding to the second ciphertext distance respectively.
[0242] In this way, the size of the squares of two Euclidean distances can be compared without revealing the Euclidean distance of the plaintext.
[0243] The image retrieval method supporting the sharing of multiple data sources provided by the present invention supports the sharing of multiple data sources, adopts the improved Paillier cryptographic system as the feature vector encryption method, ensures the security of privacy information during retrieval, and realizes the complete encryption of calculations during the retrieval process. By adopting the ciphertext re-encryption protocol, the balance problem between the rich image resources in the cloud environment and different retrieval requests is solved. By assigning different keys to different data owners and query users, any query user can retrieve the image sets from multiple data owners stored on the cloud server. In addition, by constructing a hierarchical index, the re-encrypted clusters from different data owners are aggregated according to certain rules, thereby achieving low retrieval calculation overhead and timely user query response.
[0244] Combine the following Figure 2 The image retrieval method supporting multiple data source sharing provided by the present invention is further described. Figure 2 This is the second flow chart of the image retrieval method supporting multi-data source sharing provided by the present invention, comprising the following steps:
[0245] Step 201: The trusted third party generates an image key and an index key pair for the data owner, generates a retrieval key pair for the query user, and generates a re-encryption key pair for the cloud server group;
[0246] Step 202: Each data owner extracts a feature vector of each image in the image set, clusters the feature vectors to obtain a cluster, uses an improved Paillier cryptographic system, uses the public key in the index key pair to encrypt the cluster to obtain a ciphertext cluster; and uses the image key to encrypt the image set to obtain a ciphertext image set;
[0247] Step 203: The cloud server group adopts a ciphertext re-encryption protocol, uses the public key in the re-encryption key pair to re-encrypt the ciphertext cluster to obtain a new ciphertext cluster, and constructs a hierarchical index based on the new ciphertext cluster;
[0248] Step 204: the query user extracts a feature vector of the query image, and uses an improved Paillier cryptographic system to encrypt the feature vector using the public key in the retrieval key pair to obtain a query trapdoor;
[0249] Step 205: The cloud server group adopts a ciphertext re-encryption protocol, uses the public key in the re-encryption key pair to re-encrypt the query trapdoor to obtain a new query trapdoor, and searches the hierarchical index according to the new query trapdoor to obtain similar ciphertext images as retrieval results;
[0250] Step 206: The querying user uses the image key to decrypt the ciphertext image to obtain a plaintext image.
[0251] The image retrieval system supporting multiple data source sharing provided by the present invention is described below. The image retrieval system supporting multiple data source sharing described below and the image retrieval method supporting multiple data source sharing described above can be referred to each other. Figure 3 As shown, Figure 3 The structure diagram of the image retrieval system supporting multiple data source sharing provided by the present invention includes:
[0252] A trusted third party 301, at least one data owner 302, at least one query user 303, and a cloud server group, wherein the cloud server group includes a first cloud server 304 and a second cloud server 305;
[0253] The trusted third party 301 is used to generate an index key pair and an image key for each of the data owners 302, generate a retrieval key pair for each of the query users 303, generate a re-encryption key pair for the cloud server group, and distribute keys;
[0254] Each of the data owners 302 is configured to extract a feature vector of each image in the respective image set, and process each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; encrypt each of the images in the image set based on the image key to obtain a ciphertext image set, and send each of the ciphertext image set and the ciphertext cluster to the first cloud server 304;
[0255] The second cloud server 305 is used to assist the first cloud server 304 in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and construct a hierarchical index based on each of the new ciphertext clusters;
[0256] The query user 303 is used to extract the feature vector of the query image and generate a query trapdoor, and send the query trapdoor to the first cloud server 304;
[0257] The second cloud server 305 is further configured to assist the first cloud server 304 in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; and search the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image;
[0258] The first cloud server 304 is used to send the at least one ciphertext image and the identifier of the corresponding data owner 302 to the query user 303;
[0259] The query user 303 is further configured to decrypt each ciphertext image using the corresponding image key based on the identifier to obtain the queried image.
[0260] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0261] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image retrieval method supporting sharing of multiple data sources, characterized in that: include: The trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, generates a re-encryption key pair for a cloud server group, and distributes the keys, wherein the cloud server group includes a first cloud server and a second cloud server; Each of the data owners extracts a feature vector of each image in the respective image set, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; Encrypting each of the images in the image set based on the image key to obtain a ciphertext image set, and sending each of the ciphertext image sets and the ciphertext cluster to the first cloud server; The second cloud server assists the first cloud server in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and constructs a hierarchical index based on each of the new ciphertext clusters; The query user extracts a feature vector of the query image and generates a query trapdoor, and sends the query trapdoor to the first cloud server; The second cloud server assists the first cloud server in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; Searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image; The first cloud server sends the at least one ciphertext image and the identifier of the corresponding data owner to the query user; The querying user decrypts each ciphertext image using the corresponding image key based on the identifier to obtain the queried image; The step of constructing a hierarchical index based on each of the new ciphertext clusters includes: Step 1: For each new ciphertext cluster, take the cluster center of the new ciphertext cluster as the representative vector, take the feature vector in the new ciphertext cluster as the stored data, and construct the initial node ,in, is the representative vector of the initial node, is the level number of the initial node in the hierarchical index, The data stored in the initial node is set to a first set whose content is empty; Step 2: Add all the initial nodes to the first set; Step 3: If the first set is not empty, randomly select an initial node from the first set and record it as the current node, and add the current node to the preset list; if the first set is empty, execute step 8; Step 4: Calculate the distance between the current node and each other node, and record the other node with the smallest distance to the current node as the first node, where the other node is any initial node except the current node; Step 5: If the first node already exists in the list, calculate the average vector of the representative vector of the current node and the representative vector of the first node, and use the average vector as the second node; record the current node as a child node of the second node, and record the first node as an auxiliary child node of the second node, and obtain a retrieval tree composed of all the initial nodes in the list, wherein the number of layers of the root node of the retrieval tree is the sum of the number of layers of the current node and one; return to step 3 to continue execution; Step 6: If the first node already exists in the search tree, record the current node as a child node of the first node, and merge all the initial nodes in the list as subtrees into the search tree; return to step 3 to continue execution; Step 7: If the first node exists in the first set, the first node is taken out from the first set, and the current node is recorded as a child node of the first node, and the first node is recorded as a new current node, and the process returns to step 4 to continue. Step 8: If the number of all current root nodes is greater than 1, each root node is recorded as a new initial node, and the process returns to step 2 to continue execution; if the number of all current root nodes is 1, the hierarchical index construction is completed.
2. The image retrieval method supporting multiple data source sharing according to claim 1, characterized in that: The calculating the distance between the current node and each other node includes: For any of the other nodes, perform the following steps: The first cloud server calculates the difference of ciphertext data at the same position in each dimension between the representative vector of the current node and the representative vectors of the other nodes to obtain a ciphertext difference; Based on the blinding method, a first random number is introduced into the ciphertext difference to obtain a first ciphertext difference, and a second random number is introduced into the ciphertext difference to obtain a second ciphertext difference, wherein the first random number is different from the second random number; the first ciphertext difference and the second ciphertext difference are partially decrypted to obtain a first intermediate ciphertext difference and a second intermediate ciphertext difference; The second cloud server processes the first ciphertext difference and the second ciphertext difference respectively based on a partial decryption algorithm to obtain a third intermediate ciphertext difference and a fourth intermediate ciphertext difference; processes the first intermediate ciphertext difference and the third intermediate ciphertext difference based on a decryption sharing algorithm to obtain first blinded data, and processes the second intermediate ciphertext difference and the fourth intermediate ciphertext difference to obtain second blinded data; Calculating a first ciphertext product of the first blinded data and the second blinded data, and encrypting the first ciphertext product according to an encryption algorithm based on a public key in the re-encryption key pair to obtain a second ciphertext product; The first cloud server encrypts the product of the first random number and the second random number based on the encryption algorithm to obtain a third ciphertext product; Based on the second ciphertext product, the third ciphertext product and the ciphertext difference, the ciphertext of the square of the Euclidean distance between the current node and the other nodes is obtained; and the ciphertext of the square of the Euclidean distance between the current node and the other nodes is used as the distance between the current node and each other node.
3. The image retrieval method supporting multiple data source sharing according to claim 1, characterized in that: The step of searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image includes: Step I: Create an empty second set and initialize the root node of the hierarchical index as the entry node, the second set is used to record the minimum distance between the new query trapdoor and each layer node of the hierarchical index under the current path during the search process; Step II: Obtain all descendant nodes of the next layer of the entry node, calculate the distance between the new query trapdoor and each descendant node, and determine the minimum distance between the new query trapdoor and each descendant node; Step III: If the second set is empty, all the descendant nodes are recorded as target nodes, and the number of layers and the minimum distance of each target node are added to the second set, and step VI is executed; Step IV: If the second set is not empty, filter out the descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, and record them as target nodes; if there are no descendant nodes whose distance is less than the minimum distance of the number of layers of the entry node in the second set, record the auxiliary child nodes of the entry node and the child nodes of the next layer as target nodes; Step V: Based on a size comparison algorithm, select the smaller of the minimum distance and the minimum distance corresponding to the number of layers of the entry node in the second set, and record it as the minimum distance corresponding to the number of layers of the target node; Step VI: If the target node has a node with a layer number greater than 2, the target node is used as a new entry node and the process returns to step II to continue; otherwise, the target node is added to the intermediate node list; Step VII: When the intermediate node list is no longer updated, obtain the descendant nodes of all nodes in the intermediate node list, calculate the distance between the new query trapdoor and each descendant node, and select a set number of descendant nodes with the smallest distance to the new query trapdoor, and record them as the final nodes obtained by searching the hierarchical index, where the set number is the number of data owners; Step VIII: searching for the feature vector corresponding to the new query trapdoor from the feature vectors stored in the final node, and using the ciphertext image corresponding to the feature vector corresponding to the new query trapdoor as the ciphertext image corresponding to the query image.
4. The image retrieval method supporting multiple data source sharing according to claim 3, characterized in that: The distances between the new query trapdoor and each of the descendant nodes are all ciphertext distances; The determining of the minimum distance between the new query trapdoor and each of the descendant nodes comprises: Comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes, and determining the minimum plaintext distance among the plaintext distances; Taking the ciphertext distance corresponding to the minimum plaintext distance as the minimum distance; The step of comparing the plaintext distances corresponding to the ciphertext distances between each of the new query trapdoors and each of the descendant nodes includes: For any two ciphertext distances between the new query trapdoor and each of the descendant nodes, the following steps are performed, where the two ciphertext distances include a first ciphertext distance and a second ciphertext distance: The first cloud server encrypts one based on the encryption algorithm and the public key in the re-encryption key pair to obtain a ciphertext of one; Multiply the ciphertext of the first ciphertext by the square of the first ciphertext distance and the second ciphertext distance respectively to obtain first ciphertext data and second ciphertext data; Generate a third random number and a fourth random number, and encrypt the fourth random number to obtain a ciphertext random number; the third random number is greater than the fourth random number, and the binary length of the third random number is less than a first set length, the first set length is one quarter of the binary length of a public parameter, and the public parameter is a part of a public key in a re-encryption key pair; Calculate based on the ciphertext random number, the third random number, the first ciphertext data and the second ciphertext data to obtain third ciphertext data; and partially decrypt the third ciphertext data to obtain fourth ciphertext data; The second cloud server partially decrypts the third ciphertext data based on a partial decryption algorithm to obtain fifth ciphertext data, and decrypts the fourth ciphertext data and the fifth ciphertext data based on a decryption sharing algorithm to obtain plaintext data; If the binary length of the plaintext data is greater than the second set length, the plaintext distance corresponding to the first ciphertext distance is less than the plaintext distance corresponding to the second ciphertext distance; if the binary length of the plaintext data is less than or equal to the second set length, the plaintext distance corresponding to the first ciphertext distance is greater than or equal to the plaintext distance corresponding to the second ciphertext distance, and the second set length is half of the binary length of the public parameter.
5. The image retrieval method supporting multiple data source sharing according to claim 1, characterized in that: The trusted third party generates an index key pair and an image key for each data owner, generates a retrieval key pair for each query user, and generates a re-encryption key pair for the cloud server group, including: For each of the data owners , each of the query users And any one of the cloud server groups performs the following steps: The trusted third party Generate four different odd prime numbers , , , ,in, , , , Calculate common parameters and private key , choose the order Generators of ,and ,in, , Representation model The multiplicative group under ; Generate a key pair , The private key in the key pair Randomly split into two parts and get the split partial private key pair , and at the same time satisfy the following formula: ; Wherein, m is 1 or 2; At the same time, for the data owner Generate Image Key ,in Indicates Data owners; Traverse each of the data owners , each of the query users and cloud server group, obtain the data owner Image key With index key pair , the query user Retrieval key pair , and the re-encryption key pair of the cloud server ,in, Indicates Query users.
6. The image retrieval method supporting multiple data source sharing according to claim 5, characterized in that: The key distribution comprises: The trusted third party sends the image key to with the index key pair Sent to the data owner , the image key Retrieve the key pair with Sent to the query user , the re-encryption key pair In and , the index key pair In and , and the retrieval key pair In and Send the re-encryption key pair to the first cloud server In and , the index key pair In and , and the retrieval key pair In and Send it to the second cloud server.
7. The image retrieval method supporting multiple data source sharing according to claim 6, characterized in that: Each of the data owners extracts a feature vector of each image in the respective image set, and processes each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster, including: The data owner Extract the feature vectors of each image in the image set, and use a clustering algorithm to cluster the feature vectors of each image in the image set to obtain Cluster ; Each of the clusters Each dimension of the data in each cluster center and its characteristic vector is processed in a unified format, and based on the public key in the index key pair , for the cluster The cluster centers and their characteristic vectors are encrypted after the unified format to obtain the ciphertext clusters. .
8. The image retrieval method supporting multiple data source sharing according to claim 7, characterized in that: The second cloud server assists the first cloud server to re-encrypt each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, including: The first cloud server determines the index key pair based on the fifth random number and the public key of the index key pair. , for the ciphertext cluster Perform blinding processing to obtain a new ciphertext cluster , and based on the partial private key in the index key pair For the new ciphertext cluster Perform partial decryption to obtain the first intermediate ciphertext cluster ; The second cloud server uses a partial private key in the index key pair , for the new ciphertext cluster Perform partial decryption to obtain the second intermediate ciphertext cluster ; Based on the decryption sharing algorithm, the first intermediate ciphertext cluster and the second intermediate ciphertext cluster Processing is performed to obtain blind clusters ; Based on the encryption algorithm and the public key in the re-encryption key pair For the blinded cluster Encrypt and get the third intermediate ciphertext cluster ; The first cloud server uses the public key in the re-encryption key pair and the fifth random number, to the third intermediate ciphertext cluster Perform deblinding processing to obtain the re-encrypted new ciphertext cluster.
9. An image retrieval system supporting sharing of multiple data sources, characterized in that: include: A trusted third party, at least one data owner, at least one query user, and a cloud server group, wherein the cloud server group includes a first cloud server and a second cloud server; The trusted third party is used to generate an index key pair and an image key for each of the data owners, generate a retrieval key pair for each of the query users, generate a re-encryption key pair for the cloud server group, and distribute the keys; Each of the data owners is configured to extract a feature vector of each image in the respective image set, and process each of the feature vectors based on a clustering algorithm and a public key in the index key pair to obtain a ciphertext cluster; Encrypting each of the images in the image set based on the image key to obtain a ciphertext image set, and sending each of the ciphertext image sets and the ciphertext cluster to the first cloud server; The second cloud server is used to assist the first cloud server in re-encrypting each of the ciphertext clusters based on the public key in the re-encryption key pair to obtain a new ciphertext cluster, and construct a hierarchical index based on each of the new ciphertext clusters; The query user is used to extract a feature vector of the query image and generate a query trapdoor, and send the query trapdoor to the first cloud server; The second cloud server is further used to assist the first cloud server in re-encrypting the query trapdoor based on the public key in the re-encryption key pair to obtain a new query trapdoor; Searching the hierarchical index based on the new query trapdoor to obtain at least one ciphertext image corresponding to the query image; The first cloud server is used to send the at least one ciphertext image and the identifier of the corresponding data owner to the query user; The query user is further configured to decrypt each ciphertext image using the corresponding image key based on the identifier to obtain the queried image; The step of constructing a hierarchical index based on each of the new ciphertext clusters includes: Step 1: For each new ciphertext cluster, take the cluster center of the new ciphertext cluster as the representative vector, take the feature vector in the new ciphertext cluster as the stored data, and construct the initial node ,in, is the representative vector of the initial node, is the level number of the initial node in the hierarchical index, The data stored in the initial node is set to a first set whose content is empty; Step 2: Add all the initial nodes to the first set; Step 3: If the first set is not empty, randomly select an initial node from the first set and record it as the current node, and add the current node to the preset list; if the first set is empty, execute step 8; Step 4: Calculate the distance between the current node and each other node, and record the other node with the smallest distance to the current node as the first node, where the other node is any initial node except the current node; Step 5: If the first node already exists in the list, calculate the average vector of the representative vector of the current node and the representative vector of the first node, and use the average vector as the second node; record the current node as a child node of the second node, and record the first node as an auxiliary child node of the second node, and obtain a retrieval tree composed of all the initial nodes in the list, wherein the number of layers of the root node of the retrieval tree is the sum of the number of layers of the current node and one; return to step 3 to continue execution; Step 6: If the first node already exists in the search tree, record the current node as a child node of the first node, and merge all the initial nodes in the list as subtrees into the search tree; return to step 3 to continue execution; Step 7: If the first node exists in the first set, the first node is taken out from the first set, and the current node is recorded as a child node of the first node, and the first node is recorded as a new current node, and the process returns to step 4 to continue. Step 8: If the number of all current root nodes is greater than 1, each root node is recorded as a new initial node, and the process returns to step 2 to continue execution; if the number of all current root nodes is 1, the hierarchical index construction is completed.
Citation Information
Patent Citations
Fuzzy keyword search method oriented to multi-server and multi-user
CN108062485A
Ciphertext image retrieval method and system under a cloud environment
CN108959478A