Secure verifiable cross-modal retrieval method
By constructing verifiable B+ tree index structure and feature extraction, the authenticity and efficiency of search results in cross-modal retrieval are solved, and efficient and reliable retrieval of cross-modal data is achieved.
Patent Information
- Application Number
- CN202510226054.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art cannot guarantee the authenticity and efficiency of the search results in cross-modal retrieval, especially when multimodal data is outsourced to the cloud platform, data transmission problems caused by data tampering and network instability.
Build a verifiable B+ tree index structure, generate high-dimensional vectors through feature extraction, use hash functions and digital signatures to ensure data integrity, and use KNN search algorithms during query to improve retrieval efficiency and accuracy.
The efficiency of cross-modal retrieval and the authenticity of the results are achieved. The search space is reduced through clustering and indexing structures, and signature verification is used to ensure that the data has not been tampered with, which improves the query efficiency and the reliability of the results.
Smart Images

Figure CN120296191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to a secure and verifiable cross-modal retrieval method. Background Art
[0002] In many applications, data can have multiple modalities. For example, in addition to image content, there is also text information, such as tags of images in Flickr and many other social websites; this kind of data is usually called multi-modal data. Therefore, cross-modal retrieval technology is needed to effectively organize and manage these massive and heterogeneous multi-modal data, mine the potential associations between different modal data, help users quickly and accurately find valuable information from complex data, and avoid the trouble brought by information overload. Specifically, cross-modal retrieval aims to retrieve semantically relevant samples from another modality through a query request from one modality. Specifically, cross-modal retrieval can query images through text, or query text through images; for example, in the medical field, cross-modal retrieval can enable users to find semantically relevant medical images using text descriptions.
[0003] To alleviate the requirements of a large amount of storage, computing, and communication, large-scale multi-modal data is usually outsourced to a cloud platform. However, the cloud platform stores a large amount of data from different users, and there is a risk that the data may be maliciously tampered with, leaked, or damaged; and in the case of outsourcing multi-modal data, the data needs to be transmitted between the user and the cloud platform, and the instability of the network environment may cause problems such as packet loss, errors, or delays during data transmission; therefore, it has caused great concerns about the authenticity of the results during the retrieval process. The authenticity of the results includes integrity and correctness. Integrity is usually guaranteed through an authenticated query scheme, and correctness requires that all returned objects come from the data owner and have not been changed.
[0004] Existing authenticated query schemes can only handle single-modal queries and cannot be directly applied to cross-modal scenarios; while traditional cross-modal retrieval cannot be verified, and it is impossible to determine whether the returned objects have been changed; therefore, the prior art cannot guarantee both the authenticity of the query results and the query efficiency in cross-modal retrieval. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem in the prior art that the authenticity and retrieval efficiency of the retrieval results cannot be guaranteed simultaneously.
[0006] To solve the above technical problem, the present invention provides a secure and verifiable cross-modal retrieval method, including:
[0007] The data owner extracts features from the text modality and image modality of the multi-modal data in the multi-modal dataset, obtains high-dimensional vectors corresponding to the modalities, and constructs a verifiable B+ tree for the corresponding modality, including:
[0008] Cluster the high-dimensional vectors, use the clustering clusters as leaf nodes, and construct the internal nodes and root nodes of the B+ tree based on the range of the index key values of all the high-dimensional vectors in all the leaf nodes to form a basic B+ tree;
[0009] Calculate the center point and radius of each node based on the membership relationship of each node;
[0010] Use a hash function to generate a hash chain in a bottom-up manner and calculate the digest of each node;
[0011] Store the center point, radius, digest of the node, and the index key values of all the high-dimensional vectors into the basic B+ tree as a verifiable B+ tree;
[0012] Digitally sign the root node of the verifiable B+ tree for each modality to obtain the corresponding root node signature value;
[0013] Encrypt each multi-modal data in the multi-modal dataset to obtain encrypted multi-modal data; digitally sign the high-dimensional vectors corresponding to the modalities of each encrypted multi-modal data to obtain the digital signature values corresponding to the modalities;
[0014] Upload the verifiable B+ tree, root node signature value, all the encrypted multi-modal data, and digital signature values to the data service provider for storage;
[0015] The data service provider obtains the query vectors generated by the query user based on different modalities; based on the modality of the query vector, traverse the verifiable B+ tree corresponding to the other modality:
[0016] If the current traversed point is an internal node, traverse its child nodes;
[0017] If the current traversed point is a leaf node, calculate the distance between the center point of the leaf node and the query vector, obtain k vectors with the smallest distances, and construct a query result set; based on the query result set, the center point, radius, digest of the node, and the root node signature value, generate a verification object set for the query result set so that the query user can obtain a query result set that passes the verification and obtain the true retrieval result of the query vector.
[0018] Preferably, calculating the center point and radius of each node based on the membership relationship of each node includes:
[0019] Calculate the center point of the leaf node based on all the high-dimensional vectors in the clustering cluster represented by the leaf node Expressed as: m represents the total number of high-dimensional vectors in the cluster represented by leaf node L, V i represents the i-th high-dimensional vector in the cluster represented by leaf node L, where 1 ≤ i ≤ m;
[0020] Based on the distances between the center point of the leaf node and all the high-dimensional vectors in the cluster represented by the leaf node, obtain the radius r of the leaf node L , expressed as: S represents the set of all high-dimensional vectors in the cluster represented by the leaf node, represents the center point of the leaf node the distance between the center point of the leaf node and the i-th high-dimensional vector in the cluster represented by the leaf node;
[0021] Calculate the geometric center of the center points of all the child nodes of the internal node as the center point of the internal node expressed as: k represents the total number of child nodes of internal node N, represents the center point of the j-th child node of the internal node, where 1 ≤ j ≤ k;
[0022] Calculate the radius r of the minimum hypersphere that can cover all the child nodes of the internal node as the radius of the internal node N , expressed as: represents the center point of the internal node the distance between the center point of the internal node and the j-th child node N j therebetween, represents the radius of the j-th child node of the internal node;
[0023] Calculate the geometric center of the center points of all the child nodes of the root node as the center point of the root node
[0024] Calculate the radius r of the minimum hypersphere that can cover all the child nodes of the root node as the radius of the root node root .
[0025] Preferably, use a hash function to generate a hash chain in a bottom-up manner and calculate the digest of each node, including:
[0026] Calculate the digest d of leaf node L L , expressed as:
[0027]
[0028] Calculate the digest d of internal node N N , expressed as:
[0029]
[0030] Calculate the digest d of the root node root , which is expressed as:
[0031]
[0032] where h() represents a hash function, and (id L |K L |V i ) represents the <identifier, index key value / high-dimensional vector> of the leaf node; K1 to K j respectively represent the index key values of the 1st to jth child nodes of the internal node N, to respectively represent the digests of the 1st to jth child nodes of the internal node N; K′1 to K′ n respectively represent the index key values of the 1st to nth child nodes of the root node, to respectively represent the digests of the 1st to nth child nodes of the root node, and n represents the total number of child nodes of the root node.
[0033] Preferably, calculating the index key value of the high-dimensional vector includes:
[0034] Based on the distance dist(O i to the high-dimensional vector p in the ith clustering cluster i , p), calculate the index key value K p of the high-dimensional vector, which is expressed as:
[0035] K p = i * ρ + dist(O i , p);
[0036] where ρ is a preset partitioning parameter.
[0037] Preferably, obtaining the high-dimensional vectors corresponding to the text modality and the image modality in each multimodal data includes:
[0038] Input the image data in the multimodal data into the image modality deep neural network to obtain the high-dimensional image vector corresponding to the multimodal data;
[0039] Use the bag-of-words model to convert the text data in the multimodal data into a bag-of-words vector, and input the bag-of-words vector into the text modality deep neural network to obtain the high-dimensional text vector corresponding to the multimodal data.
[0040] Preferably, the image modality deep neural network is a convolutional neural network, including, connected in series in sequence: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a first fully connected layer, a second fully connected layer, and a third fully connected layer.
[0041] Preferably, the text-modal deep neural network includes, connected in series in sequence: a first fully connected layer with the ReLU function as the activation function, and a second fully connected layer with the identity function as the activation function.
[0042] Preferably, for the modality based on the query vector, traversing the verifiable B+ tree corresponding to another modality includes:
[0043] Initializing an auxiliary hash table H with a size of h' and a queued sequence Q with a size of 2k;
[0044] Using the KNN search algorithm to traverse the verifiable B+ tree, and determining whether the search sphere centered on the query vector with a preset query radius intersects the data sphere of the currently traversed node:
[0045] If It indicates that the search sphere intersects the data sphere of the currently traversed node, and determine whether the currently traversed node is a leaf node:
[0046] If it is not a leaf node, continue to traverse its child nodes until reaching the leaf node;
[0047] If it is a leaf node, check all <identifier, index key value, high-dimensional vector> of the high-dimensional vectors contained in this leaf node, and determine whether the identifier of the high-dimensional vector is in the auxiliary hash table:
[0048] If the identifier of this high-dimensional vector is in the auxiliary hash table, obtain the preset distance value in the auxiliary hash table as the retrieval distance of this high-dimensional vector;
[0049] If the identifier of this high-dimensional vector is not in the auxiliary hash table, calculate the distance between the query vector and this high-dimensional vector as the retrieval distance of this high-dimensional vector, and add <identifier, high-dimensional vector, retrieval distance> to the auxiliary hash table based on the identifier and the hash value;
[0050] Select the high-dimensional vectors with retrieval distances less than the preset query radius, and add <identifier, high-dimensional vector, retrieval distance> to the queued sequence;
[0051] If It indicates that the search sphere does not intersect the data sphere of the currently traversed node, and continue to search for the next node;
[0052] Until the total number of high-dimensional vectors in the queued sequence is not less than the preset number, then obtain the identifiers of all high-dimensional vectors in the queued sequence;
[0053] Based on the identifiers of all high-dimensional vectors in the queued sequence, obtain the corresponding multi-modal data, store it in the form of <identifier, multi-modal data, high-dimensional vector>, and construct the query result set R.
[0054] Preferably, based on the query result set, the center point of the node, the radius, the summary, and the root node signature value, a verification object set is generated for the query result set, including:
[0055] Initialize the verification object set to be empty;
[0056] Use the KNN search algorithm to traverse the verifiable B+ tree, and determine whether the search sphere centered on the query vector and the data sphere of the currently traversed node intersect:
[0057] If Then it indicates that the search sphere intersects the data sphere of the currently traversed node, and determine whether the currently traversed node is a leaf node
[0058] If it is a leaf node, then determine whether the identifier of the leaf node exists in the query result set:
[0059] If it exists, add the center point, radius of the node, and <identifier, index key value, high-dimensional vector> of all high-dimensional vectors in the node to the verification object set;
[0060] If it does not exist, add the summary of the node, and <identifier, high-dimensional vector> of all high-dimensional vectors in the node to the verification object set;
[0061] If it is not a leaf node, add the center point, radius of the node, and <index key value> of all high-dimensional vectors in the node to the verification object set;
[0062] If Then it indicates that the search sphere does not intersect the data sphere of the currently traversed node, then add the center point, radius and summary of the node to the verification object set;
[0063] Add the root node signature value and the digital signature values corresponding to all high-dimensional vectors in all nodes in the query result set to the verification object set, and obtain the verification object set VO of the query object set.
[0064] Preferably, the query user obtains the query result set that passes the verification to obtain the true retrieval result of the query vector, including:
[0065] Determine whether the nodes in the result query set are all valid objects:
[0066] If for all id α ∈R and id β ∈VO - R, there exists a situation where dist(q, V α ) < dist(q, V β) nodes, not all nodes in the result query set are valid objects, so the current retrieval fails;
[0067] If for all id α ∈R and id β ∈VO-R, all satisfy dist(q, V α ) < dist(q, V β ), then all nodes in the result query set are valid objects, and based on the information contained in the verification object set, the minimum subtree is reconstructed, and the summaries of all nodes in the minimum subtree are calculated to reconstruct the summary of the root node to obtain the reconstructed root summary;
[0068] Use the public key to verify whether the reconstructed root summary matches the root node signature value:
[0069] If they do not match, the verification of the verification object set fails, and the current retrieval fails;
[0070] If they match, use the public key to verify whether the high-dimensional vectors in each node of the minimum subtree are consistent with the digital
[0071] signature value:
[0072] If they are not consistent, the current retrieval fails;
[0073] If they are consistent, the query result set is true and the retrieval is successful.
[0074] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0075] The secure and verifiable cross-modal retrieval method described in the present invention is based on a multi-modal data set. After extracting the high-dimensional vectors corresponding to the image modality and text modality of all multi-modal data, clustering is used to generate a unified index structure for each node, and the corresponding image-verifiable B+ tree and text-verifiable B+ tree are constructed, so as to perform retrieval on the verifiable B+ tree of the corresponding type based on the type of the query request of the query user to achieve cross-modal retrieval. The verifiable B+ tree constructed by the present invention sets a unified index structure for leaf nodes, internal nodes and root nodes, which can effectively organize multi-modal data and map high-dimensional data to a one-dimensional space. In this way, during querying, it is possible to quickly locate the area that may contain the target data, reduce the search space, and improve the query efficiency; at the same time, by calculating the center point, radius and summary of each node, it is convenient to quickly judge the distance between the node and the query vector during subsequent queries, obtain the subtree that meets the conditions, and avoid unnecessary subtree traversal, thereby improving the query efficiency. Moreover, this application performs a signature operation on the summary of the root node of the verifiable B+ tree, and the query user can judge whether the data has been tampered with through signature verification, so as to ensure the authenticity of the query result.
[0076] After the present invention receives a query vector sent by a querying user, it traverses a verifiable B+ tree using the KNN search algorithm, with the query vector as the center and a search sphere with a preset verification radius, to obtain vectors that meet the distance requirements and form a query result set, avoiding meaningless excessive searches, saving computing resources and time costs, and improving the retrieval efficiency. At the same time, the present invention strategically utilizes shared tree nodes and prunes unnecessary subtrees, adopts different processing methods for intersection and non-intersection situations, comprehensively records relevant information of relevant nodes, and generates a corresponding verification object set for the query result set, so as to quickly and effectively determine whether the query results are true.
[0077] When the present invention verifies the authenticity of the query result set, it checks whether all valid objects are in the query result set, reconstructs the minimum subtree based on the verification object set, and verifies whether the reconstructed root digest of the minimum subtree matches the root node signature, and uses the public key to verify whether the high-dimensional vectors in the nodes of the minimum subtree are consistent with the digital signature values, so as to determine whether the query results in the query result set are accurate and have not been tampered with, further ensuring the integrity and correctness of the query results in the entire cross-modal retrieval process. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the drawings, where:
[0079] Figure 1 is a flowchart of the steps of the secure verifiable cross-modal retrieval method provided by the present invention;
[0080] Figure 2 is a system structure diagram;
[0081] Figure 3 is a flowchart of verifiable cross-modal retrieval of VCMR implementation results taking two modalities of image and text as an example;
[0082] Figure 4 is a multi-modal neural network model diagram;
[0083] Figure 5 is a schematic diagram of the construction process of a verifiable B+ tree;
[0084] Figure 6 is a schematic diagram of the generation of a verification object set. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] The following further illustrates the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.
[0086] Refer to Figure 1As shown in the figure, the flowchart of the steps of the secure and verifiable cross-modal retrieval method provided by the present invention specifically includes the following steps:
[0087] S101: The data owner extracts features from the text modality and image modality of the multi-modal data in the multi-modal dataset, obtains high-dimensional vectors corresponding to the modalities, and constructs a verifiable B+-tree for each modality, including:
[0088] Cluster the high-dimensional vectors, use the clustering clusters as leaf nodes, and construct the internal nodes and root nodes of the B+-tree based on the range of the index key values of all high-dimensional vectors in all leaf nodes to form a basic B+-tree;
[0089] Calculate the center point and radius of each node based on the membership relationship of each node;
[0090] Use a hash function to generate a hash chain in a bottom-up manner and calculate the digest of each node;
[0091] Store the center point, radius, digest of the node, and the index key values of all high-dimensional vectors
[0092] into the basic B+-tree as a verifiable B+-tree;
[0093] S102: Digitally sign the root node of the verifiable B+-tree for each modality to obtain the corresponding root node signature value
[0094] S103: Encrypt each multi-modal data in the multi-modal dataset to obtain encrypted multi-modal data; digitally sign the high-dimensional vectors corresponding to each modality of each encrypted multi-modal data to obtain the digital signature values corresponding to the modalities;
[0095] S104: Upload the verifiable B+-tree, root node signature value, all encrypted multi-modal data, and digital signature values to the data service provider for storage;
[0096] S105: The data service provider obtains the query vectors generated by the query user based on different modalities; based on the modality of the query vector, traverse the verifiable B+-tree corresponding to the other modality:
[0097] If the current traversed point is an internal node, traverse its child nodes;
[0098] If the current traversed point is a leaf node, calculate the distance between the center point of the leaf node and the query vector, obtain k vectors with the smallest distances, and construct a query result set; based on the query result set, the center point, radius, digest of the node, and the root node signature value, generate a verification object set for the query result set so that the query user can obtain a query result set that passes the verification and obtain the true retrieval result of the query vector.
[0099] In step S101, high-dimensional vectors corresponding to the text modality and the image modality in each multimodal data are obtained, including:
[0100] S101-1: Input the image data in the multimodal data into the deep neural network of the image modality to obtain the high-dimensional image vector corresponding to the multimodal data;
[0101] S101-2: Use the bag-of-words model to convert the text data in the multimodal data into a bag-of-words vector, and input the bag-of-words vector into the deep neural network of the text modality to obtain the high-dimensional text vector corresponding to the multimodal data.
[0102] Among them, the deep neural network of the image modality is a convolutional neural network, including, connected in series in turn: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a first fully connected layer, a second fully connected layer, and a third fully connected layer. The deep neural network of the text modality includes, connected in series in turn: a first fully connected layer with the ReLU function as the activation function, and a second fully connected layer with the identity function as the activation function.
[0103] Specifically, the nodes in the B+ tree include leaf nodes, internal nodes, and root nodes. Based on the membership relationship of each node, the center point and radius of each node are calculated, including:
[0104] Based on all the high-dimensional vectors in the clustering cluster represented by the leaf node, calculate the center point of the leaf node Expressed as: m represents the total number of high-dimensional vectors in the clustering cluster represented by the leaf node L, V i represents the i-th high-dimensional vector in the clustering cluster represented by the leaf node L, 1 ≤ i ≤ m;
[0105] Based on the distance between the center point of the leaf node and all the high-dimensional vectors in the clustering cluster represented by the leaf node, obtain the radius r of the leaf node L , expressed as: S represents the set of all high-dimensional vectors in the clustering cluster represented by the leaf node, represents the center point of the leaf node and the distance between the i-th high-dimensional vector in the clustering cluster represented by the leaf node;
[0106] Calculate the geometric center of the center points of all the child nodes of the internal node as the center point of the internal node Expressed as: k represents the total number of child nodes of the internal node N, represents the center point of the j-th child node of the internal node, 1 ≤ j ≤ k;
[0107] Calculate the radius r of the minimum hypersphere that can cover all the child nodes of an internal node as the radius of this internal node N , denoted as: represents the center point of the internal node The distance between the center point of the internal node and the j-th child node N j of the internal node, represents the radius of the j-th child node of the internal node;
[0108] Calculate the geometric center of the center points of all the child nodes of the root node as the center point of the root node
[0109] Calculate the radius r of the minimum hypersphere that can cover all the child nodes of the root node as the radius of the root node root .
[0110] Specifically, use the hash function to generate a hash chain in a bottom-up manner, and calculate the digest of each node, including:
[0111] Calculate the digest d of the leaf node L L , denoted as:
[0112]
[0113] Calculate the digest d of the internal node N N , denoted as:
[0114]
[0115] Calculate the digest d of the root node root , denoted as:
[0116]
[0117] where h() represents the hash function, (id L |K L |V i ) represents the <identifier, index key value / high-dimensional vector> of the leaf node; K1 to K j respectively represent the index key values of the 1st to j-th child nodes of the internal node N, to respectively represent the digests of the 1st to j-th child nodes of the internal node N; K′1 to K′ n respectively represent the index key values of the 1st to n-th child nodes of the root node, to respectively represent the digests of the 1st to n-th child nodes of the root node, and n represents the total number of child nodes of the root node.
[0118] Specifically, calculating the index key value of a high-dimensional vector includes:
[0119] Based on the clustering center point O in the i-th clustering cluster i to the distance dist(O i , p) of the high-dimensional vector p, calculate the index key value K p , expressed as:
[0120] K p = i * ρ + dist(O i , p);
[0121] where ρ is a preset partition parameter.
[0122] When storing data, the secure and verifiable cross-modal retrieval method of the present invention, based on a multi-modal data set, extracts the high-dimensional vectors corresponding to the image modality and text modality of all multi-modal data, and then uses clustering to generate a unified index structure for each node, constructs the corresponding image-verifiable B+ tree and text-verifiable B+ tree, so as to perform retrieval on the corresponding type of verifiable B+ tree based on the type of the query request of the query user, and realize cross-modal retrieval. The verifiable B+ tree constructed by the present invention sets a unified index structure for leaf nodes, internal nodes and root nodes, can effectively organize multi-modal data, map high-dimensional data to a one-dimensional space, so that during query, it is possible to quickly locate the area that may contain the target data, reduce the search space, and improve the query efficiency; at the same time, by calculating the center point, radius and digest of each node, it is convenient to quickly judge the distance between the node and the query vector during subsequent queries, obtain the subtree that meets the conditions, avoid unnecessary subtree traversal, and improve the query efficiency. And this application signs the digest of the root node of the verifiable B+ tree, and the query user can judge whether the data has been tampered with through signature verification, so as to ensure the authenticity of the query result.
[0123] In the embodiment of the present invention, the query vector generated based on the text modality is to be searched on the image-verifiable B+ tree, and the query vector generated based on the image modality is to be searched on the text-verifiable B+ tree.
[0124] Specifically, in step S105, the step of obtaining the query result set includes:
[0125] S105-1: Initialize an auxiliary hash table H with a size of h' and a queued sequence Q with a size of 2k;
[0126] S105-2: Use the KNN search algorithm to traverse the verifiable B+ tree, and judge whether the search sphere with the query vector as the center and a preset query radius intersects the data sphere of the currently traversed node:
[0127] S105-3: If It indicates that the search sphere intersects with the data sphere of the currently traversed node, and it is determined whether the currently traversed node is a leaf node:
[0128] If it is not a leaf node, continue to traverse its child nodes until a leaf node is reached;
[0129] If it is a leaf node, check <identifier, index key value, high-dimensional vector> of all high-dimensional vectors contained in this leaf node, and determine whether the identifier of the high-dimensional vector is in the auxiliary hash table:
[0130] If the identifier of this high-dimensional vector is in the auxiliary hash table, obtain the preset distance value in the auxiliary hash table as the retrieval distance of this high-dimensional vector;
[0131] If the identifier of this high-dimensional vector is not in the auxiliary hash table, calculate the distance between the query vector and this high-dimensional vector as the retrieval distance of this high-dimensional vector, and add <identifier, high-dimensional vector, retrieval distance> to the auxiliary hash table based on the identifier and hash value;
[0132] Select high-dimensional vectors with retrieval distances less than the preset query radius, and add <identifier, high-dimensional vector, retrieval distance> to the queued sequence;
[0133] S105-4: If It indicates that the search sphere does not intersect with the data sphere of the currently traversed node, and continue to search for the next node;
[0134] S105-5: Until the total number of high-dimensional vectors in the queued sequence is not less than the preset number, obtain the identifiers of all high-dimensional vectors in the queued sequence;
[0135] S105-6: Based on the identifiers of all high-dimensional vectors in the queued sequence, obtain the corresponding multimodal data, store it in the form of <identifier, multimodal data, high-dimensional vector>, and construct the query result set R.
[0136] The data service provider generates a verification object set for the query result set based on the query result set, the center point, radius, summary of the node, and the root node signature value, including:
[0137] Initialize the verification object set to be empty;
[0138] Use the KNN search algorithm to traverse the verifiable B+ tree, and determine whether the search sphere with the query vector as the center and the preset verification radius intersects with the data sphere of the currently traversed node:
[0139] If It indicates that the search sphere intersects with the data sphere of the currently traversed node, and determine whether the currently traversed node is a leaf node
[0140] If it is a leaf node, determine whether the identifier of the leaf node exists in the query result set:
[0141] If it exists, add the center point, radius of the node, and <identifier, index key value, high-dimensional vector> of all high-dimensional vectors in the node to the verification object set;
[0142] If it does not exist, add the summary of the node and <
[0143] identifier, high-dimensional vector> of all high-dimensional vectors in the node to the verification object set;
[0144] If it is not a leaf node, add the center point, radius of the node, and <index key value> of all high-dimensional vectors in the node to the verification object set;
[0145] If it indicates that the search sphere and the data sphere of the currently traversed node do not intersect, then add the center point, radius and summary of the node to the verification object set;
[0146] Add the root node signature value and the digital signature values corresponding to all high-dimensional vectors in all nodes in the query result set to the verification object set to obtain the verification object set VO of the query object set.
[0147] The query user obtains the query result set that passes the verification to obtain the true retrieval result of the query vector, including:
[0148] Determine whether the nodes in the result query set are all valid objects:
[0149] If for all id α ∈R and id β ∈VO-r, there exists a node that does not satisfy dist(q, V α ) < dist(q, V β ), then the nodes in the result query set are not all valid objects, and the current retrieval fails;
[0150] If for all id α ∈R and id β ∈VO-R, all satisfy dist(q, V α ) < dist(q, V β ), then the nodes in the result query set are all valid objects, and based on the information contained in the verification object set, reconstruct the minimum subtree, and calculate the summaries of all nodes in the minimum subtree to reconstruct the summary of the root node to obtain the reconstructed root summary;
[0151] Use the public key to verify whether the reconstructed root summary matches the root node signature value:
[0152] If there is no match, the verification object set fails the verification, and the current retrieval fails;
[0153] If there is a match, use the public key to verify whether the high-dimensional vectors in each node of the minimum subtree are consistent with the digital signature value:
[0154] If they are inconsistent, the current retrieval fails;
[0155] If they are consistent, the query result set is authentic and the retrieval is successful.
[0156] After the query user issues a query vector, the present invention traverses the verifiable B+-tree using the KNN search algorithm, and uses a search sphere with a preset verification radius centered on the query vector to obtain vectors that meet the distance requirements, forming a query result set, avoiding meaningless over-search, saving computing resources and time costs, and improving retrieval efficiency; at the same time, the present invention strategically uses shared tree nodes and prunes unnecessary subtrees, adopts different processing methods for intersecting and non-intersecting situations, comprehensively records relevant information of relevant nodes, generates a corresponding verification object set for the query result set, so as to quickly and effectively determine whether the query result is authentic. When the present invention verifies the authenticity of the query result set, it checks whether all valid objects are in the query result set, reconstructs the minimum subtree based on the verification object set, and verifies whether the reconstructed root digest of the minimum subtree matches the root node signature, and uses the public key to verify whether the high-dimensional vectors in the nodes of the minimum subtree are consistent with the digital signature value, so as to determine whether the query results in the query result set are accurate and not tampered with, further ensuring the reliability of the query results in the entire cross-modal retrieval process and the authenticity of the data in the query results.
[0157] Based on the above embodiments, the verifiable cross-modal retrieval method VCMR provided by the embodiments of the present invention consists of 5 key algorithms, namely Π=(KeyGen, Construct, Token, Query, Verify); the specific framework includes:
[0158] ① Key generation KeyGen;
[0159] {<sk, pk>, <K pri , K pub >} ← KenGen(λ);
[0160] The data owner DO (Data Owner) takes the security parameter λ as input, generates the private key sk and public key pk for the Paillier cryptosystem, and the private key K pri and public key K pub ;
[0161] ② Index construction Construct;
[0162] I ← Construct(V, pk);
[0163] DO Obtain the image vectors and text vectors of the multimodal dataset D, encrypt and sign the dataset, and then apply k-means clustering to construct a VB + tree, and upload the encrypted dataset, signature, and tree structure to the data service provider DSP (Data Service Provider);
[0164] ③ Query token generation Token;
[0165] q ← Token(q, pk);
[0166] The query user QU (Query User) generates a query vector q and submits it to the data service provider DSP;
[0167] ④ Query processing Query;
[0168] {R, VO} ← Query(q, I);
[0169] The data service provider DSP collaborates with the data assistance provider DAP (Data Assistance Provider) to search for the k most relevant results based on the query vector and the verifiable index through secure similarity measurement, and returns the encrypted result R and the verification object VO;
[0170] ⑤ Result verification Verify;
[0171] {0, 1} ← Verify(R, VO);
[0172] The query user QU uses the information in the verification object VO to verify the integrity and correctness of the encrypted result R.
[0173] Among them, the index structure of VCMR: VB + The tree is used to ensure the authenticity of the results but does not provide privacy protection; its nodes contain pointers p j , index key K j , the center point of the cluster and radius as well as the digest and other information. The leaf nodes record multiple <id j , K j , V j > pairs, where id j is the identity of V j associated with the original data. The digest is generated through a specific hash calculation, and the node digests are calculated from bottom to top. Finally, the root node digest is signed for result verification.
[0174] A verifiable B+-tree (VB+-tree) can be formalized as follows:
[0175]
[0176] The present invention uses the k-means clustering method to partition the high-dimensional space, selects reference points through specific rules; maps all vectors to a one-dimensional space, and then calculates the center point and radius of each node; on this basis, determines the summary of each node by a bottom-up method, and finally performs a signature operation on the summary of the root node.
[0177] The cross-modal retrieval method of the present invention includes that the data owner performs data preprocessing, such as generating keys, obtaining vectors, encrypting and signing data, and constructing a tree structure to upload to the data service provider. After the query user submits a query vector, the data service provider and the data assistance provider cooperate to calculate the distance between the query vector and the objects in the verifiable B+-tree to obtain a result set. At the same time, the data service provider traverses the verifiable B+-tree again to generate a verification object containing information related to the objects in the result set and signatures. After receiving the result set and the verification object returned by the data service provider, the query user performs a series of operations to ensure the authenticity of the results, including checking whether all valid objects are in the result set, reconstructing and verifying the root summary of the verifiable B+-tree (calculating through the information on the path from the leaf node containing the object to the root node and comparing with the signature in the verification object), and confirming the correctness of the signatures of the objects in the result set, so as to judge whether the retrieval result is accurate and not tampered with, and ensure the authenticity of the results of the entire cross-modal retrieval process.
[0178] Based on the above embodiments, this embodiment uses the secure and verifiable cross-modal retrieval method provided above for data storage and cross-modal retrieval; referring to Figure 2 as shown, it is the system structure diagram; referring to Figure 3 as shown, it is the flowchart of the verifiable cross-modal retrieval of the VCMR implementation result taking two modalities of image and text as an example. The specific steps include:
[0179] S201: Use a deep neural network to generate a representation vector;
[0180] Taking two modalities of image and text as an example, use a deep learning network for feature learning;
[0181] In a system with two modalities, it includes a multi-modal data set with two modalities, expressed as:
[0182] n represents the data size;
[0183] DO starts from the multi-modal data set to obtain image vectors for similarity measurement and text vectors
[0184] Refer to Figure 4 As shown, it is a multi-modal neural network model diagram; the feature learning part contains two deep neural networks, one for the image modality and the other for the text modality.
[0185] The deep neural network for the image modality is the convolutional neural network CNN. This CNN model has 8 layers and outputs the learned image features. Its network structure is as follows:
[0186] 1) conv1: There are 64 convolutional kernels, the size of the convolutional kernel is 11×11, the convolutional stride is 4×4, the padding is 0, local response normalization (LRN) is applied and 2×2 pooling is performed;
[0187] 2) conv2: 265 convolutional kernels, size 5×5, stride 1×1, padding 2, LRN is applied and 2×2 pooling is performed;
[0188] 3) conv3: 265 convolutional kernels, size 3×3, stride 1×1, padding 1;
[0189] 4) conv4: 265 convolutional kernels, size 3×3, stride 1×1, padding 1;
[0190] 5) conv5: 265 convolutional kernels, size 3×3, stride 1×1, padding 1, 2×2 pooling is performed;
[0191] 6) Fully connected layer full6: There are 4096 nodes;
[0192] 7) Full7: There are 4096 nodes;
[0193] 8) Full8: The number of output nodes is the dimension of the high-dimensional vector, that is, the output of this layer is the high-dimensional vector of the image.
[0194] And in the first seven layers of this network structure, the rectified linear unit (ReLU) is used as the activation function, and the identity function is used as the activation function in the eighth layer.
[0195] Refer to Table 1 as shown, which are the configuration parameters of the deep neural network for the image modality;
[0196] Table 1 Configuration Parameters of the Deep Neural Network for the Image Modality
[0197]
[0198]
[0199] For feature learning from the text modality, first, each text y is represented as a vector with a bag-of-words (BOW) representation; then, the bag-of-words vector is used as the input to a deep neural network with two fully connected layers, denoted as "full1-full2". Referring to Table 2, parameters are configured for the text deep neural network, where the configuration shows the number of nodes in each layer, the activation function of the first layer is ReLU, and the activation function of the second layer is the identity function.
[0200] Table 2 Configuration Parameters of the Text Deep Neural Network
[0201] layer configuration Full1 Number of nodes and output dimension: 8192 full2 Number of nodes and output dimension: Dimension of text vector
[0202] S202: Construct a verifiable B+ tree;
[0203] After generating the high-dimensional vector and the multi-modal dataset D is encrypted, and each element is signed.
[0204] For an image, by combining its identity raw data and the vector a signature is generated using asymmetric encryption (such as RSA) Then, DO applies K-means clustering to V I and V T and constructs a verifiable B+ tree for each modality using the algorithm I←Construct(V,pk).
[0205] The index of the verifiable B+ tree implements the index of the one-dimensional key K i of the vector V i The steps include:
[0206] ① Clustering: Using a clustering method to divide the points in the high-dimensional space into different regions; the number of partitions (a key hyperparameter) directly affects the search performance;
[0207] ② Reference point selection: For each partition P i , select the clustering center point O i to set the key of the vectors within this partition;
[0208] ③ One-dimensional representation: All vectors are mapped to a one-dimensional space;
[0209] The key K p of the data point p is calculated as:
[0210] K p = i * ρ + dist(O i , p);
[0211] where \(i\) represents the \(i\)-th partition, \(\rho\) is chosen to be large enough to prevent key overlap between partitions, and \(dist(O i , p)\) is the distance from \(O i to \(p\). This key calculation method facilitates range queries.
[0212] By constructing a verifiable B+-tree and utilizing the Merkle tree property, the present invention effectively verifies the results, avoids data tampering and omission, and ensures the authenticity of retrieval results. By constructing a verifiable B+-tree index, it effectively organizes multi-modal data, maps high-dimensional data to a one-dimensional space, and facilitates the rapid positioning and retrieval of relevant data. During the query process, by using clustering information and distance calculation, it reduces the unnecessary search scope and improves the query speed.
[0213] Referring to Figure 5 shown in the figure, it is a schematic diagram of the construction process of a verifiable B+-tree. First, the space is divided into partitions \(P1\), \(P2\), and \(P3\) by applying 3-means clustering. Then, a basic B+-tree is constructed based on the \(K\) value. The specific construction process includes:
[0214] ① Using the selected clustering method, calculate the center point and radius \(r\) of each node.
[0215] For example, using the selected clustering method;
[0216] ② Determine the digest (\(d\)) of each node in a bottom-up manner. If it is a leaf node, use to calculate its digest. If it is an internal node, use is the first child node of \(N i . For example,
[0217] ③ Sign the digest of the root node for result verification; that is, \(s Root = sign(d Root , K Pri ). Finally, upload the encrypted multi-modal dataset \(D\), the data signature, and all verifiable B+-trees to the DSP.
[0218] S203: Query processing process;
[0219] Upon receiving a query, here the query is that the query user \(QU\) generates a query vector \(q\) through the algorithm Token. After the DSP calculates the Euclidean distance between the query and the objects in the verifiable B+-tree to measure the similarity. After receiving the query vector, the DSP runs \(\{R, VO\} \leftarrow Query(q, I)\) to output the result \(R\) and the verification object \(VO\).
[0220] The specific process of obtaining results and verification objects includes:
[0221] ① Input the query vector q and the search parameter k;
[0222] ② Use the KNN search algorithm in VCMR, with the query vector q, the root node Root of the verifiable B+ tree, and the search parameter k as inputs, and output the result queue R, with the initial value set to
[0223] Before the search, initialize an auxiliary hash table H and a sorted queue Q, with their sizes being |H| = n and |Q| = 2k respectively. H is used to prevent redundant dist(q, V i ) calculations, while Q is used to store the objects that meet the search conditions. The search radius r will be continuously expanded until the termination condition is met, and the matching objects will be added to Q.
[0224] During the traversal of the verifiable B+ tree, two situations may occur:
[0225] Intersection situation: If dist(q, φ node ) < r + r node , then the search sphere centered on intersects with the data sphere contained in the node, which indicates that there may be a match. For leaf nodes, we check all <id i , K i , V i > pairs. If id i is in H, we obtain dist i from H; otherwise, calculate dist i = dist(q, V i ), and use the hash value H(id i ) to insert <id i , V i , dist i > into H. If dist i < r, then add <id i , V i , dist i > to Q. For non-leaf nodes, we continue to traverse their child nodes.
[0226] Non-intersection situation: If dist(q, φ node ) ≥ r + r node , the two spheres do not intersect, and none of the in the node will meet the search conditions, so there is no need for further search. After fully expanding the search radius, once |Q| ≥ R, the KNN search ends. Use the collected identity information to retrieve semantically related multimodal data Therefore, the result set R is constructed as containing identifiers, multimodal data, and their corresponding vectors.
[0227] S204: Verification object generation;
[0228] After the retrieval results (R) obtained by traversing the verifiable B+ tree using the KNN search algorithm in the VCMR method, the DSP traverses the verifiable B+ tree again to generate VOs for all objects in R, strategically leveraging shared tree nodes and pruning unnecessary subtrees to optimize the process.
[0229] During this process, the VO generation algorithm in the VCMR method is used to generate VOs bas , and the VO generation process is explained here with the verifiable B+ tree as the input:
[0230] After the kNN search, a shared verification object VO is generated for the result set R. Using the query vector q, the result queue R, and the root node (Root) of the verifiable B+ tree as inputs, an initially empty VO is generated.
[0231] The initial search radius is set to r R , where V k is the k-th vector in R.
[0232] There are mainly two cases when traversing the verifiable B+ tree:
[0233] Intersection case:
[0234] For leaf nodes, if the id in R exists in the node i , then r node is added to the VO; if the id in R does not exist in the node i , then d node is added to the VO.
[0235] For internal nodes, r node is added to the VO and its child nodes are continued to be explored.
[0236] Non-intersection case r node , d node is added to the VO. After that, the DSP appends the root node signature s Root and the signatures of the elements in R to the VO.
[0237] Refer to Figure 6 As shown, it is a schematic diagram for generating a set of verification objects. Suppose there is a text vector q, k = 2, and the result set The DSP calculates r according to When traversing the verifiable B+ tree, it is found that nodes N1 and N2 intersect with the search radius, so their child nodes are further checked. Node N3 does not meet the intersection condition, and its center point radius and digest are added to VO. For leaf node L1, since and will be included in VO. For leaf nodes L2 and L4, because and so the corresponding tuples are also included in VO. For leaf node L3, since will be added to VO. Finally, s Root and are added to VO. The digest calculations for leaf nodes and internal nodes are as follows:
[0238]
[0239] S205: Result authenticity verification;
[0240] After receiving R and VO, QU runs {0,1} ← Verify(R,VO) using the auxiliary proof information provided in VO to verify the integrity and correctness of R. The specific process is as follows:
[0241] 1) Ensure that all valid objects are in R, that is, for all and there is established;
[0242] 2) Reconstruct and verify the root digest d + of the VB Root tree. Specifically, QU first reconstructs the minimum subtree according to the information contained in VO, which contains all the nodes related to the query. Then, QU reconstructs the root digest d Root by calculating the digests of these nodes, and uses the public key K pub to verify whether s Root matches the reconstructed root digest;
[0243] 3) Confirm the correctness of the signature . QU uses the public key K pub to verify whether each signature s i is consistent with the corresponding data object and vector. Taking the example in Figure 6 , QU first checks that for all and Whether it meets Then, reconstruct VB + The smallest subtree of the tree and verify the conditions of the pruned branches Using this subtree, QU can reconstruct d Root And use the public key K pub Verify s Root . Finally, verify the signature for correctness. Through these steps, the integrity and correctness of the retrieval results can be ensured.
[0244] The verifiable cross-modal retrieval method described in the present invention is based on a multi-modal data set. After extracting the high-dimensional vectors corresponding to the image modality and text modality of all multi-modal data, a unified index structure is generated for each node by clustering, and the corresponding image-verifiable B+-tree and text-verifiable B+-tree are constructed, so as to perform retrieval on the corresponding type of verifiable B+-tree based on the type of the query request of the query user, and realize cross-modal retrieval. The verifiable B+-tree constructed in the present invention sets a unified index structure for leaf nodes, internal nodes and root nodes, which can effectively organize multi-modal data and map high-dimensional data to a one-dimensional space. In this way, when querying, it is possible to quickly locate the area that may contain the target data, reduce the search space, and improve the query efficiency; at the same time, by calculating the center point, radius and digest of each node, it is convenient to quickly judge the distance between the node and the query vector during subsequent queries, obtain the subtree that meets the conditions, avoid unnecessary subtree traversal, and improve the query efficiency. Moreover, the present application performs a signature operation on the digest of the root node of the verifiable B+-tree, and the query user can judge whether the data has been tampered with through signature verification, so as to ensure the authenticity of the query results. After the query user issues a query vector, the present invention uses the KNN search algorithm to traverse the verifiable B+-tree, with the query vector as the center, a search sphere with a preset verification radius, to obtain vectors that meet the distance requirements and form a query result set, avoiding meaningless over-search, saving computing resources and time costs, and improving the retrieval efficiency; at the same time, the present invention strategically uses shared tree nodes and prunes unnecessary subtrees, adopts different processing methods for intersecting and non-intersecting situations, comprehensively records the relevant information of relevant nodes, generates a corresponding verification object set for the query result set, so as to quickly and effectively judge whether the query results are correct and complete. When verifying the authenticity of the query result set, the present invention checks whether all valid objects are in the query result set, reconstructs the smallest subtree based on the verification object set, and verifies whether the reconstructed root digest of the smallest subtree matches the root node signature, and uses the public key to verify whether the high-dimensional vector and the digital signature value in the nodes of the smallest subtree are consistent, so as to judge whether the query results in the query result set are accurate and have not been tampered with, further ensuring the integrity and correctness of the query results of the entire cross-modal retrieval process.
[0245] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0246] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0247] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0248] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0249] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A secure and verifiable cross-modal retrieval method, characterized in that, Including: The data owner extracts features from the text modality and image modality of the multimodal data in the multimodal dataset, obtains high-dimensional vectors corresponding to the modalities, and constructs a verifiable B+ tree for each modality, including: Clustering the high-dimensional vectors, using the clustering clusters as leaf nodes, and constructing the internal nodes and root node of the B+ tree based on the range of the index key values of all high-dimensional vectors in all leaf nodes to form a basic B+ tree; Calculating the center point and radius of each node based on the membership relationship of each node; Using a hash function to generate a hash chain in a bottom-up manner and calculating the digest of each node; Storing the center point, radius, digest of the node, and the index key values of all high-dimensional vectors into the basic B+ tree as a verifiable B+ tree; Digitally signing the root node of the verifiable B+ tree for each modality to obtain the corresponding root node signature value; Encrypting each multimodal data in the multimodal dataset to obtain encrypted multimodal data; digitally signing the high-dimensional vectors corresponding to the modalities of each encrypted multimodal data to obtain the digital signature values corresponding to the modalities; Uploading the verifiable B+ tree, root node signature value, all encrypted multimodal data, and digital signature values to the data service provider for storage; The data service provider obtains query vectors generated by the query user based on different modalities; traversing the verifiable B+ tree corresponding to the other modality based on the modality of the query vector: If the current traversed point is an internal node, traverse its child nodes; If the current traversed point is a leaf node, calculate the distance between the center point of the leaf node and the query vector, obtain k vectors with the smallest distances, and construct a query result set; generating a verification object set for the query result set based on the query result set, the center point, radius, digest of the node, and the root node signature value, so that the query user can obtain a query result set that passes the verification and get the true retrieval result of the query vector.
2. The secure and verifiable cross-modal retrieval method according to claim 1, wherein Calculating the center point and radius of each node based on the membership relationship of each node, including: Calculate the center point of the leaf node based on all the high-dimensional vectors in the cluster represented by the leaf node It is expressed as: m represents the total number of high-dimensional vectors in the cluster represented by the leaf node L, and V i represents the i-th high-dimensional vector in the cluster represented by the leaf node L, where 1 ≤ i ≤ m; Based on the distances between the center point of the leaf node and all the high-dimensional vectors in the cluster represented by the leaf node, obtain the radius r of the leaf node L , which is expressed as: S represents the set of all high-dimensional vectors in the cluster represented by the leaf node, represents the center point of the leaf node is the distance between the center point of the leaf node and the i-th high-dimensional vector in the cluster represented by the leaf node; Calculate the geometric center of the center points of all child nodes of the internal node as the center point of this internal node It is expressed as: k represents the total number of child nodes of the internal node N represents the center point of the j-th child node of the internal node, where 1 ≤ j ≤ k; Calculate the minimum hypersphere radius that can cover all the child nodes of the internal node, and use it as the radius r of this internal node N , which is expressed as: Represents the center point of the internal node And the j-th child node N of the internal node j The distance between them, Represents the radius of the j-th child node of the internal node; Calculate the geometric center of the center points of all child nodes of the root node as the center point of the root node Calculate the minimum hypersphere radius that can cover all the child nodes of the root node, and use it as the radius r of the root node root .
3. The secure and verifiable cross-modal retrieval method according to claim 2, wherein Using a hash function to generate a hash chain in a bottom-up manner and calculating the digest of each node, including: Calculate the digest d of leaf node L L , expressed as: Calculate the digest d of the internal node N N , denoted as: Calculate the digest d of the root node root , expressed as: Among them, h() represents a hash function, (id L |K L |V i ) represents the <identifier, index key value / high-dimensional vector> of a leaf node; K1 to K j respectively represent the index key values of the 1st to jth child nodes of the internal node N, to respectively represent the digests of the 1st to jth child nodes of the internal node N; K′1 to K′ n respectively represent the index key values of the 1st to nth child nodes of the root node, to respectively represent the digests of the 1st to nth child nodes of the root node, and n represents the total number of child nodes of the root node.
4. The secure and verifiable cross-modal retrieval method according to claim 1, characterized in that Calculating the index key value of the high-dimensional vector, including: Based on the clustering center point O in the i-th clustering cluster i to the distance dist(O i , p) of the high-dimensional vector p, calculate the index key value K p , which is expressed as: K p = i * ρ + dist(O i , p); where ρ is a preset partitioning parameter.
5. The secure and verifiable cross-modal retrieval method according to claim 1, wherein Obtaining the high-dimensional vectors corresponding to the text modality and image modality in each multimodal data, including: Inputting the image data in the multimodal data into the image modality deep neural network to obtain the high-dimensional image vector corresponding to the multimodal data; Using the bag-of-words model to convert the text data in the multimodal data into a bag-of-words vector, and inputting the bag-of-words vector into the text modality deep neural network to obtain the high-dimensional text vector corresponding to the multimodal data.
6. The secure and verifiable cross-modal retrieval method according to claim 5, wherein The image modality deep neural network is a convolutional neural network, including, in series: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a first fully connected layer, a second fully connected layer, and a third fully connected layer.
7. The secure and verifiable cross-modal retrieval method according to claim 5, characterized in that, The text modality deep neural network includes, in series: a first fully connected layer with the ReLU function as the activation function, and a second fully connected layer with the identity function as the activation function.
8. The secure and verifiable cross-modal retrieval method according to claim 1, wherein Traversing the verifiable B+ tree corresponding to the other modality based on the modality of the query vector, including: Initialize an auxiliary hash table H of size h’ and a queued sequence Q of size 2k; Use the KNN search algorithm to traverse the verifiable B+ tree, and determine whether the search sphere centered on the query vector with a preset query radius intersects the data sphere of the currently traversed node: If it indicates that the search sphere intersects with the data sphere of the currently traversed node, and determine whether the currently traversed node is a leaf node: If it is not a leaf node, continue to traverse its child nodes until reaching a leaf node; If it is a leaf node, check all <identifier, index key value, high-dimensional vector> of the high-dimensional vectors contained in this leaf node, and determine whether the identifier of the high-dimensional vector is in the auxiliary hash table: If the identifier of the high-dimensional vector is in the auxiliary hash table, obtain the preset distance value in the auxiliary hash table as the retrieval distance of the high-dimensional vector; If the identifier of the high-dimensional vector is not in the auxiliary hash table, calculate the distance between the query vector and the high-dimensional vector as the retrieval distance of the high-dimensional vector, and add <identifier, high-dimensional vector, retrieval distance> to the auxiliary hash table based on the identifier and hash value; Select high-dimensional vectors with retrieval distances less than the preset query radius, and add <identifier, high-dimensional vector, retrieval distance> to the queued sequence; If it indicates that the search sphere and the data sphere of the currently traversed node do not intersect, and continue to search for the next node; Until the total number of high-dimensional vectors in the queued sequence is not less than the preset number, obtain the identifiers of all high-dimensional vectors in the queued sequence; Based on the identifiers of all high-dimensional vectors in the queued sequence, obtain the corresponding multimodal data, store it in the form of <identifier, multimodal data, high-dimensional vector>, and construct a query result set R.
9. The secure and verifiable cross-modal retrieval method according to claim 8, wherein Based on the query result set, the center point, radius, summary of the node, and the root node signature value, generate a verification object set for the query result set, including: Initialize the verification object set to be empty; Use the KNN search algorithm to traverse the verifiable B+ tree, and determine whether the search sphere centered on the query vector with a preset verification radius intersects the data sphere of the currently traversed node: If it indicates that the search sphere intersects with the data sphere of the currently traversed node, and it is determined whether the currently traversed node is a leaf node If it is a leaf node, determine whether the identifier of this leaf node exists in the query result set: If it exists, add the center point, radius of this node, and all <identifier, index key value, high-dimensional vector> of the high-dimensional vectors in this node to the verification object set; If it does not exist, add the summary of this node, and all <identifier, high-dimensional vector> of the high-dimensional vectors in this node to the verification object set; If it is not a leaf node, add the center point, radius of this node, and all <index key value> of the high-dimensional vectors in this node to the verification object set; If it indicates that the search sphere and the data sphere of the currently traversed node do not intersect, then add the center point, radius, and summary of this node to the verification object set; Add the root node signature value and the digital signature values corresponding to all high-dimensional vectors in all nodes in the query result set to the verification object set to obtain the verification object set VO of the query object set.
10. The secure and verifiable cross-modal retrieval method according to claim 9, wherein The query user obtains the query result set that passes the verification to obtain the true retrieval result of the query vector, including: Determine whether the nodes in the result query set are all valid objects: If for all id α ∈ R and id β ∈ VO-R, there exists a node that does not satisfy dist(q, V α ) < dist(q, V β ), then not all nodes in the result query set are valid objects, and the current retrieval fails; If for all id α ∈R and id β ∈VO-R, dist(q, V α ) < dist(q, V β ) is satisfied, then all nodes in the result query set are valid objects. Based on the information contained in the verification object set, the minimum subtree is reconstructed, and the digests of all nodes in the minimum subtree are calculated to reconstruct the digest of the root node and obtain the reconstructed root digest; Use the public key to verify whether the reconstructed root summary matches the root node signature value: If they do not match, the verification of the verification object set fails, and the current retrieval fails; If they match, use the public key to verify whether the high-dimensional vectors in each node in the minimum subtree are consistent with the digital signature values: If they are inconsistent, the current retrieval fails; If they are consistent, the query result set is true and the retrieval is successful.
Citation Information
Patent Citations
Fuzzy query encryption method supporting dynamic verification in unreliable cloud computing environment
CN106776904A
Verifiable encrypted image retrieval method supporting dynamic updating
CN113569280A
Verifiable multi-mode spatio-temporal data index structure and spatio-temporal range query verification method
CN117194418A
Cross-modal semantic retrieval method for privacy protection based on tree structure
CN118535759A
Lightweight and privacy-protected cross-modal retrieval method in Internet of Things environment
CN119167393A