Multi-modal database ciphertext storage and retrieval method and system
By constructing encrypted Bloom filter set and hidden vector encryption technology, the problem of index expansion and low retrieval efficiency of multimodal data in the big data environment is solved, and efficient and secure multimodal data storage and retrieval is achieved, meeting the real-time requirements.
Patent Information
- Application Number
- CN202411343329.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology cannot effectively support the secure storage and retrieval of multimodal data in big data scenarios, especially when facing massive data and diverse queries, resulting in index expansion, low retrieval efficiency and inability to meet real-time requirements.
The multimodal database encrypted storage and search method is adopted to construct r empty Bloom filter sets of m lengths, and the Bloom filter is encrypted using hidden vector encryption technology. Combined with the CSC architecture and trap gate mechanism, the unified index and cross-type retrieval of multimodal data are realized, and efficient search is carried out through the ciphertext query scheduling service module.
It effectively reduces the index inflation problem, improves storage efficiency, realizes cross-type retrieval of multimodal data, meets the real-time requirements in the big data environment, and enhances data security and retrieval accuracy.
Smart Images

Figure CN120354420A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security, and particularly relates to a multi-modal database ciphertext storage and retrieval method and system thereof. Background Art
[0002] Data security is the prerequisite for the development of the national big data. The integration and open sharing of data resources are the keys to the utilization and value generation of big data. However, the accompanying security issues are becoming increasingly prominent, and ensuring data security has become a global consensus. Therefore, researching the key technologies for data security protection in an open environment is the prerequisite for ensuring the secure utilization of big data and an indispensable part of promoting the secure implementation of the big data strategy. In recent years, multiple big data centers and data exchanges have been established in various places. The big data centers and data exchanges are typical data outsourcing models. Data owners entrust their data to the cloud servers of the big data centers or data exchanges, and share the data with authorized legitimate users through the cloud servers. The data outsourced to the big data centers and data exchanges needs to be stored for a long time and continuously provide retrieval services for legitimate data users. Ensuring the security of data storage and retrieval is the primary requirement for realizing data security protection, and data encryption is the fundamental means to protect data security.
[0003] To ensure system performance, existing big data storage and retrieval systems have made compromises in terms of security and use the method of "disk encryption at rest" to protect data security. "Disk encryption at rest" means that data (and indexes) are only encrypted during storage: when data needs to be retrieved, the data (and indexes) are decrypted, and the retrieval is performed on the plaintext data or indexes; when inactive data needs to be stored on the disk, the data is encrypted before storage. This method can only resist attackers who only access disk data, but the more common attack scenario is that attackers obtain partial permissions of the cloud platform through various means and then use this permission to access active data. At this time, since the active data is not encrypted, the attacker can directly obtain the plaintext data. Therefore, the ideal way to ensure data security is to encrypt the data before it is outsourced and hosted to the cloud service center, and directly perform data retrieval on the ciphertext without decrypting the data throughout the process. At this time, the cloud platform cannot obtain any plaintext data, and attackers who invade the cloud platform also cannot obtain any plaintext data. However, current ciphertext storage and retrieval systems are mainly oriented to relational databases and are difficult to meet the security storage and retrieval requirements of massive unstructured data in the big data scenario.
[0004] Open big data is characterized by massive and multimodal data as well as diverse queries. On the one hand, existing ciphertext index construction technologies will bring huge index expansion, and the volume of ciphertext indexes is even much larger than the original data. On the other hand, they mainly target single-modal (such as text, space, image, etc.) data and cannot effectively support cross-type retrieval of multimodal data. To support secure queries of multimodal data, existing ciphertext retrieval methods use computationally complex cryptographic algorithms such as homomorphic encryption, which are difficult to meet the real-time requirements of big data retrieval services. Therefore, how to design efficient ciphertext index and retrieval technologies to support the real-time response of large-scale ciphertext retrieval services is the key to promoting the practical application of big data security protection technologies.
[0005] Bloom filters are widely used in constructing ciphertext retrieval methods because they can achieve efficient membership detection. However, traditional ciphertext retrieval methods based on Bloom filters usually need to construct a Bloom filter for each data object, which will inevitably bring a large amount of storage overhead. At the same time, retrieval methods based on Bloom filters cannot avoid the problem of false positives, which will directly affect the accuracy of retrieval. Summary of the Invention
[0006] The present invention provides a method and system for encrypted storage and retrieval of a multimodal database to solve the technical problems in the prior art that massive and multimodal data and diverse queries will cause data index expansion, cannot effectively support cross-type retrieval of multimodal data, and are difficult to meet the real-time requirements of big data retrieval services.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for encrypted storage and retrieval of a multimodal database includes the following steps:
[0009] Step 1: Construct a set of r empty Bloom filters with a length of m. Insert keywords into the set of empty Bloom filters, and calculate the insertion position loc of the binary group of each keyword to be inserted t so as to construct an initial set of Bloom filters BF;
[0010] Step 2: Assign an independent and random containment identifier C[i] to each position in the initial set of Bloom filters BF, and perform r repeated perturbations to generate a series of transformed Bloom filters BF new ;
[0011] Step 3: Encrypt the transformed Bloom filters through hidden vector encryption technology to generate a compressed encrypted index BF Enc ;
[0012] Step 4: For the keyword w′ to be retrieved, calculate the corresponding containment identifier C′ at its position i i[i], calculate the encrypted value of the position where the keyword to be retrieved needs to match according to the obtained identifier-containing value.
[0013] Step Five: Based on the encrypted value Construct a trapdoor TK based on the calculated encrypted value, and match the trapdoor and the encrypted index BF Enc If the match is successful, the server obtains a candidate result accordingly. Similarly, r candidate results are obtained from the set BF, and a set intersection operation is performed on the r candidate results to obtain the retrieval result.
[0014] Insert all keywords into the empty Bloom filter set using the CSC architecture.
[0015] Set the calculated insertion position to 1 and the remaining positions to 0.
[0016] The insertion position of the keyword is calculated through a hash function and a partitioning function.
[0017] The binary tuple for inserting the keyword is: (w, id), and the calculation method of the insertion position is: loc t =(h t (w) / m + g t (i)) / m; where w is the keyword contained in the file with identifier id, h t is the t-th hash function, g t is the t-th partitioning function, and m represents the total length of the Bloom filter.
[0018] The method of scrambling transformation is: where C′ represents the non-identifier-containing value at the i-th position in BF, and γ represents a randomly selected parameter.
[0019] Repeat the scrambling r times and use different parameters γ for each scrambling, that is, different BF new corresponding C′ uses different γ.
[0020] The encrypted value of the position where the keyword to be retrieved needs to match is calculated using the hidden vector encryption technology.
[0021] Hidden vector encryption method: where, d j0 and d j1 are the encrypted values of d j1 = Sym.Enc(α ij , 0 λ+logλ ), F0: {0,1} k ×{0,1} * →{0,1}* denotes a random function, Sym.Enc denotes a secure symmetric encryption algorithm, and α ij denotes a random number.
[0022] A multi-modal database encryption system, comprising:
[0023] A ciphertext data management service module, configured to receive raw data in the initial Bloom filter set BF, generate and manage keys for indexing and data encryption in the Bloom filter set, and perform encryption and decryption of the raw data;
[0024] A ciphertext index management service module, configured to extract data features from the encrypted data in the ciphertext data management service module and establish a ciphertext index, and manage the storage, update, and reading of the ciphertext index to assist in the smooth progress of the ciphertext retrieval process;
[0025] A ciphertext query scheduling service module, configured to perform a secure query operation according to the user's ciphertext query token and the ciphertext index stored in the ciphertext index management service module to obtain a corresponding ciphertext result set;
[0026] A service management center module, responsible for managing each service, including functions such as service monitoring, service governance, service discovery / registration, and service load balancing.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] A multi-modal database ciphertext storage and retrieval method disclosed by the present invention stores keyword information by constructing a Bloom filter set. As a probabilistic data structure with high space efficiency, the Bloom filter can represent the existence of a large number of elements in a limited space, effectively reducing the index expansion problem caused by a large amount of data in traditional indexes. Even when dealing with multi-modal data, it can maintain good storage efficiency. This method not only processes text keywords, but theoretically is also applicable to inserting feature summaries of multi-modal data such as images and videos as keywords into the Bloom filter, as long as these non-text data can be converted into computable identifiers. In this way, the limitation of a single data type is broken, and unified indexing and cross-type retrieval of multi-modal data are realized, enhancing the flexibility and generality of the system. Although the encryption process increases the computational complexity, through efficient hidden vector encryption technology and trapdoor mechanism, fast retrieval matching can be achieved while ensuring data security. Especially in step five, by parallel processing the set intersection operation of r candidate results, the retrieval speed can be significantly accelerated, meeting the real-time requirements of retrieval services in the big data environment. By assigning random inclusion identifiers to each Bloom filter position and performing multiple disruptions, and using hidden vector encryption technology to encrypt the Bloom filter set, this scheme greatly enhances the data security, effectively preventing unauthorized data access and leakage, which is particularly important for the storage and retrieval of sensitive data. Through repeated mapping operations, although the false positive problem of the Bloom filter cannot be completely eliminated, the probability of incorrect matching can be effectively reduced, thereby improving the accuracy of retrieval results. This method shows significant beneficial effects in solving key problems such as data index expansion, multi-modal data retrieval, retrieval real-time performance, and data security, and is particularly suitable for scenarios that require efficient and secure processing of a large amount of multi-modal data.
[0029] Furthermore, by using the CSC architecture to batch insert keywords, the efficiency of data preprocessing is improved. At the same time, through the cooperation of the ciphertext query scheduling service module and the ciphertext index management service module, efficient retrieval of ciphertext data is realized, ensuring fast and accurate multi-modal data matching even in the encrypted state, and meeting the real-time requirements.
[0030] Furthermore, by introducing a random parameter γ for multiple disruptions, the complexity of the encrypted index is increased, the difficulty of attackers to crack is enhanced, and at the same time, the dynamic adaptability of the system is maintained, and the encryption strength and security level can be adjusted according to actual needs.
[0031] Furthermore, the application of hidden vector encryption technology not only protects the privacy of the index, but also ensures the security of the query process, and the actual information of the query keyword will not be leaked even in the query stage. By constructing a trapdoor TK and performing matching, secure retrieval is realized, reducing the risk of information leakage.
[0032] Furthermore, in a multi-modal database encryption system, by using a set of Bloom filters, the system can store index information of a large number of keywords in a compact form, significantly reducing the storage space requirements for indexes and effectively alleviating the problem of data index bloat. Especially in a multi-modal data environment, this method helps to efficiently manage different types of index information and avoids the complication of index structures caused by diverse data types in traditional methods. The system supports extracting features from multi-modal data and establishing a unified ciphertext index, which means that whether it is text, images, or data of other modalities, they can all be retrieved through the same set of index mechanisms, effectively breaking down the barriers between modalities and improving the accuracy and efficiency of cross-type data retrieval. The ciphertext data management service module is responsible for the encryption processing of the original data, as well as the management of keys and the encryption and decryption operations of the data, ensuring the security of the data during storage and transmission. The introduction of the ciphertext index and query scheduling service further ensures data privacy during the query process, and even during the retrieval process, the user's query intent and data content are not leaked.
[0033] Furthermore, the design of the service management center module in the system realizes the centralized management and optimization of each service component, including service monitoring, governance, discovery / registration, and load balancing, ensuring the stability and scalability of the system in the face of large-scale data and high-concurrency queries, and providing a solid infrastructure support for the retrieval service in a big data environment.
[0034] Furthermore, through centralized service management, the system can dynamically monitor the status of each service module, and perform service governance and optimization in a timely manner, such as automatically discovering new services, registering services, load balancing, etc. This not only improves the reliability of the system but also facilitates seamless expansion according to business needs in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of data and index encryption logic before compressing the ciphertext multi-set filter;
[0036] Figure 2 Schematic diagram of constructing a prefix family in ciphertext space query;
[0037] Figure 3 Schematic diagram of querying prefix elements in ciphertext space query;
[0038] Figure 4 Schematic diagram of encoding and extracting key feature data of image data;
[0039] Figure 5 Schematic diagram of encoding query of key feature data related to images;
[0040] Figure 6 Schematic diagram of hash learning for cross-modal ciphertext retrieval;
[0041] Figure 7 It is a schematic diagram of cross - membrane state ciphertext retrieval query;
[0042] Figure 8 It is an architecture diagram of a multi - modal encryption database system. Specific implementation manners
[0043] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0044] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] Embodiment 1
[0046] Refer to Figure 1 , a schematic diagram of the data and index encryption logic before compressing the ciphertext multi - set filter, to construct r Bloom filters {BF0, BF1, …, BF r-1}. Then, construct the corresponding {C0, C1, …, C r-1} for each Bloom filter, and assign a random inclusion identifier to each position. After that, the data owner calculates the non - value of the inclusion identifier to obtain {C0′, C1′, …, C r ′ -1}, and exclusive - OR each Bloom filter BF i with the corresponding C′ i bit - by - bit to obtain where i ∈ [0, r - 1] represents r Bloom filters. Finally, the data owner encrypts each Bloom filter bit - by - bit to obtain the encrypted Bloom filter where i ∈ [0, r - 1] represents r Bloom filters.
[0047] During query, the searching user constructs different search traps {tk 0 , tk 1 , …, tk r-1} for the r Bloom filters and sends them to the server to initiate a search. After receiving the search traps, the server matches the i - th Bloom filter with the i - th search trap tk i to obtain the i - th matching result, where i ∈ [0, r - 1] represents r Bloom filters. Finally, the server performs a set intersection operation on the r matching results obtained from the r Bloom filters to finally obtain the search result.
[0048] The specific encryption and retrieval steps are as follows:
[0049] First, the data owner constructs a set BF = {BF0,..., BF r-1} of r empty Bloom filters of length m, and then inserts all keywords using the CSC architecture: for the binary tuple (w, id) to which a keyword needs to be inserted, where w is the keyword contained in the file with identifier id, calculate the insertion location loc t = (h t (w) % m + g t (i)) % m, set the mapped position to 1, and the remaining positions to 0, where h t is the t-th hash function and g t is the t-th partitioning function, and m represents the total length of the Bloom filter. Thus, the final set BF of Bloom filters after insertion is obtained.
[0050] Taking a single Bloom filter as an example, assign an independent and random identifier to each position in BF and calculate Repeat the scrambling r times through the above transformation method to obtain To eliminate the correlation between different , different γ are used for the corresponding C′ of different BF new . Secondly, encrypt BF new using the hidden vector encryption technology where d j1 = Sym.Enc(α ij , 0 λ+logλ ). Finally, the compressed encrypted index is obtained. When a keyword w′ needs to be retrieved, calculate the identifier contained in the corresponding matching position of w′ where 0 ≤ i ≤ b - 1. Then, according to the obtained identifier, calculate the encrypted value of each matching position using the hidden vector encryption technology 0 ≤ i ≤ r - 1, 0 ≤ j ≤ k - 1, 0 ≤ t ≤ b - 1, and obtain the trapdoor where represents the search trapdoor corresponding to the t-th position in the partition of the j-th hash function in the i-th Bloom filter, k represents the number of hash functions, and r represents the number of repetitions. Finally, match TK and BF Enc . Taking one as an example, first use h j (w) to locate the corresponding query position in , and call the hidden vector encryption for matching. If for all 0 ≤ j ≤ k - 1, c t and all match successfully, it means that c tThe query keyword in it is included in the t-th partition. Therefore, the server obtains the candidate results R of the Bloom filter of . i . Similarly, r candidate results R = {R0,..., R r-1} can be obtained. Finally, the set intersection operation is performed on these r candidate results to obtain the retrieval result R.
[0051] Embodiment 2
[0052] This method can also be used for multimodal ciphertext retrieval, and its specific implementation steps are as follows:
[0053] To achieve the secure retrieval of encrypted text data, based on the compressed ciphertext multi-set filter, an efficient and secure text retrieval method is designed to achieve the secure retrieval of large-scale encrypted text data. Each file f can be represented by a set of keyword sets W, that is, the inclusion relationship between a set of keywords and the file can be represented by the binary tuple (w, id), where id is the identifier of file f and w is the keyword included in file f. The data owner constructs an index using the ciphertext multi-set filter: First, the data owner maps all binary tuples (w, id) into the CSC-BF and constructs the encrypted index Then, when the data user needs to query the keyword w′, the data user calculates the trapdoor TK and sends TK to the server to initiate the query; finally, the server obtains the retrieval result R containing the query keyword w′ by matching BF Enc with TK.
[0054] Embodiment 3
[0055] As Figure 2 and Figure 3 shown, this method can also be applied to ciphertext space query, and its specific implementation steps are as follows:
[0056] Based on the compressed ciphertext multi-set filter, an efficient and secure spatial encrypted data retrieval mechanism is designed to achieve the secure query of spatial encrypted data. As Figure 2 shown, for each spatial data, first use Hilbert coding to cover the space, and each spatial data is converted into the corresponding Hilbert coding value. For example, as Figure 2As shown, the Hilbert coding values corresponding to the spatial objects O1: (40.55, 585.5) and O2: (55.52, 561.3) are 13 and 45 respectively, and they are located at O1 and O2 in the figure. Then, using the prefix coding technique, the Hilbert coding values of each spatial object are constructed into corresponding prefix families. For example: O1 = 13 = {******, 0******, 00****, 001***, 0011**, 00110*, 001101}, O2 = 45 = {******, 1*****, 10****, 101***, 1011**, 10110*, 101101}. Each element in the prefix coding family is regarded as a keyword of the spatial data, and all spatial data and their keywords are mapped to a Bloom filter according to the structure of the ciphertext multi-set filter, and an encrypted index BF is constructed. Enc .
[0057] When it is necessary to query which spatial data are included in a certain spatial range, the data user first converts the query range into one-dimensional data using the Hilbert coding and finds the prefix elements that can cover the entire query range. For example, Figure 3 as shown: the Hilbert coding range corresponding to the spatial range {[50.5, 57.3], [550, 600.1]} is R = [38, 47]. Then, the search user uses the prefix coding technique to construct the Hilbert coding range corresponding to the search range into corresponding prefix elements. For example, R = [38, 47] = {101***, 10011*}. Then, this prefix element is regarded as a query keyword, and a trapdoor TK is constructed and uploaded to the server to initiate a query. The prefix coding algorithm believes that if the prefix element of the query range is included in the prefix family of a data, it means that the data is included in the query range. For example: the prefix element 101*** of the search range R is included in the prefix family of the spatial object O2. Therefore, the spatial object O2 is included in the search range R. Finally, the same as the ciphertext retrieval method in Embodiment 2, the server uses hidden vector encryption to match BF Enc and TK to obtain the retrieval result R.
[0058] Embodiment 4
[0059] As Figure 4 and Figure 5 shown, this method can also be applied to ciphertext image query, and its specific implementation steps are as follows:
[0060] An accurate and efficient image data retrieval mechanism is designed based on the compressed ciphertext multi-level sum filter and the visual word bag model, which solves the problem of low efficiency of image data retrieval and realizes the secure query of encrypted image data. As Figure 3As shown in the figure, for all image data, first, the bag-of-visual-words model is used to extract the feature vectors of each image to obtain a set of feature vectors, and all the feature vectors are clustered to obtain K clustering centers. The K clustering center vectors are regarded as K keywords to obtain the keyword dictionary W = {w1,..., w K}. According to the keyword dictionary W, each feature vector can be represented by a keyword in W that is closest to the feature vector in terms of Euclidean distance, that is, each image can be represented by a set of keywords W α . Then, each keyword is associated with the image data, and a ciphertext multi-set filter is used to construct the index BF new , where When an image query needs to be initiated, the data user uses the same method to extract the set of feature vectors of the query image. At the same time, each feature vector is transformed into a keyword in W that is closest to it in terms of Euclidean distance to obtain the query keyword set. For each query keyword w′, the query user calculates (h j (w′) + t) % m and h k ((h i (w′) + t) % m) to obtain the trapdoor TK = { (h j (w′) + t) % m, h k ((h j (w′) + t) % m)}, where 0 ≤ i ≤ r - 1, 0 ≤ j ≤ k - 1, k represents the number of hash functions, and r represents the number of repetitions. Finally, the server matches TK and BF new . Taking one as an example, the server uses loc t = (h j (w′) + t) % m to locate the corresponding position in and checks whether is equal to . If they are equal, it means the match is successful. Finally, similar to the text query, the server obtains the retrieval result R
[0061] Embodiment 5 As Figure 6 and Figure 7 shown, this method can also be used for cross-modal ciphertext retrieval, and its specific implementation method is as follows To achieve the retrieval of encrypted multi-modal data in the cloud server, a ciphertext cross-modal retrieval method is proposed based on cross-modal hashing and inner product encryption, and the retrieval process is as Figure 6 and Figure 7 shown
[0064] First, use the existing training dataset to train a cross-modal hashing function \(f\) through a machine learning model, which can accurately reflect the semantic information of various types of data. v (v) and \(f\) t (t), where \(f\) v (v) is the hashing function for mapping image data, and \(f\) t (t) is the hashing function for mapping text data. Using the corresponding cross-modal hashing function, each object in the dataset is mapped to a corresponding hash code. Through cross-modal hashing, different types of data are mapped to the same space across the heterogeneous gap. When querying, only the Hamming distance between two hash codes needs to be calculated to obtain the similarity degree of the corresponding two data. On the basis of achieving accurate cross-modal retrieval, to protect data privacy, each hash code \(y\) in the database is encrypted to obtain the encrypted value \(ct=(d,c=(b,a))\), where u and S are private keys, \(e\) i , \(e\) * and \(e\) are error factors, \(a\) is a random vector, and \(q'\) and \(p\) are public parameters. Finally, the data owner outsources the encrypted set to the cloud server. When a data user needs to initiate a query, he first uses the previously trained CMH to generate the hash code \(x\) of the query data, and then calculates the trapdoor \(TK = u + t*Tx\), where \(t\) is freely selected according to the public parameters, and \(T\) is the private key. The data user sends the trapdoor \(TK\) to the CSP to initiate a query. In the query phase, the CSP matches \(ct\) with \(TK\) to obtain the Hamming distance between the two data in \(ct\) and \(TK\) where \([*]\) represents rounding to the nearest integer. Finally, the CSP returns the data with the smallest Hamming distance as the retrieval result to the data user.
[0065] Example 6
[0066] Based on the above multi-modal database ciphertext storage and retrieval method, a multi-modal encrypted database system is designed, and the system architecture is as Figure 8 shown.
[0067] The multi-modal encrypted database system adopts a storage and computing separation architecture, which is compatible with the technical architecture of existing big data services. The existing big data platform can complete security reinforcement and upgrade through microservice incremental deployment, ensuring the scalability, usability, and efficiency of the system.
[0068] The system mainly consists of a ciphertext data management service, a ciphertext index management service, a ciphertext query scheduling service, and a service management center:
[0069] (1) The ciphertext data management service generates and manages the keys for indexing and data encryption, and performs encryption and decryption of the original data;
[0070] (2) The ciphertext index management service extracts data features from the data, establishes a ciphertext index, and manages the storage, update, and retrieval of the ciphertext index to assist the smooth progress of the ciphertext retrieval process;
[0071] (3) The ciphertext query scheduling service performs a secure query operation based on the user's ciphertext query token and the stored ciphertext index to obtain the corresponding ciphertext result set;
[0072] (4) The service management center is responsible for managing each service, including functions such as service monitoring, service governance, service discovery / registration, and service load balancing.
[0073] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art. The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A multi-modal database encryption storage and retrieval method, characterized in that, including the following steps, Step 1: Construct a set of r empty Bloom filters with length m. Insert keywords into the set of empty Bloom filters, and calculate the insertion position loc of each keyword pair to be inserted, so as to construct the initial Bloom filter set BF; t , thus constructing the initial Bloom filter set BF; Step 2: Assign an independent and random inclusion identifier C[i] to each position in the initial Bloom filter set BF, and perform r repeated disruptions to generate a series of transformed Bloom filters BF new ; Step 3: Encrypt the transformed Bloom filter through the hidden vector encryption technology to generate a compressed encrypted index BF Enc ; Step 4: For the keyword w' to be retrieved, calculate the inclusion identifier C' at the corresponding position i i [i], and calculate the encrypted value at the position where the keyword to be retrieved is matched according to the obtained inclusion identifier Step 5: Based on the encrypted value Construct a trapdoor TK based on the calculated encrypted value, and match the trapdoor with the encrypted index BF Enc If the match is successful, the server obtains a candidate result accordingly. Similarly, the set BF obtains r candidate results. Perform a set intersection operation on the r candidate results to obtain the retrieval result.
2. The multimodal database encryption storage and retrieval method according to claim 1, wherein Insert all keywords into the empty Bloom filter set using the CSC architecture.
3. A multimodal database encryption storage and retrieval method according to claim 1, characterized in that, Set the calculated insertion position to 1 and the remaining positions to 0.
4. A multimodal database encryption storage and retrieval method according to claim 1, characterized in that The insertion position of the keyword is calculated through a hash function and a partitioning function.
5. A multimodal database encryption storage and retrieval method according to claim 4, characterized in that The binary tuple that needs to insert keywords is: (w, id), and the calculation method of the insertion position is: loc t =(h t (w) / m + g t (i)) / m; where w is the keyword included in the file with the identifier id, h t is the t-th hash function, g t is the t-th partitioning function, and m represents the total length of the Bloom filter.
6. A multimodal database encryption storage and retrieval method according to claim 5, characterized in that The method of disrupted transformation is as follows: where C′ represents the non-inclusion identifier at the i-th position in BF, and γ represents a randomly selected parameter.
7. A multimodal database encryption storage and retrieval method according to claim 6, characterized in that, Repeat the perturbation r times and use different parameters γ for each perturbation, i.e., different BFs new The corresponding C′ uses different γ values.
8. A multimodal database encryption storage and retrieval method according to claim 1, characterized in that, The encrypted value at the position matched by the search keyword is required. It is calculated using the hidden vector encryption technology.
9. A multimodal database encryption storage and retrieval method according to claim 8, characterized in that, The hidden vector encryption method: where d j0 and d j1 are the encrypted values of d j1 = Sym.Enc(α ij , 0 λ+logλ ), F0: {0, 1} k × {0, 1} * → {0, 1} * represents a random function, Sym.Enc represents a secure symmetric encryption algorithm, and α ij represents a random number.
10. A multimodal database encryption system, characterized in that, A multi-modal database encryption storage and retrieval method according to any one of claims 1 to 9, comprising: A ciphertext data management service module for receiving the original data in the initial Bloom filter set BF, generating and managing the keys for indexing and data encryption in the Bloom filter set, and performing encryption and decryption of the original data; A ciphertext index management service module for extracting data features from the encrypted data in the ciphertext data management service module and establishing a ciphertext index, and managing the storage, update, and reading of the ciphertext index to assist the smooth progress of the ciphertext retrieval process; A ciphertext query scheduling service module for performing a secure query operation based on the user's ciphertext query token and the ciphertext index stored in the ciphertext index management service module to obtain a corresponding ciphertext result set; A service management center module for responsible for managing each service, including functions such as service monitoring, service governance, service discovery / registration, and service load balancing.