Searchable encrypted data sharing method, system and terminal based on BK-LSH
The BK-LSH-based data sharing method addresses privacy and efficiency challenges by employing a novel clustering approach and deep learning for secure, efficient, and flexible data sharing with enhanced privacy and computational efficiency.
Patent Information
- Application Number
- CN202510638274.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-04-27
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art has privacy protection problems in cross-domain data sharing, especially the difficulty in measuring similarity of classified data and high computational complexity, and the credibility of third-party parties is difficult to guarantee when sharing data from multiple parties, and the performance efficiency of existing clustering algorithms is insufficient.
The searchable encrypted data sharing method based on BK-LSH is adopted, a new clustering method similar to k-means is used, and keyword vectors are extracted in combination with the deep language learning model. The initial clustering is quickly generated and iterative optimization is optimized through the alliance chain, and multi-keyword searchable encryption is supported.
It improves cluster search performance, enhances data security and practicality, and realizes multi-keyword searchable encryption through the BK-LSH algorithm, which improves computing efficiency and flexibility while ensuring data security.
Smart Images

Figure CN120321022A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of vector data search, searchable encryption, and consortium chain data sharing, and particularly relates to a searchable encryption data sharing method, system, and terminal based on BK-LSH. Background Art
[0002] In recent years, big data technology has gradually occupied an important position in our lives. For example, big data applications such as finance, technology, energy, transportation, logistics, informatization, network security, intelligence, visualization, wisdom, and intelligence have had a significant impact on our daily lives. However, during the continuous strengthening of data convergence in different fields, many new challenges have emerged: First, the privacy protection problem in cross-domain data sharing has also become a major hidden danger. If the data privacy protection mechanism cannot be implemented, once the data is leaked, it will cause great harm. Second, it is very difficult to measure the similarity of categorical data (such as text, images, audio, etc.) because these data types do not have a natural numerical order, and due to their large scale and high computational complexity, it is difficult to operate efficiently in practical applications.
[0003] Currently, most privacy protection for data is achieved by desensitizing sensitive information in the data to protect this information. To protect privacy during data sharing, various methods are adopted. For example, to solve the privacy leakage problem caused by link attacks, the k-anonymity method is introduced. By describing the data more generally and abstractly, it becomes impossible to distinguish specific values, thereby protecting data users from being identified by attackers. Some people propose that randomness should be introduced into the data publishing or analysis process to ensure that even if an attacker has all the information except the target individual, they cannot infer any specific information about the target individual from the data publishing results. Others propose that searches can be performed on encrypted data, allowing users to submit keywords for convenient, flexible, and efficient searches while ensuring that the cloud server responsible for storage knows nothing about the ciphertext data itself and keyword-related information. However, these privacy protection methods all have certain defects. For example, the homomorphic encryption algorithm supports fewer calculation methods and consumes more computing resources, with low practicality; searchable encryption can only perform exact searches, which does not meet the actual needs of daily life. At the same time, with the rapid development of technologies such as cloud computing today, the demand for multi-party data sharing has increased greatly. At this time, a trusted third party is needed to operate and share multi-party data. However, it is difficult to ensure the trustworthiness of the third party in reality, which has become a difficult point in the current data sharing field.
[0004] In addition, the prior art has proposed various methods to solve the problem of clustering large-scale classification data, but all of them have certain limitations. For example: the k-modes algorithm uses the mode instead of the mean in k-means and uses an overlap metric to calculate the distance between classification objects. This method is directly applicable to classification data and solves the limitations of the k-means algorithm when dealing with classification data. However, at the same time, the k-modes algorithm has instability in selecting the mode because when multiple classification values have the same highest frequency on a certain feature, the selection of the mode will affect the clustering result. Another example is the k-representatives algorithm, which uses representatives instead of the mode. Representatives are the distributions of each feature in each cluster rather than specific classification values. This method improves the stability and accuracy of clustering. There is also the k-means++ algorithm, which improves the stability and efficiency of clustering by carefully selecting the initial clustering centers; the MH-k-modes algorithm uses the local sensitive hashing (LSH) technique to accelerate the k-modes algorithm. By constructing a hash table, similar data objects are stored in the same bucket, thus reducing the number of distances that need to be calculated in each iteration. However, the performance and efficiency of the existing clustering algorithms still cannot meet the actual usage requirements.
[0005] In view of this, there is an urgent need to propose a data sharing search scheme with high clustering search performance and high security. Summary of the Invention
[0006] The purpose of the present invention is to propose a searchable encryption data sharing method, system and terminal based on BK-LSH. Using a new clustering method similar to k-means, it can quickly generate initial clusters and optimize the clusters through an iterative process, reducing the number of distance calculations required for reallocation, greatly improving the clustering search performance, and implementing the multi-keyword searchable encryption function through the BK-LSH algorithm. On the premise of ensuring data security, the practicality and flexibility of this method are also enhanced.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In the first aspect, a searchable encryption data sharing method based on BK-LSH is proposed, which is implemented based on a consortium blockchain and includes:
[0009] S1. Configure and encrypt the original data to generate a digest and a data ciphertext corresponding to the original data;
[0010] S2. Extract multiple keywords of the original data of the original document data based on a deep language learning model, construct a vector representation for each keyword, and construct a hash table of the vector representation;
[0011] S3. Determine the binary sequence of the keys of the hash buckets in the hash table, group the hash buckets into several initial clusters and perform iterative optimization until the clustering is stable or the preset number of iterations is reached, and use the clustering result as the search index for the original data;
[0012] S4. Upload the digest corresponding to the original data, the data ciphertext, and the search index to the cloud server;
[0013] S5. Configure multiple query keywords, generate a query trapdoor and upload it to the consortium blockchain, and locate the corresponding hash bucket in the hash table based on the query trapdoor. Match the query trapdoor with the keyword vectors processed by the hash function in the hash bucket to obtain a matching degree score. Sort the search index according to the matching degree score from high to low, and output the digests and data ciphertexts corresponding to the top several search indexes;
[0014] S6. Decrypt the data ciphertext to obtain the original data.
[0015] As a possible implementation, S1 includes the following sub-steps:
[0016] S10. Clean the original data, use a program to delete irrelevant characters, and obtain the cleaned original data;
[0017] S11. Generate a digest of the cleaned original data using a one-way encryption algorithm;
[0018] S12. The data owners and data searchers in the consortium blockchain request a symmetric key from the key distribution center, and use the symmetric key to encrypt the cleaned original data to obtain the data ciphertext.
[0019] As a possible implementation, S2 includes the following sub-steps:
[0020] S20. Extract multiple keywords from the original data based on a deep language learning model, and construct a vector representation for each keyword;
[0021] S21. Generate a number of random vectors that follow the Gaussian distribution N(0,1), and the quantity is the number of preset hash tables multiplied by the number of hash functions for each hash table.
[0022] S22. For each vector, multiply it by the above random vector, add an offset uniformly distributed on [0, w), and then divide by the preset bucket width. The vector is assigned to the corresponding hash bucket to obtain the hash table represented by the vector.
[0023] As a possible implementation, S3 is specifically: Select the top k buckets with a larger number of accommodated data among the 2 l buckets; The remaining buckets are merged into the top k buckets based on the Hamming distance between the binary sequences according to the principle of proximity, and the k new buckets obtained after the merger are used as the k initial clusters of the clustering algorithm.
[0024] As a possible implementation, S4 includes the following sub-steps:
[0025] S40. Determine a cluster representative in each initial cluster. A total of k cluster representatives are determined for k initial clusters. Calculate the distance between the i-th cluster representative among the k cluster representatives and the other k - 1 cluster representatives, where 1 ≤ i ≤ k. Select the k cluster representatives with relatively close distances to form the shortlist of the i-th cluster representative; ′ and form the shortlist of the i-th cluster representative;
[0026] S41. Calculate the distance between each non-cluster representative data in the m-th initial cluster and each cluster representative in the shortlist, where 1 ≤ m ≤ k. Transfer the non-cluster representative data to the initial cluster where the closer cluster representative is located to form new clusters;
[0027] S42. Repeat S40 - S41 until the clustering is stable or the preset number of iterations is reached, and then jump to S43;
[0028] S43. Output the clustering result at this time, which is the search index.
[0029] As a possible implementation, configure multiple query keywords to generate a query trapdoor, including: configuring multiple query keywords, constructing a vector representation for each query keyword based on a deep language learning model, and performing hash processing on the vector representation of each query keyword using l binary hash functions to generate a query trapdoor.
[0030] As a possible implementation, the encryption of the original data in S1 is implemented using the AES symmetric encryption algorithm.
[0031] As a possible implementation, the deep language learning model in S2 is implemented using the BERT model.
[0032] As the second aspect of the present invention, a searchable encryption data sharing system based on BK-LSH is proposed, which is characterized by including: an encryption and upload unit, a vector extraction unit, a hash table construction unit, a clustering unit, a query unit, and a data acquisition unit;
[0033] The encryption and upload unit is used to configure the original data, generate a digest of the original data, encrypt the original data to obtain a data ciphertext, and upload the digest, data ciphertext, and search index to the cloud server;
[0034] The vector extraction unit extracts multiple keywords of the original data based on a deep language learning model and constructs a vector representation for each keyword;
[0035] The hash table construction unit constructs a hash table of the vector representation;
[0036] The clustering unit determines the binary sequences of the keys of 2 l buckets in the hash table, groups the 2 l buckets into initial clusters based on the Hamming distance between the binary sequences; iteratively optimize the initial clusters until the clustering is stable or a preset number of iterations is reached, and use the clustering result at this time as the search index; upload the digest, data ciphertext, and search index to the cloud server;
[0037] The query unit configures multiple query keywords to generate a query trapdoor; locates the corresponding hash bucket in the hash table based on the query trapdoor, and performs a matching calculation between the query trapdoor and the search index in the hash bucket, sorts them in descending order of the matching degree, and outputs the digests and data ciphertexts corresponding to the top λ search indexes;
[0038] The data acquisition unit verifies the integrity of the data based on the digest, and then decrypts the data ciphertext to obtain the original data.
[0039] Advantageous Effects
[0040] A searchable encryption data sharing method, system, and terminal based on BK-LSH proposed by the present invention have the following advantageous effects compared with the prior art:
[0041] 1. The proposed searchable encryption data sharing method based on BK-LSH uses the BERT deep model to vectorize keywords, can distinguish the deep meanings of words in different contexts by combining context content, the vectorization result is more accurate, and the comparison search effect is better;
[0042] 2. The proposed searchable encryption data sharing method based on BK-LSH uses a new clustering method similar to k-means, can quickly generate initial clusters, and optimizes the clusters through an iterative process, reducing the number of distance calculations required for reallocation, and greatly improving the clustering search performance;
[0043] 3. The proposed searchable encryption data sharing method based on BK-LSH performs data sharing through the consortium chain framework, performs keyword matching calculations on the server, has high calculation efficiency and short time consumption, and due to the immutable nature of the consortium chain, the security of data sharing is guaranteed;
[0044] 4. The proposed searchable encryption data sharing method based on BK-LSH realizes the multi-keyword searchable encryption function through the BK-LSH algorithm, enhancing the practicability and flexibility of the present invention while ensuring data security. Description of the Drawings
[0045] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0046] Figure 1 It is a schematic flowchart of a searchable encryption data sharing method based on BK-LSH in an embodiment of the present invention;
[0047] Figure 2 It is a schematic diagram of extracting keywords according to context based on the BERT deep learning model in an embodiment of the present invention. Detailed implementation manners
[0048] For the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily limit to be different.
[0049] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0050] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single item (item) or multiple items (items). For example, at least one (item) of a, b or c may represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b and c may be single or multiple.
[0051] An embodiment of the present invention aims to provide a searchable encryption data sharing method, system and terminal based on BK-LSH. Using a new clustering method similar to k-means, it can quickly generate initial clusters and optimize the clusters through an iterative process, reducing the number of distance calculations required for reallocation, greatly improving the clustering search performance, and implementing the multi-keyword searchable encryption function through the BK-LSH algorithm. On the premise of ensuring data security, it also has high practicability and flexibility.
[0052] In a first aspect, an embodiment of the present invention provides a searchable encryption data sharing method based on BK-LSH, which is implemented based on a consortium blockchain. See Figure 1 , including the following steps:
[0053] S1. Configure and encrypt the original data to generate a digest corresponding to the digest of the original data and the data ciphertext;
[0054] As a possible implementation, S1 includes the following sub-steps:
[0055] S10. Clean the original data, use a program to delete irrelevant characters, and obtain the cleaned original data;
[0056] S11. Generate a digest of the cleaned original data using a one-way encryption algorithm;
[0057] As an example, the one-way encryption algorithm is the SHA-256 algorithm.
[0058] S12. The data owner and data searcher in the consortium blockchain request a symmetric key from the key distribution center, and use the symmetric key to encrypt the cleaned original data to obtain the data ciphertext.
[0059] As a possible implementation, encrypting the cleaned original data is implemented using the AES symmetric encryption algorithm.
[0060] As an example, the data owner and data searcher request a symmetric key SK from the key distribution center for encrypting and decrypting data. For example, the cleaned original data is P, the digest of the cleaned original data P generated using the SHA-256 algorithm is H(P), and the data ciphertext C is obtained by encrypting the data P using the AES symmetric encryption algorithm.
[0061] S2. Extract multiple keywords of the original data based on a deep language learning model, construct a vector representation for each keyword, and construct a hash table of the vector representation;
[0062] Exemplarily, a class k-means LSH method is used to construct a hash table of the vector representation;
[0063] As a possible implementation, S2 includes the following sub-steps:
[0064] S20. Extract multiple keywords from the original data based on a deep language learning model, and construct a vector representation for each keyword;
[0065] See Figure 2 , as an example, use the deep language learning model BERT to extract multiple keywords {w1, w2, …, w k} through the context information of the original data. When constructing the vector representation for each keyword, any vector representation method in the prior art can be used, such as word embedding. Denote the obtained multiple vector representations as X = {v1, v2, …, v k}.
[0066] S21. Generate a number of random vectors that follow a Gaussian distribution N(0, 1), and the number is the number of preset hash tables multiplied by the number of hash functions for each hash table;
[0067] As an example, determine the number of hash tables L and the number of hash functions l. Usually, L is taken as 10 - 20. Increasing L can improve the recall rate, but it will increase the memory and time overhead. l is the number of hash functions for each hash table, usually taken as 864. Increasing l can improve the precision, but it will reduce the collision probability. Generate L×l Gaussian random vectors {a1, …, a L×l}, and generate L×l uniform offsets {b1, …, b L×l}.
[0068] S22. For each vector, multiply it by the above random vectors, then add an offset uniformly distributed on [0, w), and then divide by the preset bucket width. The vector is assigned to the corresponding hash bucket to obtain the hash table of the vector representation. The number of hash functions for each hash table is l, and the number of hash buckets is 2 l ;
[0069] As an example, the obtained hash table of the vector representation is:
[0070] S3. Determine the binary sequences of the keys of the 2 l buckets in the hash table, and group the 2 l buckets into k initial clusters based on the Hamming distance between the binary sequences and perform iterative optimization until the clustering is stable or reaches the preset number of iterations, and use the clustering result at this time as the search index for the original data;
[0071] As a possible implementation manner, S3 is specifically: Among the 2 lSelect the top k buckets with a larger number of data from the buckets; the remaining buckets are merged into the top k buckets based on the Hamming distance between binary sequences according to the principle of proximity, and the k new buckets obtained after merging are used as the k initial clusters of the clustering algorithm.
[0072] As an example, by based on the hash table in 2 l the Hamming distance between the binary sequences of the keys of the buckets, group these buckets into k initial clusters: First, select the k largest buckets that contain the largest number of data objects. Then, the remaining buckets will be merged into these larger buckets according to the Hamming distance between the binary sequences of the hash values of the buckets. The final k groups obtained will be used as the k initial clusters G = {G1, G2, …, G k}.
[0073] As a possible implementation, S3 includes the following sub-steps:
[0074] S30. Determine a cluster representative in each initial cluster. A total of k cluster representatives are determined for the k initial clusters. Calculate the distance between the i-th cluster representative among the k cluster representatives and the other k - 1 cluster representatives, where 1 ≤ i ≤ k, and select the k ′ cluster representatives with closer distances to form the shortlist of the i-th cluster representative;
[0075] As an example, the k cluster representatives determined by the k initial clusters are represented as: R = {r1, r2, …, r k}, and the following method is used to calculate the distance between the i-th cluster representative among the k cluster representatives and the other k - 1 cluster representatives, where 1 ≤ i ≤ k:
[0076]
[0077] where D is the dimension of the data, r id and r jd are the values of the representative r i and r j in the d-th dimension respectively.
[0078] S31. Calculate the distance between each non-cluster representative data in the m-th initial cluster and each cluster representative in the shortlist, where 1 ≤ m ≤ k, and transfer the non-cluster representative data to the initial cluster where the closer cluster representative is located to form a new cluster;
[0079] S32. Repeat S30 - S31 until the clustering is stable or reaches the preset number of iterations, and then jump to S33;
[0080] S33. Output the clustering result at this time, which is the search index.
[0081] As an example, for each cluster G j Determine the k″ nearest neighbor clusters (k″ is an integer less than k), denoted as Shortlist(G j ). For each object in the dataset, calculate its distance from the representative of each cluster in the shortlist and update its cluster membership. Based on the updated membership, reassign the data objects to new clusters and update the representative of each cluster. Then repeat the steps until the clusters are stable or a preset number of iterations is reached, and obtain the optimized clustering result as the search index I.
[0082] S4. Upload the abstract, data ciphertext, and search index corresponding to the original data to the cloud server;
[0083] S5. Configure multiple query keywords, generate a query trapdoor, upload it to the consortium blockchain, and locate the corresponding hash bucket in the hash table based on the query trapdoor. Match the query trapdoor with the keyword vectors processed by the hash function in the hash bucket to obtain a matching degree score. Sort the search index from high to low according to the matching degree score, and output the abstracts and data ciphertexts corresponding to several search indexes.
[0084] As a possible implementation, configure multiple query keywords and generate a query trapdoor, including: configuring multiple query keywords, constructing a vector representation of each query keyword based on a deep language learning model, and using l binary hash functions to perform hash processing on the vector representation of each query keyword to generate a query trapdoor.
[0085] S6. Decrypt the data ciphertext to obtain the original data.
[0086] As an example, independently configure multiple query keywords, construct a vector representation {y1, y2, …, y k} of each query keyword according to the BERT deep model, and then perform hash processing on it through the same hash function H = [h1, …, h l to generate a query trapdoor t. Upload the query trapdoor t to the consortium blockchain (CB). The cloud server obtains the query trapdoor t from the consortium blockchain, quickly locates the corresponding hash bucket H i in the hash table according to the query trapdoor t, and perform matching calculations on the query trapdoor and the candidate points in the hash bucket. Sort them from high to low according to the matching degree, and output the abstracts H(P i ) and data ciphertexts C i corresponding to the top λ encrypted indexes.
[0087] As an example, the following method is used to calculate the matching degree:
[0088]
[0089] Where match(x,y) represents the matching degree, Hamming(x,y) represents the Hamming distance calculated between x and y, H(I) represents the hash value of the index, H(t) represents the hash value of the trapdoor, and l is the length of the binary hash.
[0090] Although the present invention has been described in connection with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the description of the drawings, etc. In the specification, the term "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of cases. A single processor or other unit can implement several functions listed in the specification. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0091] Although the present invention has been described in connection with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, this specification and the drawings are merely exemplary descriptions of the present invention and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A searchable encryption data sharing method based on BK-LSH, implemented based on a consortium blockchain, characterized in that, Including: S1. Configure and encrypt the original data to generate a digest and a data ciphertext corresponding to the original data; S2. Extract multiple keywords of the original data based on a deep language learning model, construct a vector representation for each keyword, and construct a hash table of the vector representations; S3. Determine the binary sequence of the keys of the hash buckets in the hash table, group the hash buckets into several initial clusters and perform iterative optimization until the clustering is stable or reaches a preset number of iterations, and use the clustering result as the search index of the original data; S4. Upload the digest, data ciphertext, and search index corresponding to the original data to the cloud server; S5. Configure multiple query keywords, generate a query trapdoor and upload it to the consortium blockchain, and locate the corresponding hash bucket in the hash table based on the query trapdoor. Match the query trapdoor with the keyword vectors processed by the hash function in the hash bucket to obtain a matching degree score. Sort the search index from high to low according to the matching degree score, and output the digests and data ciphertexts corresponding to the top several search indexes; S6. Decrypt the data ciphertext to obtain the original data.
2. The searchable encryption data sharing method based on BK-LSH according to claim 1, wherein, The S1 includes the following sub-steps: S10. Clean the original data, and use a program to delete irrelevant characters to obtain the cleaned original data; S11. Generate a digest of the cleaned original data using a one-way encryption algorithm; S12. The data owner and data searcher in the consortium blockchain request a symmetric key from the key distribution center, and use the symmetric key to encrypt the cleaned original data to obtain a data ciphertext.
3. The searchable encryption data sharing method based on BK-LSH according to claim 1, wherein, The S2 includes the following sub-steps: S20. Extract multiple keywords of the original data based on a deep language learning model, and construct a vector representation for each keyword; S21. Generate a number of random vectors that follow a Gaussian distribution N(0,1), and the number is the number of preset hash tables multiplied by the number of hash functions of each hash table; S22. For each vector, multiply it by the above random vectors, add an offset uniformly distributed on [0, w), and then divide by the preset bucket width. The vector is assigned to the corresponding hash bucket to obtain the hash table of the vector representation.
4. The searchable encryption data sharing method based on BK-LSH according to claim 3, wherein Specifically, S3 is as follows: select the top k buckets with a larger amount of stored data from the 2 l buckets; based on the Hamming distance between the binary sequences, the remaining buckets are merged into the top k buckets according to the principle of proximity, and the k new buckets obtained after merging are used as the k initial clusters of the clustering algorithm.
5. The searchable encryption data sharing method based on BK-LSH according to claim 1, wherein The S3 includes the following sub-steps: S30. Determine a cluster representative in each of the initial clusters. A total of k cluster representatives are determined for the k initial clusters. Calculate the distances between the i-th cluster representative among the k cluster representatives and the other k-1 cluster representatives, where 1 ≤ i ≤ k. Select the k cluster representatives with relatively short distances to form a shortlist for the i-th cluster representative; ′ cluster representatives to form a shortlist for the i-th cluster representative; S31. Calculate the distance between each non-cluster representative data in the m-th initial cluster and each cluster representative in the shortlist, 1 ≤ m ≤ k, and transfer the non-cluster representative data to the initial cluster where the closer cluster representative is located to form a new cluster; S32. Repeat S30 to S31 until the clustering is stable or reaches a preset number of iterations, and jump to S33; S33. Output the clustering result at this time, which is the search index.
6. The searchable encryption data sharing method based on BK-LSH according to claim 3, characterized in that Configuring multiple query keywords and generating a query trapdoor includes: configuring multiple query keywords, constructing a vector representation for each query keyword based on a deep language learning model, and using the l binary hash functions to perform hash processing on the vector representation of each query keyword to generate a query trapdoor.
7. The searchable encryption data sharing method based on BK-LSH according to claim 1, characterized in that, In S1, the original data is encrypted using the AES symmetric encryption algorithm.
8. The searchable encryption data sharing method based on BK-LSH according to claim 1, characterized in that In S2, the deep language learning model is implemented using the BERT model.
9. A searchable encrypted data sharing system based on BK-LSH, characterized in that, Including: An encryption and upload unit, a vector extraction unit, a hash table construction unit, a clustering unit, a query unit, and a data acquisition unit; The encryption and upload unit is used to configure the original data, generate a digest of the original data, encrypt the original data to obtain a data ciphertext, and upload the digest, the data ciphertext, and the search index to the cloud server; The vector extraction unit extracts multiple keywords of the original data based on a deep language learning model and constructs a vector representation for each keyword; The hash table construction unit constructs a hash table of the vector representations; The clustering unit determines the binary sequences of the keys of 2 l buckets in the hash table, and groups the 2 l buckets into initial clusters based on the Hamming distance between the binary sequences; iteratively optimize the initial clusters until the clustering is stable or a preset number of iterations is reached, and use the clustering result at this time as the search index; upload the digest, data ciphertext, and search index to the cloud server; The query unit configures multiple query keywords and generates a query trapdoor; Based on the query trapdoor, locate the corresponding hash bucket in the hash table, perform a matching calculation on the query trapdoor and the search index in the hash bucket, sort them in descending order of the matching degree, and output the digests and data ciphertexts corresponding to the top λ search indexes; The data acquisition unit verifies the integrity of the data based on the digest, and then decrypts the data ciphertext to obtain the original data.