Vector retrieval method and device for protecting data privacy

By introducing centroid vectors and homomorphic operations into vector search services, the client and server work together to solve the problem that vector search services in the prior art are difficult to protect data privacy and optimize communication and computing overhead, and efficient and secure vector search is achieved.

CN120030064AActive Publication Date: 2025-05-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510507288.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing vector search services are difficult to optimize communications and calculate overhead while protecting data privacy, and reduce service delays.

Method used

A vector search method that protects data privacy is adopted. Through the coordinated work of the client and the server, the center of mass vector and homomorphic operation are used to realize the similarity calculation between the query vector and the object vector, reducing communication and calculation overhead.

Benefits of technology

Effectively protect the privacy information of users and data parties, reduce communication and computing overhead, reduce service delays, and improve the efficiency and security of vector retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030064A_ABST
    Figure CN120030064A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a vector retrieval method and device for protecting data privacy. The method relates to a client and a server, the server stores plaintexts or ciphertexts corresponding to each group of object vectors, and the client stores centroid vectors of each group of object vectors. The method comprises the following steps that: a client calculates the similarity between a corresponding query vector and each centroid vector according to the query input of a user to obtain a target group identifier corresponding to the highest similarity, so that a ciphertext corresponding to the query vector and a plurality of group identifiers containing the target group identifier are sent to a server; the server performs homomorphic operation based on local storage and received data to obtain a first ciphertext, and a plaintext corresponding to the first ciphertext is used for measuring the similarity between the query vector and each object vector in the p groups; and the client side decrypts the inner product ciphertext sent by the server side, and determines the object vectors of which the similarity with the query vector is ranked at the top k in the target group according to a decryption result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of data security processing technology, and in particular, to a vector retrieval method and apparatus for protecting data privacy, a computer-readable storage medium, and a computing device. Background Art

[0002] In today's various applications, vector retrieval services have become key components and are widely used in large language model (LLM) prompt word engineering, recommendation systems, search engines, and drug discovery. Vector retrieval services achieve efficient retrieval of relevant information by encoding multimodal and unstructured data into high-dimensional vector representations. In this type of workflow, the user's profile features and private queries are embedded and processed into vectors and transmitted to the vector database server. The server identifies the most relevant objects from the data owner's vector set through similarity calculations and finally returns the results to the user.

[0003] However, current vector retrieval service solutions are difficult to meet higher requirements in practical applications, such as optimizing communication and computing overhead, and reducing service latency while protecting data privacy. Summary of the invention

[0004] The embodiments of this specification describe a vector search method and device, which can meet higher requirements in practical applications.

[0005] According to a first aspect, a vector retrieval method for protecting data privacy is provided, involving a client and a server, wherein the server stores plaintext or ciphertext corresponding to each group of object vectors in m groups of object vectors, and the client stores the centroid vector of each group of object vectors. The method is applied to the client, comprising:

[0006] According to the user's query input, calculate the similarity between the corresponding query vector and the centroid vector, and obtain the target group identifier corresponding to the highest similarity. Send the ciphertext corresponding to the query vector and p group identifiers to the server; wherein p<m, the p group identifiers include the target group identifier. Receive a first ciphertext, whose corresponding plaintext is used to measure the similarity between the query vector and each object vector in the p groups; the first ciphertext is determined by the server based on its locally stored vectors and the received vectors by performing homomorphic operations. According to the decryption result of the first ciphertext, determine the similarity between the query vector and each object vector in the target group, and use it to obtain the query result corresponding to the query input.

[0007] In one embodiment, the m groups of object vectors and centroid vectors are determined by the data party based on the following steps: clustering is performed based on the object vector set to obtain m clusters and the centroid vectors of each cluster, wherein the number of vectors of each cluster is less than or equal to a preset value n. The m groups of object vectors are initialized using the m clusters, and for each group, a nearest neighbor search algorithm is used to determine a batch of vectors whose similarity with the centroid vector of the group is within a preset range from the object vector set, and the group is filled based on the batch of vectors until the number of vectors in the group reaches n and there are no duplicate vectors.

[0008] Furthermore, in a specific embodiment, clustering processing is performed based on the object vector set to obtain m clusters, including: performing a first clustering processing on the object vector set to obtain multiple first clusters; for each first cluster, when it is determined that the number of vectors therein is greater than n, performing a second clustering processing on it, and classifying the multiple second clusters obtained into the m clusters, otherwise, directly classifying the first cluster into the m clusters.

[0009] In one embodiment, the determination of the p group identifiers includes: based on the target partition identifier, randomly selecting p-1 other partition identifiers from the m group identifiers and collectively including them into the p group identifiers.

[0010] In one embodiment, the server is implemented as a third-party server, which stores the ciphertext corresponding to each group of object vectors, and the ciphertext is provided by the data party.

[0011] Furthermore, in a specific embodiment, the method also includes: receiving the ciphertext of the centroid vector from the third-party server, which is obtained by the data party encrypting the centroid vector using a symmetric key; decrypting the ciphertext of the centroid vector using the symmetric key, and storing the decrypted plaintext.

[0012] In one embodiment, the server is implemented as a private server of the data party, in which the plain texts corresponding to the m groups of object vectors are stored.

[0013] Furthermore, in a specific embodiment, the method further includes: receiving m centroid vectors corresponding to the m groups of object vectors from the private server and storing them.

[0014] In one embodiment, the first ciphertext includes an inner product ciphertext, and its corresponding plaintext includes the inner product between the query vector and each object vector in the p groups.

[0015] In one embodiment, determining the similarity between the query vector and each object vector in the target group based on the decryption result of the first ciphertext includes: for each object vector in the target group, extracting the inner product between it and the query vector from the decryption result as the corresponding similarity.

[0016] In one embodiment, the client also stores the norm of each object vector in the m groups; wherein, according to the decryption result of the first ciphertext, the similarity between the query vector and each object vector in the target group is determined, including: for each object vector in the target group, extracting the inner product between it and the query vector from the decryption result, and determining the similarity between the object vector and the query vector based on the inner product, the norm of the object vector and the norm of the query vector.

[0017] In one embodiment, after determining the similarity between the query vector and each object vector in the target group, the method further includes: determining k object IDs corresponding to the top k similarities as the query result.

[0018] According to the second aspect, a vector retrieval method for protecting data privacy is provided, involving a client and a server, wherein the server stores plaintext or ciphertext corresponding to each group of object vectors in m groups of object vectors, and the client stores the centroid vector of each group of object vectors. The method is applied to the server, comprising: receiving the ciphertext corresponding to the query vector and p group identifiers from the client; wherein p<m, the p group identifiers include a target group identifier; the centroid vector of the target group has a higher similarity to the query vector than other groups, and the query vector is determined based on the user's query input. A homomorphic operation is performed based on the locally stored vector and the received vector to obtain a first ciphertext, and its corresponding plaintext is used to measure the similarity between the query vector and each object vector in the p groups. The first ciphertext is sent to the client, so that the client determines the similarity between the query vector and each object vector in the target group based on the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

[0019] In one embodiment, the server is implemented as a third-party server; the method further includes: receiving ciphertext corresponding to each group of object vectors from a data source.

[0020] Furthermore, in a specific embodiment, the server is implemented as a private server of the data party, in which the plaintext corresponding to each group of object vectors is stored.

[0021] In one embodiment, the first ciphertext includes an inner product ciphertext, and its corresponding plaintext includes the inner product between the query vector and each object vector in the p groups.

[0022] Furthermore, in a specific embodiment, the server stores the plaintext or ciphertext corresponding to each group of object vectors; each group of object vectors includes n object vectors, and the corresponding plaintext is the first rearranged vector of the group of object vectors, and the first rearranged vector includes d groups of elements, wherein the jth element in any i-th group is the i-th dimension element of the j-th object vector in the group of object vectors; or, the ciphertext corresponding to each group of object vectors is obtained by encrypting the first rearranged vector using the public key of the user. The ciphertext corresponding to the query vector is obtained by encrypting the second rearranged vector using the public key, and the second rearranged vector includes d groups of elements, wherein the elements of any i-th group are the i-th dimension elements of the query vector.

[0023] Based on this, the homomorphic operation includes: for any group in the p groups, performing a homomorphic multiplication operation on the plaintext or ciphertext of the first rearranged vector of the group and the ciphertext of the second rearranged vector to obtain the current vector ciphertext; performing a logarithmic multiplication operation on the current vector ciphertext 2 d iterations, wherein any sth iteration includes: performing a step of n*d / 2 on the current vector ciphertext s cyclic rotation; updating the current vector ciphertext to the result of homomorphic addition operation between the current vector ciphertext and the result of rotation of the vector; determining the inner product ciphertext based on the iterated current vector ciphertext, wherein the first n elements of the plaintext corresponding to the current vector ciphertext are the inner products between the query vector and each object vector in the corresponding group.

[0024] Furthermore, in one example, the inner product ciphertext is determined based on the iterated current vector ciphertext, including: performing homomorphic multiplication operation on the current vector ciphertext using a mask vector to obtain a mask result ciphertext; the first n bits of the mask vector are predetermined non-zero values, and the other bits are zero. The p mask result ciphertexts corresponding to the p groups are merged, specifically including: for the i-th mask result ciphertext, performing a cyclic rotation with a step length of (i-1)*n bits; performing a homomorphic addition operation on the p rotation results corresponding to the p mask result ciphertexts to obtain the inner product ciphertext.

[0025] According to a third aspect, a vector retrieval device for protecting data privacy is provided, which is integrated in a client, and its corresponding server stores the plaintext or ciphertext corresponding to each group of object vectors in m groups of object vectors, and the client stores the centroid vector of each group of object vectors; the device includes:

[0026] The target group determination unit is configured to calculate the similarity between the query vector corresponding to the user's query input and the centroid vector, and obtain the target group identifier corresponding to the highest similarity. The query data sending unit is configured to send the ciphertext corresponding to the query vector and p group identifiers to the server; wherein p<m, the p group identifiers include the target group identifier. The inner product ciphertext receiving unit is configured to receive a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; the first ciphertext is determined by the server based on the locally stored vector and the received vector by performing homomorphic operations. The query result determination unit is configured to determine the similarity between the query vector and each object vector in the target group according to the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

[0027] According to a fourth aspect, a vector retrieval device for protecting data privacy is provided, which is integrated in a server, wherein the server stores the plaintext or ciphertext corresponding to each group of object vectors in m groups of object vectors, and a client corresponding to the server stores the centroid vector of each group of object vectors. The device comprises:

[0028] The query data receiving unit is configured to receive the ciphertext corresponding to the query vector and p group identifiers from the client; wherein p<m, the p group identifiers include a target group identifier; the similarity between the centroid vector of the target group and the query vector is higher than that of other groups, and the query vector is determined based on the user's query input. The first ciphertext determination unit is configured to perform homomorphic operations based on the locally stored vectors and the received vectors to obtain a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups. The first ciphertext sending unit is configured to send the first ciphertext to the client, so that the client determines the similarity between the query vector and each object vector in the target group based on the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

[0029] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method provided in the first aspect or the second aspect.

[0030] According to a sixth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method provided in the first aspect or the second aspect is implemented.

[0031] In summary, the above method and device disclosed in the embodiments of this specification are adopted to introduce FHE to realize data privacy protection for users and data parties, and it is proposed to move the selection of branches to the user side, so that the range of vector similarity search can be determined before reaching the server. Through this design, the user side and the server side jointly construct a complete vector index structure, which effectively makes up for the shortcomings of FHE in complex logical judgment. Furthermore, it is proposed to simplify the distance or similarity calculation in FHE to the inner product calculation, and optimize the inner product calculation mode to effectively reduce the calculation complexity. Furthermore, it is proposed to compress and merge the result ciphertext corresponding to each partition to obtain the inner product ciphertext, thereby improving the space utilization of the ciphertext slot and effectively reducing the communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 Schematic diagram of the interactive architecture of the vector search service in plain text mode; Figure 2 The illustrated data is the interactive architecture of the vector search service in the dense form mode; Figure 3 It illustrates the trusted wall design implemented by using encryption algorithm in the improved solution disclosed in the embodiment of this specification; Figure 4 The basic mode of calculating the inner product between vector ciphertexts based on FHE disclosed in the embodiments of this specification is illustrated; Figure 5 The client-server collaborative structure disclosed in the embodiments of this specification is illustrated; Figure 6 The service model architecture of the vector security search service disclosed in the embodiments of this specification is illustrated; Figure 7 It illustrates the optimization mode of calculating the inner product between vector ciphertexts based on FHE disclosed in the embodiments of this specification; Figure 8 The embodiment disclosed in this specification Figure 7 The process of merging and compressing the result ciphertext in ; Fig. 9 This is a functional structure diagram of a vector search device integrated in a client disclosed in an embodiment of this specification; Fig.10 This is a functional structure diagram of a vector search device integrated in a server according to an embodiment of this specification. DETAILED DESCRIPTION

[0034] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0035] As mentioned above, the current vector retrieval (or search) service solutions are difficult to meet the higher requirements in practical applications. Specifically, the current service form poses serious security and privacy risks to users and data owners, or data owners, data providers, data parties, etc., especially when these systems and products are deployed in the form of hosted cloud services.

[0036] See also Figure 1 ,There are three main types of data leakage risks in the existing vector search system architecture: ,First, user privacy queries are submitted to the public cloud server in plain text; ,Second, the query data is transmitted to the data owner's private server, ,both of which will result in sensitive user information being exposed to the service ,provider. Third, when outsourcing the database, the data owner's vector collection (usually containing ,sensitive or proprietary information) will be leaked to the public cloud service provider, ,such disclosure will pose a potential threat to the data owner's intellectual property ,and commercial interests.

[0037] However, there is no practical solution that can protect the privacy information of users and data owners at an acceptable cost and practical performance to build such Figure 2 Security vector database for untrusted scenarios shown.

[0038] To build Figure 2 The vector security search architecture shown in the figure proposes the following two improvement directions: 1) Vector retrieval based on secure multiparty computation (MPC). MPC is a method for building secure protocols with low computational requirements but high communication costs due to the large number of interactions using oblivious transfers and garbled circuits.

[0039] 2) Vector retrieval based on Fully Homomorphic Encryption (FHE). FHE supports direct calculations on encrypted data without redundant interactions. Although it can significantly reduce communication overhead, it will bring a huge computational burden.

[0040] Furthermore, the embodiments of this specification propose an improved FHE-based vector retrieval scheme (or improved scheme for short), which can provide privacy protection for users and data parties, and at the same time, can effectively reduce communication and computing overheads.

[0041] Next, we will introduce the actual application scenarios and define the threat model. We will then discuss the cost and practical performance challenges faced when applying FHE to vector retrieval. We will then explain the design considerations and overall architecture of the improvement solution. Finally, we will give the specific implementation process of the improvement solution.

[0042] 1. Actual Application Scenarios and Threat Models 1. Actual application scenarios The improvement plan aims to protect the privacy and security of users and data holders in the vector retrieval service. Its architecture includes the following three core roles.

[0043] 1) User side (or client side): The user submits a query vector to the system to obtain similar object vectors. Since the leakage of the query vector may cause the user to be harassed by advertisements or suffer economic losses, it is necessary to ensure that the user's query vector is always kept confidential and cannot be obtained by other parties.

[0044] 2) Data Provider: Pharmaceutical, financial, e-commerce and other companies provide data sets as data providers, which can be deployed on the data provider's private server or external server. Especially when data is deployed on external servers, the confidentiality of sensitive data must be ensured to prevent unauthorized access. It should be understood that external servers can also be referred to as public servers, third-party servers, etc.; in addition, the third party in this article refers to parties other than users and data providers.

[0045] 3) Server: Responsible for vector retrieval processing and data storage. The server can be a server managed by a third-party service provider (such as a cloud server). In this case, for the data subject, there is a risk of data leakage caused by deploying its data on a third party. Alternatively, the server can be the data subject's private server, in which case there is no such risk. The following introduction focuses on the more common situation in actual applications, that is, the former situation with leakage risk.

[0046] 2. Threat Model The following two potential threats are considered in vector security retrieval.

[0047] 1) User privacy threats: User privacy faces dual threats from data providers and public servers. When users send query vectors to public servers for processing, both public servers and data providers may use these data to infer sensitive information about users. This may lead to the risk of users’ search history and related preferences being exposed.

[0048] 2) Confidentiality threats to data parties: Data parties’ datasets are primarily threatened by public servers and external attackers. When these datasets are hosted on public servers, they are vulnerable to unauthorized access by third-party service providers or other malicious attackers, and there is a risk of leaking proprietary or sensitive data. For example, a dataset includes object vectors, or encoding vectors, corresponding to multiple business objects.

[0049] The improvement scheme aims to adopt effective encryption measures to deal with the above threats and protect the interests of users and data parties, while maintaining efficient and secure vector similarity search capabilities.

[0050] 2. Basic solutions and challenges 1. Solution design like Figure 3 As shown, the design goal is to establish two trusted barriers, or two trusted walls, for users and data parties.

[0051] Table 1 below briefly shows the process of Algorithm 1, ie, the Approximate Nearest Neighbor Search (ANNS) algorithm.

[0052] Table 1

[0053] In order to protect the privacy of user queries, when the user needs to send its query vector When it reaches the server, it will first use its public key pk to Encrypted to ciphertext . Complete homomorphic FHE operation on the cloud server After that, the result of the operation in ciphertext form It is returned to the user end, and can be decrypted into plain text using the private key sk held only by the user end. Thus, the user's query information is protected.

[0054] It should be understood that evk represents the evaluation key, rk represents the relinearization key, and gk represents the Galois key. These keys are used for FHE calculations. For a detailed introduction to these keys, please refer to the introduction to FHE algorithms such as the CKKS (Cheon-Kim-Kim-Song) algorithm or the BFV (Brakerski-Fan-Vaikuntanathan) algorithm, which will not be expanded here.

[0055] To protect the data privacy of the data party, the data party first encrypts the data set into ciphertext using the user's public key pk before deploying the collected data set to an untrusted public server. It should be understood that if a service is provided to h users, encryption is required h times.

[0056] Based on the above design, two trusted walls are established to ensure the data security of users and data parties. In addition, the cloud server holds the keys: evk, rk, and gk, so that FHE calculations can be performed to obtain the distance ciphertext or similarity ciphertext between vector ciphertexts without leaking any information.

[0057] 2. Challenges Although FHE provides strong security guarantees and can perform computations directly on ciphertext, it introduces additional overhead in many aspects of the process of Algorithm 1, including the computational overhead in line 4, the communication overhead in line 5, and the index structures involved in lines 3 and 7. Next, these challenges will be discussed in turn.

[0058] 1) Computational overhead Introducing FHE into ANNS will result in limited programmability and incur huge computational overhead.

[0059] The current FHE library is not designed for specific applications and can only be used to implement FHE operation primitives, including homomorphic addition, homomorphic multiplication, and circular rotation operations, which are all single instruction multiple data (SIMD) instructions. In addition, individual encrypted elements in the ciphertext cannot be indexed, and any function acting on a vector ciphertext will act on all its vector elements at the same time.

[0060] Inner Product is not only a common type of distance calculation, but also a major prerequisite for other complex distances (such as L2 norm and cosine distance). When providing vector retrieval services under the FHE architecture, each service involves converting the inner product calculation into a series of FHE operation primitives on large-scale vectors. Figure 4 The basic pattern of calculating the inner product between vector ciphertexts is shown in Figure 1 to intuitively illustrate the overhead incurred by FHE in calculating the inner product.

[0061] Figure 4 Assume that the vector dimension is 4. Due to the characteristics of SIMD, the query vector ( ) is copied into multiple copies and filled into the ciphertext with the same length as the ciphertext of other vectors. Figure 4 In the figure, other vectors are n object vectors: }, , , it can be understood that the object vector has the same dimension as the query vector. When in the dense form, the calculation Figure 4 When calculating the inner product between the query vector shown in the first line and the n object vectors shown in the second line, the following steps are included:

[0062] Mult(), i.e. homomorphic multiplication operation: First, a homomorphic multiplication operation is performed between the query ciphertext and the object vector ciphertext. The result of the operation can be seen in Figure 4 The third line, or the first blue line, shows that the plain text corresponding to the result of the Mult() operation is: .

[0063] Rotate(), i.e., the circular rotation operation of the vector: Then, based on the above Mult() operation result, three circular rotation operations are performed to convert each , All elements in are moved to the same index positions as the vector ciphertext, which are in Figure 4 Circled in red.

[0064] Add(), i.e. homomorphic addition operation: Then, the four blue vector ciphertexts are homomorphically added, and the elements are accumulated into the bottom ciphertext to obtain all , ,like Figure 4 As shown in line 7 of .

[0065] The above describes the calculation process in the basic mode. This process can be extended from 4 dimensions to any d dimensions. Each execution of this calculation process requires at least one Mult(), Rotate() and Add() operation.

[0066] Experiments show that the overhead of a single FHE operation is in the microsecond to millisecond range. However, for ANNS queries, since the similarity calculation under a vector search involves tens of thousands of FHE operations, the overhead of the calculation process under the above basic mode is relatively large.

[0067] 2) Communication overhead In FHE algorithms such as CKKS, the size of the vector expands several times when it is converted from plaintext to encrypted. Figure 4 The IP distance calculation result at the bottom is Figure 4 In line 7, the valid slots of the ciphertext are only a part of the total slots, which leads to a huge waste of space and thus a waste of communication overhead.

[0068] 3) Index structure dilemma In ANNS, almost all vector index structures contain unavoidable logical branches, which lead to the contradiction between dynamics and confidentiality. For clustering or tree-based ANNS indexes, when a query request arrives, the system selects the cluster closest to the query vector or the region where the query may be located in multiple hyperplanes, calculates the similarity distance, and scans nearby vectors to obtain the top k most similar results.

[0069] However, in an encrypted scenario, it is difficult to find neighboring clusters and possible hyperplanes, because this process requires selecting and judging directions based on the known distances between the query vector and each logical branch. Specifically, selecting neighboring clusters requires comparing the distances between all centroid vectors and the query vector, while finding possible hyperplanes requires knowing each dimension of the query vector. It is important to understand that the centroid vector of a cluster can be the average vector of all vectors in the cluster. However, the query vector and distance should not be visible to the server, which is a dilemma that is difficult to reconcile. The graph-based ANNS index makes the situation more complicated. When the query vector arrives, the search process starts at the entrance of the graph. Each time the search path passes through a point in the graph, it selects the nearest point among all the connected points and continues to move forward in a greedy manner. This means that each time a new point is reached through a connecting edge, a direction selection must be made among multiple branches, and this selection depends on the known distance and the query data in plain text.

[0070] Word-wise FHE only provides SIMD vector computing capabilities and cannot handle branching issues. Although bit-wise FHE theoretically supports the implementation of any logic through circuit gates, it is difficult to implement complex choices in various dynamic ANNS strategies by relying solely on static Boolean circuits.

[0071] Therefore, when faced with unavoidable branches, one can either choose to perform brute force distance calculations on the query vector and all encrypted vectors in the database, or send the encrypted data back to the user every time a branch is encountered, allowing the user to decrypt and make a choice. However, both methods will increase the computational and communication overheads.

[0072] The above introduces the challenges faced by the process of Algorithm 1. Next, the design considerations and overall architecture of the improvement plan are explained.

[0073] 3. Design considerations for the improvement plan In order to address the three major challenges of communication, computing, and indexing structure when FHE is applied to ANNS, the applicants formulated the following design considerations when developing a privacy-preserving vector retrieval system.

[0074] 1. Optimize distance calculation based on FHE Figure 4 The calculation process in the basic mode is shown, and its computational complexity increases linearly with the increase of vector dimension. Therefore, it is proposed to reduce the number of FHE operations in a single vector search.

[0075] 2. Improve the utilization rate of ciphertext slots during communication Due to the inherent characteristics of FHE technology, the expansion of message size is inevitable after its introduction. In order to reduce the communication burden, it is proposed to make full use of all slots in the ciphertext. dimensional query vector and indivual When the inner product between dimensional object vectors is generated, the generated inner product ciphertext contains slots, but in reality only The inner product result scalar means a serious waste of space in the ciphertext. Therefore, it is proposed to fill more inner product results into the ciphertext slots to achieve high utilization of the ciphertext slots, thereby reducing the number of ciphertexts and alleviating the communication burden. Experiments show that the communication burden can be reduced to less than 1 / 10.

[0076] 3. Delegate the branching of the index structure to the user end The above analysis of the logical branches of various ANNS index types points out that it is not feasible to process unknown distances in ciphertext for branch selection. Therefore, it is proposed to move the branch selection to the user side, so that the range of vector similarity search can be determined before reaching the server. Through this design, the user side and the server side jointly build a complete vector index structure, which effectively makes up for the shortcomings of FHE in complex logical judgment. After actual measurement and evaluation, the load performance of the user side is good, and the memory usage and CPU running time are both within the actual acceptable range.

[0077] 4. Overall Architecture of the Improvement Plan The following introduces the client-server collaborative architecture and service model in the improvement plan.

[0078] 1. Client-server collaborative architecture. Considering the inevitable branch selection and large communication overhead, it is proposed to move the branch selection to the user side and reduce the communication rounds between the user side and the server side. Although the graph-based ANNS index shows better performance in plaintext mode, its dynamic greedy strategy requires multiple rounds of communication when determining the next operation direction, or huge overhead is generated by comparing the ciphertext.

[0079] Therefore, in the improved scheme, it is proposed to use a partition-based ANNS index, which is similar to the cluster-based and tree-based ANNS index schemes. It should be understood that the partitions can also be called groups, and similar object vectors are pre-divided into the same partition or the same group. The m groups divided based on the object vector set (or all object vectors) of the data square can be instantiated as multiple clusters in the cluster-based ANNS index, or multiple hyperplanes in the tree-based ANNS index.

[0080] like Figure 5 As shown, the user side and the server side ( Figure 5 A trusted wall based on FHE is established between the two nodes (shown as cloud servers in the figure). The user terminal holds multiple partition IDs, for example, {1,2,...,n-1,n}, and the centroid vectors in plain text corresponding to each partition ID. When the user terminal starts an ANNS query, the nearest partition is selected by calculating the similarity between the query vector and the n centroid vectors, for example, {3,i,n-2}. The number of partitions selected p can be 1 or more than 1.

[0081] In the case where the number of selected partitions p ≥ 2, in one embodiment, p centroid vectors with the top p similarities to the query vector can be determined, and the corresponding p partitions can be selected. In another embodiment, in addition to selecting a nearest neighbor partition, a definable coefficient is provided to the user to introduce randomly generated p-1 misleading (or confusing) partition IDs into the final selected p partitions. In this case, the selected p partitions reveal the fuzzy range related to the query, which can better protect the privacy of the user's query.

[0082] After encrypting the query vector to obtain the corresponding query ciphertext, the user sends the query ciphertext and the selected p partition identifiers to the server. The server reads the plaintext or ciphertext corresponding to the object vector under each of the p partitions from the local according to the received p partition identifiers. The inner product ciphertext can be calculated using FHE and then returned to the user.

[0083] Finally, the user can decrypt the returned inner product ciphertext locally to obtain the k objects closest to his query.

[0084] It should be noted that the following mainly introduces the first ciphertext calculated by the server using FHE and returned to the client as the inner product ciphertext mentioned above. In practice, it is not limited to the inner product ciphertext. It only needs the corresponding plaintext to be used to measure the similarity between the query vector and each object vector in the p groups. The specific design can be based on the existing similarity index. In addition, it should be noted that the "first" in the "first ciphertext" and similar terms such as "second" in other places in the text are all for distinguishing similar things and do not have other limiting functions such as sorting.

[0085] The above introduces the client-server collaborative architecture in the improvement plan.

[0086] 2. Service model like Figure 6 As shown, the FHE algorithm is illustrated as the CKKS algorithm, and the CKKS algorithm is used to encrypt the vectors involved in homomorphic operations and perform inner product calculations on the ciphertext. In addition, the Advanced Encryption Standard (AES) is used to encrypt the binary bytes of the centroid vector corresponding to each partition to protect the information security of the partition. It should be understood that the FHE algorithm is not limited to the CKKS algorithm; the key used to encrypt and decrypt the centroid vector is not limited to the symmetric key under AES, although it can reduce the encryption and decryption burden for the centroid vector compared to reusing the key of the CKKS algorithm.

[0087] The data party first Figure 6 The entire vector index is deployed on a cloud server (shown in the figure). If the server is not a private server but a public server, the index data will be uploaded to the public server after encryption. Then, a connection is established between the client and the server to start the vector retrieval service. During the connection establishment process, the centroid vector is transmitted to the client. If the centroid vector comes from a public server, its corresponding ciphertext will be decrypted into the centroid vector plaintext on the client. If the distance (or similarity) is calculated using the L2 norm or cosine distance, the size of the centroid vector will also be transmitted to the client to reduce the computational overhead.

[0088] After the connection is established, the user can encrypt the query vector using the CKKS public key pk and determine the nearest partitions (referring to one or more) locally. Then, the query ciphertext and the IDs of the partitions will be provided to the server. On the server, the FHE engine calculates the inner product ciphertext between the query ciphertext and the object vector ciphertext. Figure 6 The inner product ciphertext is shown as CKKS (distance) in the figure, and the inner product ciphertext is returned to the user. Finally, the user decrypts the inner product ciphertext to obtain the k objects most similar to the query vector.

[0089] The above introduces the overall architecture of the improvement plan.

[0090] V. Specific implementation of the improvement plan In this section, we first discuss the design of indexes that reflect the characteristics of ciphertexts. Figure 4 The basic mode of calculating the inner product ciphertext based on FHE is shown in the figure, an optimization mode is proposed, and it is explained how to compress the inner product ciphertext to reduce the traffic overhead of network communication.

[0091] 1. Index building to achieve balanced and fully utilized partitions.

[0092] As analyzed above, it is proposed to apply partition-based index as the basis for the improvement scheme. Furthermore, it is proposed that the simple partitioning method can be tuned to better cooperate with FHE calculation.

[0093] For multiple object vectors held by the data party, such as multiple commodity vectors corresponding to multiple commodities, or multiple molecular vectors corresponding to multiple drug molecules, the conventional method of grouping similar vectors into the same partition or the same group may cause the problem of unbalanced partitions. When combining partition-based ANNS with FHE calculation, the unbalanced partitions will lead to serious waste of ciphertext slots and invalid calculations in subsequent SMIDs.

[0094] In order to make the partitions balanced and fully filled, it is proposed to use a clustering algorithm, especially a hierarchical K-Means clustering algorithm, to obtain balanced partitions and fill the partitions with nearby vectors until the maximum ciphertext length is reached.

[0095] Specifically, clustering can be performed based on the object vector set to obtain m clusters and the centroid vector of each cluster, wherein the number of vectors of each cluster is less than or equal to a preset value n (or denoted as max); then, the m clusters are used to initialize m groups of object vectors, and for each group, a nearest neighbor search algorithm is used to determine a batch of vectors whose similarity to the centroid vector of the group is within a preset range from the object vector set, and the group is filled based on the batch of vectors until the number of vectors in the group reaches n and there are no duplicate vectors. It should be understood that the centroid vector of a group can be the average vector of all object vectors in the group.

[0096] Further, in a specific embodiment, the clustering process may include: performing a first clustering process on the object vector set to obtain a plurality of first clusters; for each first cluster, when it is determined that the number of vectors therein is greater than n, performing a second clustering process on it, and classifying the obtained plurality of second clusters into the m clusters, otherwise, directly classifying the first cluster into the m clusters. Exemplarily, the clustering algorithms used in the first and second clustering processes may both be K-Means algorithms. In this way, m clusters with balanced vector distribution are obtained.

[0097] In a specific embodiment, the nearest neighbor search algorithm used when filling each partition may be a K-Nearest Neighbors Algorithm (KNN), etc. In addition, during filling, some heuristic filling methods may be used for optimization.

[0098] Table 2 below gives an example of an algorithm for determining index partitions.

[0099] Table 2

[0100] In Algorithm 2, the first line of code indicates that K-Means clustering is performed on the object vector set V to obtain the initial partition result P. Lines 2-5 of code indicate that for each partition p∈P, if the number of vectors in the partition is greater than max, K-Means clustering is performed on p, and p is replaced by the cluster obtained by clustering it. In this way, the size of all partitions can be controlled within max, achieving partition balance.

[0101] Lines 6-13 of the code indicate that all partitions are filled with the neighboring vectors of the centroid of each partition, and the size of each filled partition is max. Lines 7-8 of the code indicate that when the number of vectors in p is less than max, the centroid of p is determined from the set V based on the KNN algorithm. The max nearest vectors form a set F. Lines 9-11 of code indicate that p is filled with F, and the filled vector f is a vector that was not originally in p to prevent subsequent repeated calculations. In addition, the filled p contains max vectors.

[0102] 2. FHE Computing Optimization against Figure 4 The heavy burden of FHE computation and communication in the basic mode shown above is proposed to optimize the mode. First, most similarity or distance metrics commonly used in vector retrieval are based on inner product (IP), including L2 distance and cosine distance. When obtaining the first k nearest neighbors, only the inner product, or the inner product and norm, is needed in the comparison process. Therefore, the improved scheme proposes to transmit the inner product as FHE ciphertext to the user end for the determination of k neighboring objects.

[0103] To help understand the algorithm flow of the optimized mode, Algorithm 3 is shown below, which shows the basic process of inner product calculation using FHE primitives in the basic mode. Given a d-dimensional query vector (the first line of code) and a single vector partition containing 𝑛 d-dimensional vectors (the second line of code), the inner product between the query vector and these 𝑛 vectors is calculated.

[0104] Table 3

[0105] In Algorithm 3, multiple copies of the query vector are needed to fill the query ciphertext for subsequent operations. First, the Mult multiplication operation is performed (line 4 of code). Then, in order to obtain the inner product result Due to the SIMD characteristics of FHE, the elements of each dimension of the product of the query vector and each vector need to be rotated to the same slot index position. In the specific implementation: it takes one Rotate operation to move a single-dimensional element of each vector to the target slot. After all dimensions of the vector are shifted to the specified slots, you need to pass The final result is obtained by performing additions and rotations. Given that a single operation has a millisecond latency, an optimized computing mode must be designed to reduce the number of FHE operations.

[0106] In the optimized mode, the data layout of the vector ciphertext and the FHE inner product calculation mode are optimized to reduce the number of calls to the FHE SIMD primitive. As shown in Table 4 below, the basic mode has the following vector dimensions: The optimized mode reduces the number of Add() and Rotate() operations to Second-rate.

[0107] Table 4: Statistics of the number of FHE operations during inner product calculation

[0108] Both modes call Mult() once for each ciphertext. The optimized FHE calculation process is as follows: Figure 7 shown.

[0109] Unlike the basic mode, the optimized mode restructures the data layout to arrange the elements of each vector dimension compactly instead of in order. Figure 7 In the example, the query vector is rearranged into the second rearranged vector shown in the first row, which includes d groups of elements arranged consecutively, and any element in the i-th group is the query vector The i-th element of It should be noted that the second rearrangement vector is determined at the user end, and after the user end encrypts the second rearrangement vector using its public key (such as the user public key under the CKKS algorithm), the user end provides the obtained ciphertext (or query ciphertext) to the server end.

[0110] exist Figure 7 In , for any group of object vectors, the n object vectors contained therein are rearranged into the first rearranged vector shown in the second row, which also includes d groups of elements, and the j-th element in any i-th group is the i-th element of the j-th object vector in the group of object vectors It should be noted that the first rearrangement vector may be stored in plain text on the private server of the data party, or the data party may determine the first rearrangement vector locally, encrypt it using the user's public key, and then provide it to the public server for storage.

[0111] The above optimized data layout can make the FHE calculation between encryption vectors more efficient.

[0112] Specifically, first, perform homomorphic multiplication on the query ciphertext and the plaintext (or ciphertext) of the first rearranged vector to obtain the current vector ciphertext. For this, see Figure 7 Then, log the current vector ciphertext 2 d iterations, wherein any sth iteration includes: 1) performing a step of n*d / 2 on the current vector ciphertext s 1) updating the current vector ciphertext to the result of performing a homomorphic addition operation on the current vector ciphertext and the result of the rotation of the vector.

[0113] Figure 7 In the first iteration, the second half of the homomorphic multiplication ciphertext in row 3 is rotated using a rotation step of Move it to the beginning, then add the original ciphertext on line 3 to the rotated ciphertext on line 4. After that, this process is repeated, with the rotation step length of each step adjusted to half of the previous rotation length. After each rotation, add the ciphertext before and after the rotation again. When , the iteration process ends and we get Figure 7 The first n elements of the plaintext corresponding to the ciphertext in the second to last row are the n inner product scalars between the query vector and the n object vectors corresponding to the second row.

[0114] Furthermore, the inner product ciphertext to be returned to the user terminal can be determined based on the iterated current vector ciphertext. In one embodiment, the iterated p current vector ciphertexts corresponding to the p groups can be directly returned to the user terminal as the inner product ciphertext.

[0115] In another implementation, the mask vector may be used to extract valid information from the current vector ciphertext, so as to subsequently compress the ciphertext, thereby using the compressed ciphertext as the inner product ciphertext. Specifically, Figure 7 The figure shows a homomorphic multiplication operation between the second to last row and the mask vector, retaining only the elements , and set all other ciphertext slots to 0 for subsequent compression. It should be noted that the first n bits of the mask vector are predetermined non-zero values ​​(such as 1), and the other bits are zero.

[0116] 3. Ciphertext compression for Figure 7For each calculated distance ciphertext, or mask result ciphertext, IP distance ciphertext, IP similarity ciphertext, rotate its valid elements from the starting position to other slots, and fill the remaining positions with zeros. Specifically, first, for the mask result ciphertext corresponding to any i-th group in the p groups, perform a cyclic rotation with a step size of (i-1)*n bits; then, perform a homomorphic addition operation on the p rotation results corresponding to the p mask result ciphertexts to obtain the inner product ciphertext.

[0117] like Figure 8 As shown in , in order to fully utilize all slots in a single ciphertext, the Rotate() operation uses an alternating permutation pattern and is performed across 𝑛 steps. Finally, to achieve compression, FHE Add() is applied to all rotated ciphertexts to merge them into a single ciphertext containing multiple masked results. This compression method improves space utilization from It is increased to nearly 100%, significantly reducing the waste of ciphertext slots and alleviating communication overhead.

[0118] It should be noted that after receiving the inner product ciphertext sent by the server, the user end can use the private key to decrypt the inner product ciphertext, regardless of whether it is compressed or not, and obtain the decryption result including the inner product of the query vector and each object vector in the p groups. Thus, the inner product of the query vector and each object vector in the target group can be extracted, and then the object vector in the target group with the top k similarity with the query vector can be determined based on this inner product as the query result.

[0119] Regarding how to calculate the corresponding similarity or distance based on the inner product, in one implementation, the calculated similarity is set to be cosine similarity. Further, in one embodiment, if the query vector and the object vector are both normalized vectors (that is, the modulus of the vector is 1), at this time, the inner product can be directly used as the corresponding cosine similarity. In another embodiment, if one of the query vector and the object vector is not a normalized vector, at this time, it is necessary to use the inner product divided by the norm of the object vector corresponding to the inner product and the query vector (or the modulus of the vector, the vector size), and the obtained quotient is used as the corresponding cosine similarity. It should be understood that the norm of the query vector can be calculated locally by the user end, and the norm of the object vector can be provided by the server when providing the centroid vector to the client.

[0120] In another embodiment, the calculated distance is set to be the L2 norm. In this case, according to the definition of the L2 norm between two vectors: (1) Assume that a and b are the query vector and object vector respectively, then That is, the square of the size of the query vector, is the inner product between two vectors, is the square of the magnitude of the object vector.

[0121] Therefore, based on formula (5), the L2 norm between the query vector and the object vector can be calculated according to their respective norms and the inner product between the two vectors.

[0122] It can be understood that it is feasible to calculate both similarity and distance. The distance is negatively correlated with the similarity. Selecting the object vector with the top k similarities means selecting the object vector with the bottom k distances.

[0123] From the above, the user end can complete the determination of the top-k similar object vectors as the query result, or further determine the corresponding k object identifiers as the query result.

[0124] In summary, the vector retrieval service method proposed in the embodiment of this specification is adopted, and FHE is introduced to realize data privacy protection for users and data parties. In addition, it is proposed to move the selection of branches to the user side, so that the range of vector similarity search can be determined before reaching the server. Through this design, the user side and the server side jointly construct a complete vector index structure, which effectively makes up for the shortcomings of FHE in complex logical judgment. Furthermore, it is proposed to simplify the distance or similarity calculation in FHE to the inner product calculation, and optimize the inner product calculation mode to effectively reduce the calculation complexity. Furthermore, it is proposed to compress and merge the result ciphertext corresponding to each partition to obtain the inner product ciphertext, thereby improving the space utilization of the ciphertext slot and effectively reducing the communication overhead.

[0125] Corresponding to the above-mentioned vector search method, the embodiment of this specification also discloses a vector search device. The details are as follows:

[0126] Fig. 9 The vector retrieval device 900 shown in FIG. 1 is integrated into a client, and the server corresponding to the client stores the plaintext or ciphertext corresponding to each group of object vectors in the m groups of object vectors, and the client stores the centroid vector of each group of object vectors. Fig. 9 As shown, the vector search device 900 includes the following functional modules:

[0127] The target group determination unit 910 is configured to calculate the similarity between the query vector corresponding to the user's query input and the centroid vector, and obtain the target group identifier corresponding to the highest similarity. The query data sending unit 920 is configured to send the ciphertext corresponding to the query vector and p group identifiers to the server; wherein p<m, the p group identifiers include the target group identifier. The inner product ciphertext receiving unit 930 is configured to receive a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; the inner product ciphertext is determined by the server based on its locally stored vectors and the received vectors by performing homomorphic operations. The query result determination unit 940 is configured to determine the similarity between the query vector and each object vector in the target group according to the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

[0128] In one embodiment, the m groups of object vectors and centroid vectors are determined by the data source based on the following steps:

[0129] Based on the object vector set, clustering processing is performed to obtain m clusters and the centroid vector of each cluster, wherein the number of vectors of each cluster is less than or equal to a preset value n. The m clusters are used to initialize m groups of object vectors, and for each group, a nearest neighbor search algorithm is used to determine a batch of vectors whose similarity with the centroid vector of the group is within a preset range from the object vector set, and the group is filled based on the batch of vectors until the number of vectors in the group reaches n and there is no duplicate vector.

[0130] Further, in a specific embodiment, performing clustering processing based on the object vector set to obtain m clusters includes: performing a first clustering processing on the object vector set to obtain multiple first clusters. For each first cluster, if it is determined that the number of vectors therein is greater than n, performing a second clustering processing on it, and classifying the obtained multiple second clusters into the m clusters, otherwise, directly classifying the first cluster into the m clusters.

[0131] In one embodiment, the vector retrieval device 900 further includes: a query data determination module configured to randomly select p-1 other partition identifiers from the m group identifiers based on the target partition identifier, classify them into the p group identifiers, and perform encryption processing based on the query vector to obtain a query ciphertext.

[0132] In one embodiment, the server is implemented as a third-party server, which stores the ciphertext corresponding to each group of object vectors, and the ciphertext is provided by the data party.

[0133] Furthermore, in a specific embodiment, the vector retrieval device 900 also includes: a centroid ciphertext receiving unit, configured to receive the ciphertexts of the respective centroid vectors from the third-party server, which are obtained by the data party encrypting the respective centroid vectors using a symmetric key; a centroid ciphertext decryption unit, configured to decrypt the ciphertexts of the respective centroid vectors using the symmetric key, and store the decrypted plaintext.

[0134] In one embodiment, the server is implemented as a private server of the data party, in which the plain texts corresponding to the m groups of object vectors are stored.

[0135] Furthermore, in a specific embodiment, the vector retrieval device 900 further includes: a centroid receiving unit configured to receive m centroid vectors corresponding to the m groups of object vectors from the private server and store them.

[0136] In one embodiment, the first ciphertext includes an inner product ciphertext, and the plaintext corresponding to the inner product ciphertext includes the inner product between the query vector and each object vector in the p groups.

[0137] In one embodiment, the query result determination unit 940 is specifically configured to: for each object vector in the target group, extract the inner product between the object vector and the query vector from the decryption result as the corresponding similarity.

[0138] In one embodiment, the client also stores the norm of each object vector in the m groups; the query result determination unit 940 is specifically configured as: for each object vector in the target group, extracting the inner product between it and the query vector from the decryption result, and determining the similarity between the object vector and the query vector based on the inner product, the norm of the object vector and the norm of the query vector.

[0139] In one embodiment, the query result determination unit 940 is further configured to: determine k object IDs corresponding to the top k similarities as the query result.

[0140] Fig.10 The vector retrieval device 1000 shown in FIG. 1 is integrated into a server, which stores the plaintext or ciphertext corresponding to each group of object vectors in m groups of object vectors, and the client corresponding to the server stores the centroid vector of each group of object vectors. Fig.10 As shown, the vector search device 1000 includes the following functional modules:

[0141] The query data receiving unit 1010 is configured to receive the ciphertext corresponding to the query vector and p group identifiers from the client; wherein p < m, the p group identifiers include a target group identifier; compared with other groups, the similarity between the centroid vector of the target group and the query vector is higher, and the query vector is determined based on the user's query input. The first ciphertext determining unit 1020 is configured to perform homomorphic operations based on the locally stored and received data to obtain a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups. The first ciphertext sending unit 1030 is configured to send the first ciphertext to the client, so that the client determines the similarity between the query vector and each object vector in the target group based on the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

[0142] In one embodiment, the server is implemented as a third-party server. The vector search device 1000 further includes a group ciphertext receiving unit configured to receive ciphertexts corresponding to each group of object vectors from a data source.

[0143] In one embodiment, the server is implemented as a private server of the data party, in which the plain texts corresponding to the m groups of object vectors are stored.

[0144] In one embodiment, the first ciphertext includes an inner product ciphertext, and the plaintext corresponding to the inner product ciphertext includes the inner product between the query vector and each object vector in the p groups.

[0145] Further, in a specific embodiment, each group of object vectors includes n object vectors, and the corresponding plaintext is the first rearranged vector of the group of object vectors, and the first rearranged vector includes d groups of elements, wherein the jth element in any i-th group is the i-th dimension element of the j-th object vector in the group of object vectors; or, the ciphertext corresponding to each group of object vectors is obtained by encrypting the first rearranged vector using the public key of the user. The ciphertext corresponding to the query vector is obtained by encrypting the second rearranged vector using the public key, and the second rearranged vector includes d groups of elements, wherein the elements in any i-th group are the i-th dimension elements of the query vector.

[0146] The first ciphertext determination unit 1020 is specifically configured to: for any group in the p groups, perform a homomorphic multiplication operation on the plaintext or ciphertext of the first rearranged vector of the group and the ciphertext of the second rearranged vector to obtain a current vector ciphertext; perform a logarithmic multiplication operation on the current vector ciphertext 2 d iterations, wherein any sth iteration includes: performing a step of n*d / 2 on the current vector ciphertext scyclic rotation; updating the current vector ciphertext to the result of homomorphic addition operation between the current vector ciphertext and the result of rotation of the vector; determining the inner product ciphertext based on the iterated current vector ciphertext, wherein the first n elements of the plaintext corresponding to the current vector ciphertext are the inner products between the query vector and each object vector in the corresponding group.

[0147] Furthermore, in one example, the first ciphertext determining unit 1020 is configured to determine the inner product ciphertext based on the iterated current vector ciphertext, specifically including:

[0148] The mask vector is used to perform a homomorphic multiplication operation on the current vector ciphertext to obtain a masked result ciphertext; the first n bits of the mask vector are predetermined non-zero values, and the other bits are zero. The p masked result ciphertexts corresponding to the p groups are merged, specifically including: for the i-th masked result ciphertext, a cyclic rotation with a step length of (i-1)*n bits is performed; the p rotation results corresponding to the p masked result ciphertexts are homomorphically added to obtain the inner product ciphertext.

[0149] It should be noted that for the introduction of the above functional modules, reference can also be made to the relevant introduction of the process method in the above embodiments.

[0150] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0151] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A vector retrieval method for protecting data privacy, wherein a client stores the centroid vector of each group of object vectors; The method is applied to the client, comprising: According to the query input of the user, the similarity between the corresponding query vector and the centroid vector is calculated to obtain the target group identifier corresponding to the highest similarity; Sending the ciphertext corresponding to the query vector and p group identifiers to the server; the p group identifiers include the target group identifier; receiving a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; the first ciphertext is determined by the server by performing a homomorphic operation based on the locally stored vector and the received vector; According to the decryption result of the first ciphertext, the similarity between the query vector and each object vector in the target group is determined to obtain a query result corresponding to the query input.

2. The method according to claim 1, wherein: Each group of object vectors belongs to m groups of object vectors; wherein the m groups of object vectors and the centroid vector are determined by the data party based on the following steps: Performing clustering processing based on the object vector set to obtain m clusters and the centroid vector of each cluster, wherein the number of vectors of each cluster is less than or equal to a preset value n; The m clusters are used to initialize m groups of object vectors. For each group, a nearest neighbor search algorithm is used to determine a batch of vectors whose similarity with the centroid vector of the group is within a preset range from the object vector set, and the group is filled based on the batch of vectors until the number of vectors in the group reaches n and there are no duplicate vectors.

3. The method according to claim 2, wherein: Based on the object vector set, clustering is performed to obtain m clusters, including: Performing a first clustering process on the object vector set to obtain a plurality of first clusters; For each first cluster, when it is determined that the number of vectors therein is greater than n, a second clustering process is performed on it, and the obtained multiple second clusters are classified into the m clusters; otherwise, the first cluster is directly classified into the m clusters.

4. The method according to claim 1, wherein: Each group of object vectors belongs to m groups of object vectors; wherein the determination of the p group identifiers includes: Based on the target partition identifier, p-1 other partition identifiers are randomly selected from the m group identifiers and are collectively included in the p group identifiers.

5. The method according to claim 1, wherein: The server is implemented as a third-party server, which stores the ciphertext corresponding to each group of object vectors, and the ciphertext is provided by the data party.

6. The method according to claim 5, further comprising: Receive the ciphertext of the centroid vector from the third-party server, which is obtained by the data party encrypting the centroid vector using a symmetric key; The symmetric key is used to decrypt the ciphertext of the centroid vector, and the decrypted plaintext is stored.

7. The method according to claim 1, wherein: The server is implemented as a private server of the data party, in which the plain text corresponding to each group of object vectors is stored.

8. The method according to claim 7, further comprising: The centroid vector is received from the private server and stored.

9. The method according to claim 1, wherein: The first ciphertext includes an inner product ciphertext, and the corresponding plaintext includes the inner product between the query vector and each object vector in the p groups.

10. The method according to claim 1 or 9, wherein: Determining the similarity between the query vector and each object vector in the target group according to the decryption result of the first ciphertext includes: For each object vector in the target group, the inner product between the object vector and the query vector is extracted from the decryption result as the corresponding similarity.

11. The method according to claim 1 or 9, wherein: The client further stores the norm of each object vector in each group of object vectors; wherein, according to the decryption result of the first ciphertext, determining the similarity between the query vector and each object vector in the target group includes: For each object vector in the target group, an inner product between the object vector and the query vector is extracted from the decryption result, and a similarity between the object vector and the query vector is determined based on the inner product, a norm of the object vector and a norm of the query vector.

12. The method according to claim 1, wherein: After determining the similarity between the query vector and each object vector in the target group, the method further includes: The k object IDs corresponding to the top k similarities are determined as the query result.

13. A vector retrieval method for protecting data privacy, wherein the client stores the centroid vector of each group of object vectors; The method is applied to the server, and includes: Receiving, from the client, a ciphertext corresponding to the query vector and p group identifiers; the p group identifiers include a target group identifier; the centroid vector of the target group has a higher similarity to the query vector than other groups, the query vector being determined based on a query input by a user; Performing homomorphic operations based on the locally stored vector and the received vector to obtain a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; The first ciphertext is sent to the client, so that the client determines the similarity between the query vector and each object vector in the target group according to the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

14. The method according to claim 13, wherein: The server is implemented as a third-party server; the method further includes: The ciphertext corresponding to each group of object vectors is received from the data party.

15. The method according to claim 13, wherein: The server is implemented as a private server of the data party, in which the plain text corresponding to each group of object vectors is stored.

16. The method according to claim 13, wherein: The first ciphertext includes an inner product ciphertext, and the corresponding plaintext includes the inner product between the query vector and each object vector in the p groups.

17. The method according to claim 16, wherein: The server stores the plaintext or ciphertext corresponding to each group of object vectors; each group of object vectors includes n object vectors, and the corresponding plaintext is the first rearranged vector of the group of object vectors, and the first rearranged vector includes d groups of elements, wherein the jth element in any i-th group is the i-th dimension element of the j-th object vector in the group of object vectors; or, the ciphertext corresponding to each group of object vectors is obtained by encrypting the first rearranged vector using the public key of the user; The ciphertext corresponding to the query vector is obtained by encrypting a second rearrangement vector using the public key, the second rearrangement vector includes d groups of elements, and any element of the i-th group is an i-th dimension element of the query vector; The homomorphic operation includes: for any group in the p groups, performing a homomorphic multiplication operation on the plaintext or ciphertext of the first rearranged vector of the group and the ciphertext of the second rearranged vector to obtain the current vector ciphertext; performing log2d iterations on the current vector ciphertext, wherein any sth iteration includes: performing a step length of n*d / 2 on the current vector ciphertext s cyclic rotation; updating the current vector ciphertext to the result of homomorphic addition operation between the current vector ciphertext and the result of rotation of the vector; determining the inner product ciphertext based on the iterated current vector ciphertext, wherein the first n elements of the plaintext corresponding to the current vector ciphertext are the inner products between the query vector and each object vector in the corresponding group.

18. The method according to claim 17, wherein: Determining the inner product ciphertext based on the iterated current vector ciphertext includes: Performing a homomorphic multiplication operation on the current vector ciphertext using a mask vector to obtain a masked result ciphertext; the first n bits of the mask vector are predetermined non-zero values, and the other bits are zero; The p mask result ciphertexts corresponding to the p groups are merged, specifically comprising: for the i-th mask result ciphertext, performing a cyclic rotation with a step length of (i-1)*n bits; and performing a homomorphic addition operation on the p rotation results corresponding to the p mask result ciphertexts to obtain the inner product ciphertext.

19. A vector retrieval device for protecting data privacy, integrated in a client, wherein the client stores the centroid vector of each group of object vectors; the device comprises: A target group determination unit is configured to calculate the similarity between the query vector corresponding to the user's query input and the centroid vector, and obtain the target group identifier corresponding to the highest similarity; A query data sending unit, configured to send the ciphertext corresponding to the query vector and p group identifiers to a server; the p group identifiers include the target group identifier; an inner product ciphertext receiving unit, configured to receive a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; the first ciphertext is determined by the server by performing a homomorphic operation based on the locally stored vector and the received vector; The query result determination unit is configured to determine the similarity between the query vector and each object vector in the target group according to the decryption result of the first ciphertext, so as to obtain the query result corresponding to the query input.

20. A vector retrieval device for protecting data privacy, integrated in a server, wherein a client corresponding to the server stores the centroid vector of each group of object vectors; the device comprises: a query data receiving unit configured to receive the ciphertext corresponding to the query vector and p group identifiers from the client; The p group identifiers include a target group identifier; The centroid vector of the target group has a higher similarity to the query vector than the other groups, wherein the query vector is determined based on the query input of the user; A first ciphertext determination unit is configured to perform a homomorphic operation based on the locally stored vector and the received vector to obtain a first ciphertext, the corresponding plaintext of which is used to measure the similarity between the query vector and each object vector in the p groups; The first ciphertext sending unit is configured to send the first ciphertext to the client, so that the client determines the similarity between the query vector and each object vector in the target group according to the decryption result of the inner product ciphertext, so as to obtain the query result corresponding to the query input.

21. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 18.

22. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 18 is implemented.

Citation Information

Patent Citations

  • Semantic-based multi-keyword sorting search privacy protection system and method

    CN108647529A

  • Data hiding query method and device, electronic equipment and storage medium

    CN118551122A

  • Semantic perception cross-modal encryption retrieval method

    CN119760188A

  • Method for fully homomorphic encryption using multivariate cryptography

    US20130329883A1

  • Metric Recommendations in an Event Log Analytics Environment

    US20160335260A1