A method and system for privacy information retrieval based on a CSD array
By integrating Flash storage and FPGA computing units into the CSD array, and performing homomorphic addition and multiplication operations in parallel, the I/O bottleneck problem in PIR technology is solved, enabling efficient privacy information retrieval and improving the system's concurrent processing capabilities.
Patent Information
- Application Number
- CN202511094985.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing PIR technologies suffer from I/O bottlenecks, especially in single-server PIR models, where the I/O burden caused by fully homomorphic encryption is difficult to meet the requirements, thus limiting the practical application of PIR technologies.
The database is sharded and stored in multiple CSDs of the CSD array. Each CSD integrates a Flash storage unit and an FPGA computing unit. Homomorphic addition and multiplication operations are performed in parallel using the FPGA. Data transmission is optimized through a customized I/O scheduling algorithm and PCIe channel, enabling the computing tasks to be pushed down to the storage device.
It significantly reduced I/O overhead, improved computing efficiency, reduced the latency of frequent data transfer to the host, achieved efficient privacy information retrieval, and reduced query latency by two orders of magnitude.
Smart Images

Figure CN120597303B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer storage and encryption technology, for example, to a privacy information retrieval method and system based on a CSD array. Background Art
[0002] With the deepening of digital transformation, cloud computing has become core infrastructure for privacy-critical industries such as healthcare and finance. In these industries, protecting the privacy of user queries is crucial, and cloud service providers should not be able to access user query information. To address this, Private Information Retrieval (PIR) technology has emerged. Its core concept is to allow users to retrieve specific information from a database while ensuring that the database server remains unaware of the information being retrieved, thereby protecting user privacy.
[0003] Early PIR models, which relied on multiple servers, were limited in practical application due to high deployment and maintenance costs, low communication efficiency, and difficulty ensuring server non-collusion. With advances in fully homomorphic encryption (FHE), single-server PIR models have emerged. These models, based on FHE, achieve privacy protection by performing a dot product of an encrypted query vector and a database vector. However, this approach suffers from significant performance bottlenecks, particularly in I / O (Input / Output). Because the PIR protocol requires scanning the entire database, it results in a large number of random read operations, which even the 4K random read bandwidth of high-performance NVMe SSDs cannot meet. Homomorphic encryption operations disrupt data spatial locality and prevent data reuse, further exacerbating the I / O burden. This I / O bottleneck has become a major obstacle to the practical application of PIR technology. Therefore, a private information retrieval method is urgently needed to address these issues.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0006] The embodiments of the present disclosure provide a privacy information retrieval method and system based on a CSD array to solve the technical problem of I / O bottleneck in PIR technology.
[0007] In some embodiments, database shards are stored in multiple CSDs in a CSD array, each CSD internally integrating a Flash storage unit and an FPGA computing unit. The method includes the following steps:
[0008] Receive an encrypted query vector containing several elements sent by the client, where the target element requested by the user in the encrypted query vector is 1 and the remaining elements are 0;
[0009] Execute the expansion algorithm on the CPU to expand the encrypted query vector into a number of ciphertexts with the same number of elements. The ciphertext of the target element is the encryption of 1, and the ciphertext of the remaining elements is the encryption of 0.
[0010] The expanded ciphertext and plaintext databases are sent to the CSD array. Each CSD uses the internal FPGA to perform homomorphic addition and multiplication operations on the ciphertext and the corresponding shard data in parallel, and the encryption results are summarized.
[0011] The encryption result is fed back to the client so that the client can decrypt and obtain the target data.
[0012] In some embodiments, in each CSD performing homomorphic addition and multiplication operations of the ciphertext and the corresponding shard data in parallel through an internal FPGA, the method includes:
[0013] Based on the SIMD characteristics within the vector, homomorphic addition and multiplication are decomposed into modular addition and modular multiplication operations. The calculation formula for modular addition is as follows:
[0014] ;
[0015] The calculation formula for modular multiplication is as follows:
[0016] ;
[0017] Where, and Represent ciphertext and plaintext respectively, and are the moduli in homomorphic encryption.
[0018] In some embodiments, when performing homomorphic addition and multiplication operations on an FPGA, the method further includes performing pipeline optimization on the homomorphic addition and multiplication operations. The pipeline optimization process includes:
[0019] The computation process is broken down into multiple pipeline stages, each of which processes different data in parallel. Each stage includes loading data, performing calculations, and storing results.
[0020] In some embodiments, the pipeline optimization process further includes:
[0021] A multi-stage pipeline structure is adopted to control each stage to complete operations within an independent clock cycle, so as to achieve parallel execution of homomorphic encryption operations.
[0022] In some embodiments, when performing homomorphic addition and multiplication operations on the FPGA, the method further includes performing unroll optimization on the homomorphic addition and multiplication operations. The unroll optimization process includes:
[0023] The loops in homomorphic addition and multiplication operations are disassembled and assigned to independent FPGA computing units for parallel execution, so that CSD can simultaneously process homomorphic addition and multiplication operations on different data within multiple clock cycles.
[0024] In some embodiments, the method further comprises:
[0025] Formulate a task allocation strategy based on the number of non-zero plaintext entries in the database;
[0026] According to the distribution strategy, entries with a large number of non-zero plaintexts are evenly distributed to each CSD.
[0027] In some embodiments, the Flash storage unit uses a customized I / O scheduling algorithm to dynamically adjust data sharding and task allocation strategies.
[0028] In some embodiments, when each CSD performs homomorphic addition and multiplication operations of ciphertext and corresponding sharded data in parallel through an internal FPGA, modular addition and modular multiplication operations are implemented using hardware circuits, which include comparators, multipliers, adders, subtractors, and multiplexers.
[0029] In some embodiments, each CSD has an independent PCIe channel for direct communication between the Flash storage unit and the FPGA computing unit.
[0030] In some embodiments, a CSD array-based private information retrieval system is used to execute any of the above-mentioned CSD array-based private information retrieval methods.
[0031] The embodiments of the present disclosure provide a method and system for private information retrieval based on a CSD array, which can achieve the following technical effects:
[0032] By storing database shards in multiple CSDs with integrated Flash storage and FPGA computing units, the multiple CSDs can process different data shards in parallel. After receiving the client's encrypted query vector and executing the expansion algorithm on the processor, the ciphertext and plaintext database are sent to the CSD array, where homomorphic addition and multiplication operations are performed in parallel using the FPGA within each CSD. During this process, the expansion algorithm is executed by the processor, leveraging its large memory to replicate the ciphertext and expand the original encrypted query vector into a corresponding number of ciphertexts, so that each database entry corresponds to a ciphertext. Furthermore, the computation is pushed down to the storage device for execution, allowing the computation to be completed within the storage device itself. This avoids the I / O bottleneck of frequent data transfers to the host in traditional architectures. Parallel processing by multiple CSDs also breaks the limitations of centralized computing and improves the efficiency of homomorphic computing. Finally, the server returns the encrypted results aggregated from the CSD array to the client, which decrypts the target data using its private key. Throughout this process, the server remains completely inaccessible to the plaintext, allowing users to retrieve information from the server without exposing their query content. As described above, this application reduces the I / O overhead in PIR technology and improves computing efficiency while ensuring privacy.
[0033] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,
[0035] Figure 1 is a flowchart of a privacy information retrieval method based on a CSD array provided by an embodiment of the present disclosure;
[0036] Figure 2 is a schematic diagram of a traditional computing architecture provided by an embodiment of the present disclosure;
[0037] Figure 3 is a schematic diagram of a CSD-enabled near data processing architecture provided by an embodiment of the present disclosure;
[0038] Figure 4 is a hardware circuit diagram for implementing modular multiplication operations provided by an embodiment of the present disclosure;
[0039] Figure 5 is a hardware circuit diagram for implementing analog-addition operation provided by an embodiment of the present disclosure;
[0040] Figure 6 This is a deployment diagram of a CSD array-based privacy information retrieval system provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0041] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0042] The terms "first," "second," and the like in the embodiments of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to facilitate the description of the embodiments of the present disclosure herein. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.
[0043] Unless otherwise stated, the term "plurality" means two or more.
[0044] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0045] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0046] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0047] The core principle of PIR is "all-for-one," which means treating every entry in the database uniformly to prevent the client's actual query interests from being leaked. This principle ensures absolute protection of user privacy, but it also brings significant technical challenges. In early PIR models that relied on multiple servers, users split query requests into multiple sub-queries and sent them to different servers. The target data was obtained by XORing the response results of each server. However, due to high deployment and maintenance costs, low communication efficiency, and the difficulty in ensuring that there is no collusion between servers, the practical application of PIR models that rely on multiple servers is limited. For the single-server PIR model, its PIR scheme based on fully homomorphic encryption performs well in terms of privacy protection, but the I / O bottleneck has become a major obstacle to the practical application of PIR.
[0048] Current optimization solutions have significant limitations in improving I / O performance. For example, traditional storage optimization techniques (such as data prefetching and cache optimization) struggle to cope with the full-database scan nature of PIR, while software-based I / O scheduling strategies fail to fully leverage the parallel capabilities of modern storage hardware. Furthermore, while current hardware acceleration solutions (such as GPUs and FPGAs) can improve computing efficiency, their impact on I / O performance is limited. This is especially true in high-concurrency query scenarios, where I / O bottlenecks still dominate system performance.
[0049] To address the aforementioned I / O bottleneck, embodiments of the present disclosure provide a privacy information retrieval method based on a CSD array. This privacy information retrieval method pushes the PIR computational tasks down to a composite storage device (CSD) array. The CSD array can integrate different types of storage media (such as SSDs, HDDs, and tapes) through a layered architecture. For example, a SmartSSD array can fully utilize its parallel computing capabilities and internal high-bandwidth characteristics, avoiding the data movement overhead of traditional architectures and significantly reducing I / O costs.
[0050] In some embodiments, database shards can be stored across multiple CSDs in a CSD array. Each CSD integrates Flash storage and FPGA computing units, allowing each SSD to independently handle homomorphic operations on local data, thus avoiding the centralized I / O bottlenecks found in traditional architectures. Furthermore, each CSD features an independent PCIe channel for direct communication between the Flash storage and FPGA computing units. This leverages PCIe's high-speed bandwidth to increase data transfer speeds, reduce communication latency through direct transmission paths, and avoid resource contention thanks to independent channels.
[0051] In some embodiments, the Flash storage unit utilizes a customized I / O scheduling algorithm that dynamically partitions data based on the number of non-zero plaintext entries in the database, evenly distributing entries with high computational complexity to different CSDs. The algorithm also monitors load metrics such as FPGA utilization and I / O queue length across each device in real time, automatically migrating data shards to lower-load nodes when a device's load exceeds the specified threshold. Regarding task allocation, the dominant inner product operation in the SealPIR protocol is pushed down to the CSD's internal FPGA for execution, utilizing a PCIe switch for point-to-point data transmission between CSDs. Simultaneously, the Flash storage unit and the FPGA computing unit communicate directly via internal PCIe channels, balancing the computational load across multiple devices while reducing data transmission to the host, significantly improving the system's concurrent processing efficiency and I / O performance.
[0052] The following describes in detail the privacy information retrieval method based on the CSD array provided by the embodiments of the present disclosure with reference to the accompanying drawings.
[0053] Figure 1 This is a flowchart of a privacy information retrieval method based on a CSD array provided by an embodiment of the present disclosure, combined with Figure 1 As shown, the privacy information retrieval method includes the following steps:
[0054] S101: Receive an encrypted query vector containing several elements sent by a client.
[0055] In some embodiments, when a user has a search requirement, an encrypted query vector can be generated by the client. The client constructs an encrypted query vector containing several elements, wherein the target element requested by the user in the encrypted query vector is 1 and the remaining elements are 0. For example, a server maintains a database D, wherein , the user wants to request the kth entry, then the encrypted query vector constructed by the client can contain N elements, where the kth element is 1 and the other elements are 0. The client sends the generated encrypted query vector to the server.
[0056] S102: Execute an expansion algorithm on the CPU to expand the encrypted query vector into a number of ciphertexts having the same number of elements.
[0057] In some embodiments, an expansion algorithm (Expand) is executed by a processor, utilizing its large memory to replicate ciphertexts and expand the original encrypted query vector into a corresponding number of ciphertexts, such that each database entry corresponds to a ciphertext. The ciphertext of the target element is an encryption of 1, and the ciphertexts of the remaining elements are encryptions of 0. Continuing with the example in S101, the server executes the expansion algorithm on a central processing unit (CPU) to expand the encrypted query vector into N ciphertexts, where the kth ciphertext is Enc (1) and the others are Enc (0).
[0058] S103: Send the expanded ciphertext and plaintext database to the CSD array. Each CSD performs homomorphic addition and multiplication operations on the ciphertext and the corresponding shard data in parallel through the internal FPGA, and summarizes the encryption results.
[0059] Figure 2 is a schematic diagram of a traditional computing architecture provided by an embodiment of the present disclosure. Figure 3 This is a schematic diagram of a CSD-enabled near data processing architecture provided by an embodiment of the present disclosure. Figure 2 and Figure 3 This application introduces the process of efficiently implementing the PIR protocol deployment based on the Near Data Processing (NDP) architecture.
[0060] Figure 2 This diagram illustrates a von Neumann architecture centered around the CPU and memory. When the CPU needs to process data stored on an SSD (Solid State Drive), the data must first be read from the SSD (Flash), passed through the SSD controller (Controller), and then transferred via the PCIe bus to DRAM (Dynamic Random Access Memory). The CPU then reads the data from DRAM for processing. To restore the processed data, it must be written back from DRAM to the SSD controller via the PCIe bus, ultimately storing it in Flash. Similarly, when a GPU (Graphics Processing Unit) or FPGA (Field-Programmable Gate Array) processes stored data, the data must also flow between the SSD, DRAM, and the GPU or FPGA via the PCIe bus. This strict separation of compute and storage means that all data processing relies on the CPU accessing data in storage via the PCIe bus, leading to significant I / O bottlenecks, latency, and energy inefficiency.
[0061] See also Figure 3 On the left, DRAM acts as memory, temporarily storing data, interacting with the CPU through data paths, and providing data support for CPU operations. The CPU is the computing core, connected to other components through the PCIe bus, coordinating and scheduling data processing tasks. GPU, NIC (network interface controller), FPGA, and CSD all have data exchanges with the CPU via the PCIe bus. Each CSD is connected to a CSD controller (Ctlr), which is responsible for managing Flash. Data can be read and written between the CSD internal controller and the flash memory, and can also be exchanged with other components such as the CPU via the PCIe bus to achieve near-data processing. In addition, each CSD in the CSD array uses the internal FPGA to execute the product of the ciphertext and plaintext databases in parallel, minimizing the amount of data flowing from the storage layer to the CPU end, thereby avoiding the centralized I / O bottleneck in the traditional architecture and reducing "Light competition". Look again Figure 3On the right, a CSD array, taking the SmartSSD Array as an example, is shown. The host is the starting and ending point for data exchange within the entire SmartSSD system, initiating operations such as SSDRead / Write. The PCIe Switch forwards data between the host, the SSDController, and the FPGA (KU15P). It supports SSD read / write data transmission between the host and the SSD controller, as well as data exchange between the host and the FPGA and DRAM (memory on the FPGA) during FPGA & DRAM Read / Write operations. It also supports peer-to-peer (P2P) data transmission, enabling direct data exchange between the SSD controller and the FPGA.
[0062] As described above, the computing tasks in this application are pushed down to the storage device side (CSD array) for execution, so that the calculations are completed inside the storage device, avoiding the I / O bottleneck of frequent data transfer to the host in traditional architectures. At the same time, the parallel processing of multiple CSDs breaks the limitations of centralized computing and improves the efficiency of homomorphic computing.
[0063] In some embodiments, in each CSD performing homomorphic addition and multiplication operations of ciphertext and corresponding sliced data in parallel through an internal FPGA, the method includes: decomposing the homomorphic addition and multiplication into modular addition and modular multiplication operations based on SIMD characteristics within the vector, wherein the calculation formula of the modular addition is as follows:
[0064] ;
[0065] The calculation formula of the modular multiplication is as follows:
[0066] ;
[0067] Where, and Represent ciphertext and plaintext respectively, and are the moduli in homomorphic encryption.
[0068] SIMD (Single Instruction Multiple Data) is a parallel computing technology that allows a single instruction to operate on multiple data elements simultaneously. Using SIMD in homomorphic encryption allows for simultaneous homomorphic addition and multiplication of multiple data elements, improving computational efficiency.
[0069] In some embodiments, in the homomorphic addition and multiplication operations of the ciphertext and the corresponding sharded data performed in parallel by each CSD through the internal FPGA, the modular addition and modular multiplication operations are implemented using a hardware circuit, and the hardware circuit includes a comparator, a multiplier, a subtractor and a multiplexer.
[0070] In the modular multiplication operation, first calculate , that is, the ciphertext and plain text Then, by integer division, we can calculate , and finally through This process ensures that the result of the multiplication operation on the ciphertext is consistent with the result of the multiplication operation on the plaintext after decryption, thus achieving homomorphic multiplication. Figure 4 This is a hardware circuit diagram for implementing modular multiplication operations provided by an embodiment of the present disclosure. Figure 4 In the example, op1 / op2 are inputs, (×) is the multiplier, Upper half / Lower half is the high / low bit split, (-) is the subtractor, CMP is the comparator, and MUX is the multiplexer. First, op1 and op2 enter the multiplier first, and we get Then, in some paths, the product will be split into the upper half (high bit) and the lower half (low bit), and combined with other inputs (such as u, p) to perform multiplication and subtraction to assist in implementing the logic of integer division and correction. The intermediate results are compared through multiple groups of CMP to determine whether a subtractor is needed for correction. Finally, the correct result is selected through MUX and output .
[0071] Combine Figure 4 To illustrate the modular multiplication process, assume op1 = 4, op2 = 5, modulus p = 3, and auxiliary parameter u = 1. First, perform a preliminary multiplication to calculate the product of op1 and op2, 4×5 = 20; calculate the product of u and op2, 1×5 = 5; and calculate the product of u and p, 1×3 = 3. Then, subtract the product of u and p from the product of op1 and op2, 20 - 3 = 17. Since 17 ≥ 3 holds true, subtract 3 again, obtaining 17 - 3 = 14. Then, determine 14 ≥ 3 and subtract 3 again, obtaining 14 - 3 = 11. Repeat this process. After multiple subtractions, the final result is 2, where 2 < 3. Finally, after the above adjustments, the final modular multiplication result is 2, that is, (4×5) mod 3 = 2.
[0072] In the modular addition operation, the result of adding two numbers under modular arithmetic is defined. The value is less than the modulus When itself; when Greater than or equal to When minus In homomorphic encryption, the ciphertext and plaintext are added together to maintain homomorphism through this modular addition operation, so that the result of the addition operation on the ciphertext is consistent with the result of the addition operation on the plaintext after decryption. Figure 5 This is a hardware circuit diagram for implementing analog-addition operation provided by an embodiment of the present disclosure. Figure 5 In the example, op1 / op2 are inputs, (+) is the adder, (-) is the subtractor, CMP is the comparator, and MUX is the multiplexer. First, op1 and op2 enter the adder and get the sum. . Then, and and the modulus p (that is, the modulus addition formula ) Input CMP to compare the magnitudes. Finally, if the sum is less than the modulus, the MUX selects the adder output and directly outputs the sum. If the sum is greater than or equal to the modulus, the MUX selects the subtractor output, which subtracts the modulus from the sum.
[0073] Combine Figure 5 Let's use an example to illustrate the modular addition process. Assume op1 = 7, op2 = 5, and modulus p = 3. First, add the two operands, i.e., 7 + 5 = 12. Then, determine the relationship between the result and the modulus and adjust it. Divide the result of the addition by the modulus p to see the remainder. Subtract the modulus p from the addition result, obtaining 12 - 3 = 9. Then, determine if 9 ≥ 0 (because the modular addition result must be within the range [0, 3)). Since the result is greater than or equal to 0, subtract the modulus p again, 9 - 3 = 6. Then, determine if 6 ≥ 0. Continue the operation, 6 - 3 = 3, determine if 3 ≥ 0, and repeat the operation again, 3 - 3 = 0. Now, the determination 0 ≥ 0 holds. Finally, after the above adjustments, the final modular addition result is 0, i.e., (7 + 5) mod 3 = 0.
[0074] In some embodiments, since multiple homomorphic addition and multiplication operations can be performed simultaneously in different stages during the calculation process of the SealPIR protocol, thereby realizing pipeline processing, the method also includes Pipeline optimization of homomorphic addition and multiplication operations in the FPGA during homomorphic addition and multiplication operations. The Pipeline optimization process includes: decomposing the calculation process into multiple pipeline stages, each stage processing different data in parallel, wherein each stage includes loading data, performing calculations, and storing results. For example, in SmartSSD, each FPGA computing unit can process multiple computing tasks simultaneously by designing multiple pipeline stages. Modular multiplication operations can be divided into loading data, performing multiplication calculations, storing results, etc. These stages can work in parallel, and each stage can process different data, thereby improving computing efficiency.
[0075] In some embodiments, the pipeline optimization process also includes: using a multi-stage pipeline structure to control each stage to complete operations within independent clock cycles, so as to achieve parallel execution of homomorphic encryption operations. For example, the FPGA computing unit of the SmartSSD can be configured as a multi-stage pipeline structure, so that the operations of each stage can be completed within an independent clock cycle. By breaking down complex homomorphic encryption operations into multi-stage pipeline stages, the computing process can achieve more efficient parallel execution.
[0076] In some embodiments, when performing homomorphic addition and multiplication operations on an FPGA, the method further includes performing Unroll optimization on the homomorphic addition and multiplication operations. The Unroll optimization process includes: disassembling the loops in the homomorphic addition and multiplication operations and assigning them to independent FPGA computing units for parallel execution, so that the CSD can simultaneously process homomorphic addition and multiplication operations on different data within multiple clock cycles. For example, when calculating the inner product, each entry in the database performs multiplication and addition operations with the encrypted query vector, and these operations can be parallelized by unrolling the loop. The calculation of each database entry can be assigned to an independent FPGA computing unit, thereby achieving more efficient parallel processing. By unrolling the loop, the FPGA computing unit can execute each homomorphic addition and multiplication operation in parallel, and different computing units process different data, so that the computing task can be executed simultaneously within multiple clock cycles.
[0077] The above optimization process can significantly improve the processing speed of encrypted data in the PIR protocol, reduce the time required for calculation, and thus speed up the query response.
[0078] In some embodiments, if some CSDs in a CSD array process large amounts of data while other CSDs handle lighter workloads, this can lead to performance bottlenecks and resource waste. Therefore, the method further includes formulating a task allocation strategy based on the number of non-zero plaintext entries in the database, and evenly distributing entries with a high number of non-zero plaintext entries across the CSDs according to the allocation strategy. Generally, data entries with a high number of non-zero plaintext entries increase computational complexity. Therefore, by evenly distributing these data entries across multiple CSDs, the computational load of each CSD can be controlled within a reasonable range, thereby achieving load balancing optimization.
[0079] S104: Feedback the encryption result to the client so that the client can decrypt and obtain the target data.
[0080] In some embodiments, after obtaining the summarized encryption results, the server feeds them back to the client. The client receives the encryption results and performs a decryption operation to finally obtain the required retrieved items.
[0081] The deployment process of privacy information retrieval in this application is described below with reference to the accompanying drawings.
[0082] Figure 6 This is a deployment diagram of a privacy information retrieval system based on a CSD array provided by an embodiment of the present disclosure. Figure 6 Taking the SmartSSD Array as an example, on the host side, a user initiates a query through the client. The query is first processed by the SealPIR::Expand() algorithm, which converts and expands the encrypted query vector into ciphertext. These ciphertexts are then stored as query data in the host-side QueriesBuffer. This buffer utilizes the fast read and write speeds of DRAM to temporarily store the query data for later transmission. The query data in the host-side query buffer is then transferred to the SmartSSD via the PCIe bus. Meanwhile, the SmartSSD's SSD stores the entire Database D, which is subsequently accessed as needed for computation. Once in the SmartSSD, the query data is first stored in the QueriesBuffer, while the relevant data in Database D is read into the Data Buffer. Both buffers are based on the FPGA's dynamic random access memory (DRAM), providing local temporary data storage for the FPGA computing module and reducing data transfer latency. The internal PCIe bus is responsible for data transmission between SmartSSD components (such as the SSD and FPGA DRAM), ensuring efficient data flow. Query data and database data stored in the FPGA DRAM are then fed into the on-chip storage / resources within the FPGA computing unit for computational operations. Figure 6 The figure demonstrates multi-instance parallel computing, with homomorphic addition (Add) and multiplication (PMult) operations performed within each instance. Multi-instance parallel computing fully utilizes the hardware parallelism of the FPGA compute unit, improving computational efficiency and accelerating the generation of query results. Finally, after the FPGA compute unit completes the computation, it aggregates the resulting Ans (results) and returns them to the host for the user to retrieve.
[0083] The privacy information retrieval method based on the CSD array in this application achieves the coordinated optimization of storage and computing by pushing the PIR calculation task down to the storage device and utilizing the parallel computing capability of the FPGA computing unit inside the CSD and a customized I / O scheduling algorithm. It reduces the query latency by 2 orders of magnitude at the scale of a million-level database, significantly reduces the I / O overhead, and improves the system's concurrent processing capability, providing an efficient and scalable solution for the large-scale application of PIR technology.
[0084] Based on the same inventive concept as the above-mentioned CSD array-based privacy information retrieval method, this application also discloses, in some embodiments, a CSD array-based privacy information retrieval system. This system is used to execute the CSD array-based privacy information retrieval method disclosed in any of the above-mentioned embodiments, and its process is not further described here.
[0085] The technical solutions of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, and other media that can store program code, or a transient storage medium.
[0086] The above description and the accompanying drawings sufficiently illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless expressly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only used to describe the embodiments and are not used to limit the scope of protection. As used in the description herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include the plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be referred to the description of the method part.
[0087] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0088] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices and equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units may be merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or omitting or disabling some features. In addition, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interface, or the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to implement the present embodiments according to actual needs. In addition, the functional units in the embodiments of the present disclosure may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
Claims
1. A privacy information retrieval method based on a CSD array, characterized in that: The database shards are stored in multiple CSDs of a CSD array, each CSD internally integrating a Flash storage unit and an FPGA computing unit. The method includes the following steps: Receive an encrypted query vector containing several elements sent by a client, wherein the target element requested by the user in the encrypted query vector is 1 and the remaining elements are 0; Executing an expansion algorithm on the CPU to expand the encrypted query vector into a number of ciphertexts equal to the number of elements, wherein the ciphertext of the target element is the encryption of 1, and the ciphertext of the remaining elements is the encryption of 0; The expanded ciphertext and plaintext databases are sent to the CSD array. Each CSD performs homomorphic addition and multiplication operations on the ciphertext and the corresponding shard data in parallel through the internal FPGA, and summarizes the encryption results. The encryption result is fed back to the client so that the client can decrypt and obtain the target data.
2. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: In each CSD executing homomorphic addition and multiplication operations of ciphertext and corresponding shard data in parallel through an internal FPGA, the method includes: Based on the SIMD characteristics within the vector, homomorphic addition and multiplication are decomposed into modular addition and modular multiplication operations, where the calculation formula of the modular addition is as follows: ; The calculation formula of the modular multiplication is as follows: ; Where, and Represent ciphertext and plaintext respectively, and are the moduli in homomorphic encryption.
3. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: In performing homomorphic addition and multiplication operations on the FPGA, the method further includes performing pipeline optimization on the homomorphic addition and multiplication operations. The pipeline optimization process includes: The computation process is broken down into multiple pipeline stages, each of which processes different data in parallel. Each stage includes loading data, performing calculations, and storing results.
4. The privacy information retrieval method based on the CSD array according to claim 3, characterized in that: The Pipeline optimization process also includes: A multi-stage pipeline structure is adopted to control each stage to complete operations within an independent clock cycle, so as to achieve parallel execution of homomorphic encryption operations.
5. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: In performing homomorphic addition and multiplication operations on the FPGA, the method further includes performing unroll optimization on the homomorphic addition and multiplication operations. The unroll optimization process includes: The loops in homomorphic addition and multiplication operations are disassembled and assigned to independent FPGA computing units for parallel execution, so that CSD can simultaneously process homomorphic addition and multiplication operations on different data within multiple clock cycles.
6. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: The method further comprises: Formulate a task allocation strategy based on the number of non-zero plaintext entries in the database; According to the allocation strategy, entries with a large number of non-zero plaintexts are evenly distributed to each CSD.
7. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: The Flash storage unit adopts a customized I / O scheduling algorithm to dynamically adjust data sharding and task allocation strategies.
8. The privacy information retrieval method based on the CSD array according to claim 2, characterized in that: In the homomorphic addition and multiplication operations of the ciphertext and the corresponding sharded data performed in parallel by each CSD through the internal FPGA, the modular addition and modular multiplication operations are implemented using a hardware circuit, and the hardware circuit includes a comparator, a multiplier, an adder, a subtractor and a multiplexer.
9. The privacy information retrieval method based on the CSD array according to claim 1, characterized in that: Each CSD has an independent PCIe channel for direct communication between the Flash storage unit and the FPGA computing unit.
10. A privacy information retrieval system based on a CSD array, characterized in that: The system is used to execute the CSD array-based privacy information retrieval method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for using clock structure for FPGA organized into multiple clock regions
CN113270125A
Keyword privacy information retrieval method and system
CN119046302A