Configuration Method and Device for Field Programmable Gate Array (FPGA)
By obtaining encryption parameters and FPGA specification information, the FPGA configuration is optimized to improve fully homomorphic encryption performance, solving the problem of performance limitations in the existing technology and achieving more efficient performance improvement.
Patent Information
- Application Number
- CN202411492604.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-10-23
AI Technical Summary
In the prior art, the performance of fully homomorphic encryption is affected by the limitations of FPGA hardware resources, resulting in poor performance in practical applications.
By obtaining the encryption parameter information of the target fully homomorphic encryption application and the specification information of the FPGA, determining the split dimension of the cache unit and the number of interfaces of the HP interface, optimizing the configuration of the FPGA to improve the performance of fully homomorphic encryption.
By rationally configuring FPGAs, the performance of fully homomorphic encryption applications can be effectively improved, transmission efficiency can be improved, and transmission bandwidth can be better utilized.
Smart Images

Figure CN119182509B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computers, and in particular, to a configuration method and apparatus for a field programmable gate array (FPGA). Background Art
[0002] Fully homomorphic encryption (FHE) is a privacy computing technology that allows general computations to be directly performed on ciphertexts to obtain ciphertext-based computation results. Its security is based on mathematical hard problems, and it can achieve "usable, controllable, and invisible" data without additional trust assumptions. However, fully homomorphic encryption has deficiencies in terms of performance. Its computing performance usually drops by 4 to 5 orders of magnitude compared to plaintext computing. Therefore, in order to enable fully homomorphic encryption to be used in practical scenarios, its algorithms and applications need to be accelerated.
[0003] The homomorphic operations of fully homomorphic encryption include homomorphic addition, homomorphic multiplication, rescaling, relinearization, and rotation. Through the combination of these homomorphic operations, various homomorphic applications can be realized, including homomorphic matrix multiplication, homomorphic neural networks, etc. Homomorphic operations are also composed of more fundamental operations performed on a polynomial in the ciphertext. The fundamental operations include modular multiplication, modular addition, modular subtraction, automorphism, number-theoretic transform (NTT), and its inverse transform (inverse number-theoretic transform, INTT).
[0004] In the prior art, optimization of the underlying fundamental operations is usually adopted to produce a certain acceleration effect on homomorphic operations and homomorphic applications. Due to the lack of a reasonable configuration method for FPGAs and being restricted by the hardware resources of FPGAs, the performance of fully homomorphic encryption applications is still not good. Summary of the Invention
[0005] One or more embodiments of this specification describe a configuration method and apparatus for an FPGA, which can effectively improve the performance of fully homomorphic encryption applications.
[0006] In a first aspect, a configuration method for a field programmable gate array (FPGA) is provided, including:
[0007] Obtain encryption parameter information of a target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of homomorphic ciphertext data to be processed;
[0008] Obtain the specification information of the FPGA, including the cache parameters of each cache unit and the interface bit width of the high-performance (HP) interface;
[0009] For any target cache unit, determine the maximum parallelism of each operation among several homomorphic basic operations that use the cache unit in the pipeline call mode as the segmentation dimension of the cache unit;
[0010] According to the data bit width and the interface bit width, determine the number of splicings for a single HP interface to transmit spliced data, and according to the segmentation dimension and the number of splicings, determine the number of interfaces of the HP interface allocated to the target cache unit to transmit spliced data;
[0011] Configure the FPGA according to the number of interfaces to execute the target fully homomorphic encryption application.
[0012] In a possible implementation manner, the determining the number of interfaces of the HP interface allocated to the target cache unit to transmit spliced data includes:
[0013] Divide the segmentation dimension by the number of splicings and round up to obtain the first number of HP interfaces allocated to the target cache unit;
[0014] Based on the first number, determine the number of interfaces.
[0015] Further, the determining the number of interfaces based on the first number includes:
[0016] In a given first pipeline stage, obtain the estimated latency time for each homomorphic basic operation to access the target cache unit in parallel;
[0017] If the estimated latency time does not exceed the preset upper limit, use the first number as the number of interfaces in the first pipeline stage;
[0018] If the estimated latency time exceeds the preset upper limit, respectively count the first number of read operations and the second number of write operations on the target cache unit in the first pipeline stage for each homomorphic basic operation, select the larger value from the first number and the second number, and multiply the larger value by the first number as the number of interfaces in the first pipeline stage.
[0019] In a possible implementation, the homomorphic ciphertext data is the polynomial coefficients of a polynomial, and the splicing data includes several consecutive polynomial coefficients of at least one polynomial.
[0020] In a possible implementation, the method further includes:
[0021] Obtain the correspondence between the polynomials of each modulus in the homomorphic ciphertext data and the HP interface;
[0022] Determine the target splicing method of the splicing data according to the correspondence.
[0023] Further, the correspondence indicates that among multiple HP interfaces, one interface is specifically responsible for the transmission of the polynomial of the first modulus; the target splicing method of the splicing data includes:
[0024] Splice the first number of polynomial coefficients of the residue number system (RNS) polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data and transmit it through the first HP interface; the sum of the first number and the second number is equal to the splicing number.
[0025] Further, the correspondence indicates that among multiple HP interfaces, two interfaces jointly are responsible for the transmission of the polynomial of the first modulus; the target splicing method of the splicing data includes:
[0026] Splice the splicing number of polynomial coefficients of the residue number system RNS polynomial of the first modulus of the first ciphertext polynomial into a splicing data and transmit it through the first HP interface;
[0027] Splice the splicing number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data and transmit it through the second HP interface.
[0028] Further, the correspondence indicates that one interface is responsible for the transmission of the polynomial of the first modulus and the transmission of the polynomial of the second modulus; the target splicing method of the splicing data includes:
[0029] Splice the first number of polynomial coefficients of the residue number system RNS polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data and transmit it through the first HP interface within the first time period; the sum of the first number and the second number is equal to the splicing number;
[0030] Concatenate the first number of polynomial coefficients of the residue number system (RNS) polynomial of the second modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the second modulus of the second ciphertext polynomial into a concatenated data, and transmit it through the first HP interface within the second time period.
[0031] In a possible implementation, the method further includes:
[0032] On the premise of minimizing calculation errors, maximizing security levels, and minimizing application latency, determine the number of storage arrays on-chip storage allocated to the target fully homomorphic encryption application according to the segmentation dimensions of each cache unit, the amount of data in each cache unit, the data bit width of the homomorphic ciphertext data, the data depth and data width of a storage array in the on-chip storage.
[0033] Furthermore, the method further includes:
[0034] On the premise of minimizing calculation errors, maximizing security levels, and minimizing application latency, determine the number of digital signal processing (DSP) units allocated to the target fully homomorphic encryption application according to the number of DSP units used by each homomorphic basic operation respectively under the data bit width and the parallelism of each operation.
[0035] In a possible implementation, the method further includes:
[0036] Obtain the usage information corresponding to each cache unit; the usage information is used to indicate the type of cache data stored in the cache unit.
[0037] If the throughput of any cache unit exceeds a preset ratio of the memory capacity and the usage information of this cache unit indicates that the stored cache data is sequentially accessible, then write the cache data of this cache unit from the on-chip storage unit to the off-chip storage unit.
[0038] Furthermore, the usage information includes at least one of the following:
[0039] For storing ciphertext input data, for storing ciphertext data obtained after an automorphism operation, for storing data of the inverse number theoretic transform in a modular addition operation, for storing data of the number theoretic transform in a modular addition operation, for storing data of the number theoretic transform in a modular subtraction operation, for storing the output data of the inner product operation of the last modulus, for storing the output data of the inner product operation of other moduli, for storing rotation factors of the number theoretic transform or the inverse number theoretic transform.
[0040] In a second aspect, a configuration device for a field programmable gate array (FPGA) is provided, including:
[0041] A first acquisition unit, configured to acquire encryption parameter information of a target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of homomorphic ciphertext data to be processed;
[0042] A second acquisition unit, configured to acquire specification information of an FPGA, including cache parameters of each cache unit and the interface bit width of a high-performance HP interface;
[0043] A first determination unit, configured to, for any target cache unit, determine the maximum parallelism of each operation among several homomorphic basic operations using the cache unit in a pipeline call mode as the segmentation dimension of the cache unit;
[0044] A second determination unit, configured to determine the splicing number of splicing data transmitted by a single HP interface according to the data bit width acquired by the first acquisition unit and the interface bit width acquired by the second acquisition unit, and determine the number of HP interfaces allocated to the target cache unit to transmit the splicing data according to the segmentation dimension and the splicing number obtained by the first determination unit;
[0045] A configuration unit, configured to configure the FPGA according to the number of interfaces obtained by the second determination unit to execute the target fully homomorphic encryption application.
[0046] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method in the first aspect.
[0047] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method in the first aspect is implemented.
[0048] Through the methods and devices provided in the embodiments of this specification, first obtain the encryption parameter information of the target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of the homomorphic ciphertext data to be processed; then obtain the specification information of the FPGA, including the cache parameters of each cache unit and the interface bit width of the HP interface; then for any target cache unit, determine the maximum parallelism of each operation in several homomorphic basic operations that use this cache unit in the pipeline call mode as the segmentation dimension of this cache unit; then determine the number of splicings for a single HP interface to transmit the spliced data according to the data bit width and the interface bit width, and determine the number of interface of the HP interface allocated to the target cache unit to transmit the spliced data according to the segmentation dimension and the number of splicings; finally, configure the FPGA according to the number of interfaces to execute the target fully homomorphic encryption application. As can be seen from the above, in the embodiments of this specification, for the case where the interface bit width of the HP interface is greater than the data bit width of the homomorphic ciphertext data, a plurality of homomorphic ciphertext data with small bit widths are spliced into a large-bit-width spliced data, and this spliced data is transmitted through the HP interface, so that when the data stored off-chip is transmitted to on-chip, the transmission bandwidth can be better utilized and the transmission efficiency can be improved. Based on this transmission method, the allocation of the HP interface is more reasonable, and a Pareto configuration solution for a given FPGA and a given application can be obtained. Based on this configuration, the performance of the fully homomorphic encryption application can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0050] Figure 1 Schematic diagram of the implementation scenario of an embodiment disclosed in this specification;
[0051] Figure 2 Flowchart showing the configuration method of the FPGA according to an embodiment;
[0052] Figure 3 Schematic diagram showing the splicing method of the spliced data according to an embodiment;
[0053] Figure 4 Schematic diagram showing the splicing method of the spliced data according to another embodiment;
[0054] Figure 5 Schematic diagram showing the splicing method of the spliced data according to another embodiment;
[0055] Figure 6A schematic diagram showing memory utilization and throughput according to an embodiment;
[0056] Figure 7 A schematic block diagram showing a configuration device of an FPGA according to an embodiment. Detailed implementation manners
[0057] The solution provided in this specification will be described below with reference to the accompanying drawings.
[0058] Figure 1 A schematic diagram of an implementation scenario of an embodiment disclosed in this specification. This implementation scenario involves the configuration of an FPGA. It can be understood that the configured FPGA is used to execute a target fully homomorphic encryption application, and the above configuration can be carried out in combination with existing toolchains. Among them, fully homomorphic encryption: is an encryption technology that has special properties and allows computational operations to be performed on encrypted data in the encrypted state without decrypting the data. This means that calculations can be performed while keeping the data encrypted without revealing the plaintext content of the data. In the embodiments of this specification, the homomorphic ciphertext data can specifically be the polynomial coefficients of a homomorphic ciphertext polynomial, and the homomorphic ciphertext polynomial can be obtained by processing through a residue number system (RNS). RNS: A method of representing a large integer as a combination of several small integers by the Chinese Remainder Theorem, which is commonly used in the representation of large numbers in some cryptographic algorithms. In FHE, a polynomial with large integer coefficients is usually divided into several polynomials with small integer coefficients through RNS.
[0059] Refer to Figure 1, showing the architecture diagram of the optimized framework for the configuration of the FPGA. Among them, the homomorphic encryption application can specifically be a fully homomorphic encryption application. In the embodiments of this specification, with memory management as the optimization core, the performance, security level and other indicators of the homomorphic encryption application are optimized through software and hardware cooperation, which can greatly improve its performance. When facing a fully homomorphic encryption application (HE Application), according to the optimization requirements of the fully homomorphic encryption application, the configuration of the FPGA needs to meet the following design goals: minimizing the application latency while minimizing the computational error and maximizing the security level. According to this design goal, appropriate encryption parameters are initially selected. Through the optimized design of the parameterized homomorphic encryption module, the resources and time of a single module are obtained by configuring the parallelism of the homomorphic encryption module and resource modeling, so as to obtain the resources and time of the entire application. The resource part is restricted by the given FPGA specifications, including DSP resources, on-chip storage BRAM / URAM resources, and the bandwidth of off-chip storage. The encryption parameters also affect the computational error and security level. Here, an evaluation model is established based on existing work, which can also be called a multi-objective optimization model. Through the overall multi-objective optimization model, Pareto solutions can be obtained, including homomorphic encryption parameters, the parallelism of the homomorphic module, the number of HP ports connecting the off-chip storage to the on-chip, and the allocation results. The configured and implemented homomorphic encryption application is transformed into register transfer level (RTL) language through high-level synthesis (HLS) technology, and then deployed on the FPGA through the Vivado tool chain.
[0060] Based on this framework, the specific optimization measures taken in the embodiments of this specification include the optimization of the HE module and the multi-objective design space exploration model. In terms of the resource management of the model, a memory allocation scheme for on-chip and off-chip of the FPGA is given, and a storage resource model is given on this basis. A multi-objective optimization model for performance, security level, and computational error is established, which can automatically generate Pareto solutions for application parameters and hardware configurations.
[0061] Among them, HP ports (high-performance ports): also known as HP interfaces, refer to the interfaces provided inside the FPGA chip with high-speed data transmission and processing capabilities. These ports usually correspond to high-speed serial communication protocols (such as PCIe, Ethernet, etc.) or parallel buses (such as DDR, AXI, etc.), and support high-speed data transmission and processing.
[0062] BRAM (block random access memory): is a dedicated RAM resource in the FPGA, which is fixedly distributed at specific positions inside the FPGA.
[0063] URAM (UltraRAM): UltraRAM is a unique storage resource in UltraScale Plus chips.
[0064] In the embodiments of this specification, off-chip storage may but is not limited to include the following two types:
[0065] DDR: Double Data Rate SDRAM, usually memory chip particles located outside the FPGA chip.
[0066] HBM: High Bandwidth Memory, usually memory chip particles located outside the FPGA chip, which has a higher read / write bandwidth rate compared to DDR.
[0067] The acceleration implementation of fully homomorphic encryption on hardware is mainly restricted by storage resources, which affects performance improvement. It is mainly reflected that the on-chip storage resources of the FPGA are very limited for homomorphic ciphertext. The on-chip storage provides data for the basic operators on the chip and stores temporary data, and it is also related to the number of basic operators that are parallel at the current moment. When the computing resources are fully utilized and a large number of basic operator modules are parallelized, the on-chip storage is often insufficient, which poses requirements for the efficient utilization of on-chip storage. In the embodiments of this specification, by reducing unnecessary on-chip storage caches and using these storage spaces for parallel computing units, the performance can be effectively improved.
[0068] In addition, a large amount of homomorphic ciphertext data is usually stored in off-chip memory, and the timing of data transfer between on-chip and off-chip storage spaces and the utilization of data transfer bandwidth also become factors affecting performance. The data transfer of the FPGA is connected from the on-chip storage to the off-chip DDR or HBM memory through the HP interface, and the number of HP interfaces is limited. In the embodiments of this specification, corresponding solutions are proposed for the efficient utilization of HP interfaces throughout the computing process. Among them, the data that needs to be transferred to the on-chip can be immediately used instead of waiting for a period of time before being used. On this basis, it is also necessary to consider that the data transfer time can be masked by performing other calculations simultaneously.
[0069] In the embodiments of this specification, based on the above optimizations, combined with the parameter settings of homomorphic applications and the existing parameterized homomorphic operation modules, it is possible to perform a co-design space exploration of software and hardware, and obtain the Pareto configuration solutions for a given FPGA and a given application for multiple objectives such as performance, accuracy, and security level.
[0070] Figure 2 Show a flowchart of the configuration method of an FPGA according to an embodiment, and this method can be based on Figure 1 the shown implementation scenario. As Figure 2As shown in the figure, the configuration method of the FPGA in this embodiment includes the following steps: Step 21, obtaining the encryption parameter information of the target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of the homomorphic ciphertext data to be processed; Step 22, obtaining the specification information of the FPGA, including the cache parameters of each cache unit and the interface bit width of the high-performance HP interface; Step 23, for any target cache unit, determining the maximum parallelism of each operation among several homomorphic basic operations using this cache unit in the pipeline call mode as the splitting dimension of this cache unit; Step 24, determining the splicing number of the splicing data transmitted by a single HP interface according to the data bit width and the interface bit width, and determining the number of HP interfaces allocated to the target cache unit to transmit the splicing data according to the splitting dimension and the splicing number; Step 25, configuring the FPGA according to the number of interfaces to execute the target fully homomorphic encryption application. The specific execution methods of the above steps are described below.
[0071] First, in Step 21, obtain the encryption parameter information of the target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of the homomorphic ciphertext data to be processed. It can be understood that the homomorphic basic operations are used to implement some basic operations, such as modular multiplication, modular addition, modular subtraction, automorphism, NTT, and INTT, etc. The homomorphic ciphertext data is the polynomial coefficients of a polynomial.
[0072] In the embodiment of this specification, the data bit width of the homomorphic ciphertext data can affect the calculation error and security level of the target fully homomorphic encryption application. The data bit width of the homomorphic ciphertext data can be determined according to the requirements of the target fully homomorphic encryption application for the calculation error and security level.
[0073] Then, in Step 22, obtain the specification information of the FPGA, including the cache parameters of each cache unit and the interface bit width of the high-performance HP interface. It can be understood that the cache parameters include the cache size, etc., and the interface bit width refers to the data transmission width of the HP interface. For example, the HP interface can implement a data bit width of 32 or 64 bits.
[0074] In the embodiment of this specification, the specification information of the FPGA reflects the limitations of its hardware resources. For example, the upper limit of DSP resources, the upper limit of on-chip storage BRAM / URAM resources, and the upper limit of the bandwidth of off-chip storage. The limitations of hardware resources will affect the latency of the application. Corresponding solutions are proposed for the efficient utilization of the HP interface. Among them, since the interface bit width of the HP interface is greater than the data bit width of the homomorphic ciphertext data, multiple homomorphic ciphertext data with a small bit width are spliced into a large-bit-width splicing data to be transmitted through the HP interface, so as to effectively utilize the transmission capacity of the HP interface.
[0075] Next, in step 23, for any target cache unit, the maximum parallelism of each operation among several homomorphic basic operations that use this cache unit in the pipeline call mode is determined as the splitting dimension of this cache unit. It can be understood that in the pipeline call mode, the calls of multiple homomorphic basic operation modules are pipelined.
[0076] Among them, in a computer, the pipeline call mode refers to the application of pipeline technology, which is a technology that improves system performance by decomposing a task into multiple subtasks and enabling these subtasks to be executed in parallel at different stages. Specifically, pipeline technology divides the operations of an instruction into multiple subtasks, and each subtask is processed by a dedicated functional component in turn, so as to achieve a quasi-parallel processing implementation technology of overlapping execution of multiple instructions.
[0077] For example, the parallelism within each homomorphic basic operation is denoted as pop i , where op i represents an operation in the OP set. For the cache unit (buffer) b j used by the homomorphic basic operation, its splitting dimension is denoted as pb j . To adapt to the parallelism of different homomorphic basic operations, the maximum parallelism is specified as the splitting dimension of this cache unit, and the splitting dimension reflects the data parallelism of the cache unit. Specifically, if the cache unit b j is used by the operation op i , it is expressed as , otherwise, . The splitting dimension j of the cache unit b can be obtained.
[0078] In the embodiments of this specification, the above splitting dimension can be regarded as a cache parameter of the cache unit.
[0079] Then, in step 24, according to the data bit width and the interface bit width, the number of splicings for a single HP interface to transmit the spliced data is determined, and according to the splitting dimension and the number of splicings, the number of interfaces of the HP interface allocated to the target cache unit to transmit the spliced data is determined. It can be understood that the determination of the number of interfaces is based on the premise of efficiently using the HP interface.
[0080] In the embodiments of this specification, the data that needs to be transferred to the chip can be used immediately instead of waiting for a period of time. On this basis, it is also necessary to consider that the data transmission time can be masked while other calculations are being performed. On this basis, the number of interfaces of the HP interface allocated to the target cache unit to transmit the spliced data is determined.
[0081] In one example, determining the number of interfaces of the HP interface assigned to the target cache unit for transmitting the spliced data includes:
[0082] Dividing the segmentation dimension by the number of splicings and rounding up to obtain a first number of HP interfaces assigned to the target cache unit;
[0083] Based on this first number, determine the number of interfaces.
[0084] In this example, the first number can be directly used as the number of interfaces, or after certain detections, based on the detection results and on the basis of the first number, determine the number of interfaces.
[0085] For example, the interface bit width of the HP interface is denoted as , and the data bit width of the homomorphic ciphertext data is denoted as , and the number of digits spliced by one HP interface can be obtained as , and further the number of HP interfaces required by cache unit b j is .
[0086] Further, based on this first number, determining the number of interfaces includes:
[0087] In a given first pipeline stage, obtain the estimated latency for each homomorphic basic operation to access the target cache unit in parallel;
[0088] If the estimated latency does not exceed the preset upper limit, use the first number as the number of interfaces in this first pipeline stage;
[0089] If the estimated latency exceeds the preset upper limit, respectively count the first number of read operations and the second number of write operations on the target cache unit in this first pipeline stage for each homomorphic basic operation, select the larger value from the first number and the second number, and multiply the larger value by the first number as the number of interfaces in this first pipeline stage.
[0090] In this example, if the estimated latency exceeds the preset upper limit, it indicates that there is an access conflict. If the estimated latency does not exceed the preset upper limit, it indicates that there is no access conflict. Determine the number of interfaces according to whether there is an access conflict.
[0091] For example, since the off-chip memory DDR can read and write simultaneously, while the buffer may only be used for reading or writing, this will cause waste of the interface bandwidth of the HP interface. If the buffers used for ping-pong in the pipeline technology are connected to the same group of HP interfaces, multiple basic functions may access the same HP interface, exceeding the one-read-one-write supported by the DDR and causing access conflicts. In the embodiments of this specification, a model is constructed to detect conflicts and allocate sufficient HP interfaces in case of conflicts. Specifically: in a given pipeline stage, if the latency of all parallel operations accessing buffer b j does not exceed the upper limit Lat up , there is no access conflict, and the number of HP interfaces remains unchanged. On the contrary, if a conflict is detected, the read and write operations on buffer b j in each operation are counted. The existence of a read operation on this buffer is represented as , where i represents the index of operation op i , otherwise . Similarly, the existence of a write operation on this buffer is represented as , otherwise . Denote the number of HP interfaces for buffer b j after considering access conflicts as , and the formula can be derived accordingly:
[0092]
[0093] Denote the pipeline stage x as stage x , and the basic operation used in this stage is , then the number of interfaces of the total HP interfaces used in stage x can be obtained:
[0094] ;
[0095] The number of interfaces of the total HP interfaces used in each pipeline stage is .
[0096] In one example, the homomorphic ciphertext data is the polynomial coefficients of a polynomial, and the concatenated data includes several consecutive polynomial coefficients of at least one polynomial.
[0097] In this example, for an RNS polynomial of a ciphertext, consecutive data can be placed at one address, so as to facilitate sequential access to the homomorphic ciphertext data.
[0098] In one example, the method further includes:
[0099] Obtain the correspondence between the polynomials of each modulus in the homomorphic ciphertext data and the HP interfaces;
[0100] Determine the target splicing method of the splicing data according to the corresponding relationship.
[0101] In this example, for the RNS polynomial of a ciphertext, continuous data can be placed in one address, and the ciphertext polynomial and are spliced. The data splicing scheme can be flexibly adjusted according to the available HP interfaces and the target transmission period. The above corresponding relationship reflects the specific data splicing scheme.
[0102] Furthermore, the corresponding relationship indicates that among multiple HP interfaces, one interface is specifically responsible for transmitting the polynomial of the first modulus; the target splicing method of the splicing data includes:
[0103] Splice the first number of polynomial coefficients of the RNS polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data, and transmit it through the first HP interface; the sum of the first number and the second number is equal to the splicing number.
[0104] In this example, an independent HP interface is allocated for the RNS polynomial of each modulus. Therefore, for the RNS polynomials of two moduli, they can be transmitted through two HP interfaces.
[0105] Figure 3 Show a schematic diagram of the splicing method of splicing data according to an embodiment. Refer to Figure 3 , in this splicing method, among multiple HP interfaces, one interface is specifically responsible for transmitting the polynomial of the first modulus. There are two HP interfaces in total, denoted as HP0 and HP1 respectively. Among them, HP0 is responsible for transmitting the RNS polynomial with modulus q 0 , and HP1 is responsible for transmitting the RNS polynomial with modulus q 1 . The ciphertext polynomial c 0 can be divided into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 . The ciphertext polynomial c 1 can also be divided into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 . The polynomial coefficients of the ciphertext polynomial c 0 and the ciphertext polynomial c 1 are distinguished by rectangles of different colors. The numbers in the rectangles identify the positions of the polynomial coefficients. For example, the numbers in two adjacent rectangles are 0 and 1 respectively, representing that the positions of the two polynomial coefficients are adjacent, and these two polynomial coefficients are continuous data. The modulus of the ciphertext polynomial c 0 is q0 Two polynomial coefficients of the RNS polynomial, and the ciphertext polynomial c 1 The modulus is q 0 Two polynomial coefficients of the RNS polynomial are concatenated into a concatenated data and transmitted through HP0; the ciphertext polynomial c 0 The modulus is q 1 Two polynomial coefficients of the RNS polynomial, and the ciphertext polynomial c 1 The modulus is q 1 Two polynomial coefficients of the RNS polynomial are concatenated into a concatenated data and transmitted through HP1. This concatenation method can achieve the performance of N / 2 cycle reading of two HP interfaces.
[0106] Furthermore, the corresponding relationship indicates that among multiple HP interfaces, two interfaces are jointly responsible for the transmission of the polynomial of the first modulus; the target concatenation method of the concatenated data includes:
[0107] Concatenate the said number of concatenated polynomial coefficients of the residue number system RNS polynomial of the first modulus of the first ciphertext polynomial into a concatenated data and transmit it through the first HP interface;
[0108] Concatenate the said number of concatenated polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a concatenated data and transmit it through the second HP interface.
[0109] In this example, two different HP interfaces are allocated for the RNS polynomial of each modulus. Thus, for the RNS polynomials of two moduli, they can be transmitted through four HP interfaces.
[0110] Figure 4 Shows a schematic diagram of the concatenation method of the concatenated data according to another embodiment. Refer to Figure 4 , in this concatenation method, among multiple HP interfaces, two interfaces are jointly responsible for the transmission of the polynomial of the first modulus. There are a total of four HP interfaces, denoted as HP0, HP1, HP2, and HP3 respectively. Among them, HP0 and HP1 are responsible for transmitting the RNS polynomial with modulus q 0 , and HP2 and HP3 are responsible for transmitting the RNS polynomial with modulus q 1 , the ciphertext polynomial c 0 can be divided into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 , the ciphertext polynomial c 1 can also be divided into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 , the ciphertext polynomial c 0 and the ciphertext polynomial c1 The polynomial coefficients are distinguished by rectangles of different colors, and the numbers in the rectangles identify the positions of the polynomial coefficients. For example, the numbers in two adjacent rectangles are 0 and 1 respectively, indicating that the positions of the two polynomial coefficients are adjacent, and these two polynomial coefficients are consecutive data. The ciphertext polynomial c 0 has a modulus of q 0 The 4 polynomial coefficients of the RNS polynomial of the ciphertext polynomial c 1 with a modulus of q 0 are concatenated into a concatenated data and transmitted through HP0; the 4 polynomial coefficients of the RNS polynomial of the ciphertext polynomial c 0 with a modulus of q 1 are concatenated into a concatenated data and transmitted through HP1; the 4 polynomial coefficients of the RNS polynomial of the ciphertext polynomial c 1 with a modulus of q 1 are concatenated into a concatenated data and transmitted through HP2; the 4 polynomial coefficients of the RNS polynomial of the ciphertext polynomial c
[0111] Furthermore, the corresponding relationship indicates that one interface is responsible for the transmission of polynomials of the first modulus and also for the transmission of polynomials of the second modulus; the target concatenation method of the concatenated data includes:
[0112] Concatenate the first number of polynomial coefficients of the RNS polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a concatenated data and transmit it through the first HP interface within the first time period; the sum of the first number and the second number is equal to the concatenation number;
[0113] Concatenate the first number of polynomial coefficients of the RNS polynomial of the second modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the second modulus of the second ciphertext polynomial into a concatenated data and transmit it through the first HP interface within the second time period.
[0114] In this example, a time - shared HP interface is allocated for the RNS polynomials of two moduli, so that for the RNS polynomials of two moduli, they can be transmitted through one HP interface.
[0115] Figure 5 Shows a schematic diagram of the concatenation method of the concatenated data according to another embodiment. Refer to Figure 5, in this splicing method, one interface is responsible for the transmission of polynomials of the first modulus and the transmission of polynomials of the second modulus. There is a total of one HP interface, denoted as HP0. Among them, HP0 is responsible for transmitting the RNS polynomial with modulus q 0 and is also responsible for transmitting the RNS polynomial with modulus q 1 . The ciphertext polynomial c 0 can be split into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 . The ciphertext polynomial c 1 can also be split into the RNS polynomial with modulus q 0 and the RNS polynomial with modulus q 1 . The polynomial coefficients of the ciphertext polynomial c 0 and the ciphertext polynomial c 1 are distinguished by rectangles of different colors. The numbers in the rectangles identify the positions of the polynomial coefficients. For example, the numbers in two adjacent rectangles are 0 and 1 respectively, indicating that the positions of the two polynomial coefficients are adjacent, and these two polynomial coefficients are consecutive data. The 2 polynomial coefficients of the RNS polynomial with modulus q 0 of the ciphertext polynomial c 0 , and the 2 polynomial coefficients of the RNS polynomial with modulus q 1 of the ciphertext polynomial c 0 are spliced into a spliced data and transmitted through HP0 within the time period of [0, N / 2 - 1]; the 2 polynomial coefficients of the RNS polynomial with modulus q 0 of the ciphertext polynomial c 1 , and the 2 polynomial coefficients of the RNS polynomial with modulus q 1 of the ciphertext polynomial c 1 are spliced into a spliced data and transmitted through HP0 within the time period of [N / 2, N - 1]. This splicing method can achieve the performance of N-cycle reading of one HP interface.
[0116] Finally, in step 25, configure the FPGA according to the number of the interfaces to execute the target fully homomorphic encryption application. It can be understood that the configured and implemented HE application can be converted into RTL language through high-level synthesis technology and then deployed on the FPGA through the Vivado toolchain.
[0117] In the embodiments of this specification, in addition to the number of the interfaces, the FPGA can also be configured according to the number of storage arrays of the on-chip storage allocated to the target fully homomorphic encryption application, and / or the number of DSP units allocated to the target fully homomorphic encryption application.
[0118] In one example, the method further includes:
[0119] On the premise of minimizing the calculation error, maximizing the security level, and minimizing the application latency, determine the number of memory arrays on-chip storage allocated to the target fully homomorphic encryption application according to the segmentation dimension of each cache unit, the data quantity of each cache unit, the data bit width of the homomorphic ciphertext data, the data depth and data width of a memory array in on-chip storage.
[0120] In this example, the limitations on FPGA storage resources are considered, and the allocation and utilization of storage resources are optimized. Since the homomorphic encryption application is a memory access-intensive application, the allocation of storage resources will have a greater impact on performance.
[0121] For example, Lat represents the application latency, Err represents the calculation error, and Sec represents the security level. The design goal can be expressed as . Considering the limitations of the FPGA's hardware resources, it can be expressed as the following conditions:
[0122]
[0123] The formula for the on-chip storage BRAM resources is as follows:
[0124]
[0125] Among them, the cache set is represented as B, the set of on-chip caches after off-chip cache allocation is denoted as B_ON, and each cache b in B_ON j has a segmentation dimension of pb j . size(b j ) represents the data quantity of this cache, and q width represents the data bit width of each homomorphic ciphertext data. BRAM_depth and BRAM_width respectively represent the data depth and data width of a memory array (bank) of BRAM.
[0126] Furthermore, the method further includes:
[0127] On the premise of minimizing the calculation error, maximizing the security level, and minimizing the application latency, determine the number of DSP units allocated to the target fully homomorphic encryption application according to the number of digital signal processing DSP units respectively used by each homomorphic basic operation under the data bit width and the parallelism of each operation.
[0128] In this example, the limitations on the FPGA's DSP resources are considered, and the allocation and utilization of DSP resources are optimized to improve the performance of the homomorphic encryption application.
[0129] For example, the formula for DSP resources is as follows:
[0130]
[0131] Among them, OP is the set of homomorphic basic operations used by a given application, and op i represents one of the homomorphic basic operations. represents the basic operation op i in q width the number of DSPs used under the bit width, represents the parallelism of this basic operation.
[0132] In the embodiments of this specification, the memory management during the calculation process of the homomorphic encryption application can also be optimized, and the effect of this optimization can be integrated in multi-objective optimization. Memory management optimization includes reasonably allocating data on-chip and off-chip.
[0133] In one example, the method further includes:
[0134] obtaining the usage information corresponding to each cache unit; the usage information is used to indicate the type of cache data stored in the cache unit;
[0135] If the throughput of any cache unit exceeds a preset ratio of the memory capacity, and the usage information of this cache unit indicates that the stored cache data is sequentially accessible, then write the cache data of this cache unit from the on-chip storage unit to the off-chip storage unit.
[0136] In this example, in the case where the on-chip BRAM / URAM becomes a performance bottleneck, the data that can be sequentially accessed on-chip can be stored off-chip, and the off-chip storage can be DDR or HBM, etc.
[0137] Further, the usage information includes at least one of the following:
[0138] used to store ciphertext input data, used to store ciphertext data obtained after the automorphism operation, used to store data of the inverse number theory transform in the modular addition operation, used to store data of the number theory transform in the modular addition operation, used to store data of the number theory transform in the modular subtraction operation, used to store the output data of the inner product operation of the last modulus, used to store the output data of the inner product operation of other moduli, used to store rotation factors of the number theory transform or the inverse number theory transform.
[0139] The following Table 1 shows some examples of the correspondence between the cache number and the cache usage.
[0140] Table 1: Correspondence between cache number and cache usage
[0141]
[0142] As can be seen from Table 1, the purpose of the cache can be found through the cache number, so as to determine whether the data in the cache can be accessed sequentially.
[0143] Figure 6 FIG. shows a schematic diagram of memory utilization and throughput according to an embodiment. Refer to Figure 6 , since the throughputs of B2, B6, and B7 are relatively high, for the KeySwitch module, the caches for the automorphism operation, the output cache, and the inner product operation can be placed off-chip, without causing a significant performance degradation due to non-burst access.
[0144] In the embodiments of this specification, a new solution is proposed for on-chip / off-chip memory co-optimization. By analyzing the data characteristics in the homomorphic operation calculation process, appropriate data is placed in the off-chip memory. On the one hand, it can reduce the demand for on-chip BRAM, and on the other hand, it can increase the utilization rate of the off-chip bandwidth. By using the saved BRAM to provide data access for the increased parallel homomorphic operators, a 60% performance improvement is achieved.
[0145] Through the method provided in the embodiments of this specification, first, obtain the encryption parameter information of the target fully homomorphic encryption application, including multiple homomorphic basic operations involved and the data bit width of the homomorphic ciphertext data to be processed; then obtain the specification information of the FPGA, including the cache parameters of each cache unit and the interface bit width of the HP interface; then for any target cache unit, determine the maximum parallelism of each operation in several homomorphic basic operations that use this cache unit in the pipeline call mode as the splitting dimension of this cache unit; then according to the data bit width and the interface bit width, determine the number of splicings for a single HP interface to transmit the spliced data, and according to the splitting dimension and the number of splicings, determine the number of HP interfaces allocated to the target cache unit to transmit the spliced data; finally, configure the FPGA according to the number of interfaces to execute the target fully homomorphic encryption application. As can be seen from the above, in the embodiments of this specification, for the case where the interface bit width of the HP interface is greater than the data bit width of the homomorphic ciphertext data, multiple small-bit-width homomorphic ciphertext data are spliced into a large-bit-width spliced data, and this spliced data is transmitted through the HP interface, so that when the data stored off-chip is transmitted to the on-chip, the transmission bandwidth can be better utilized and the transmission efficiency can be improved. Based on this transmission method, the allocation of the HP interface is more reasonable, and a Pareto configuration solution for a given FPGA and a given application can be obtained. Based on this configuration, the performance of the fully homomorphic encryption application can be effectively improved.
[0146] According to an embodiment of another aspect, there is also provided a configuration device for a field programmable gate array (FPGA), and this device is used to execute the method provided in the embodiments of this specification. Figure 7Schematic block diagram showing a configuration device of an FPGA according to an embodiment. As Figure 7 shown, the device 700 includes:
[0147] A first acquisition unit 71, configured to acquire encryption parameter information of a target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of homomorphic ciphertext data to be processed;
[0148] A second acquisition unit 72, configured to acquire specification information of the FPGA, including cache parameters of each cache unit and the interface bit width of the high-performance HP interface;
[0149] A first determination unit 73, configured to, for any target cache unit, determine the maximum parallelism of each operation among several homomorphic basic operations using the cache unit in the pipeline call mode as the splitting dimension of the cache unit;
[0150] A second determination unit 74, configured to determine the splicing number of the spliced data transmitted by a single HP interface according to the data bit width acquired by the first acquisition unit 71 and the interface bit width acquired by the second acquisition unit 72, and determine the number of interfaces of the HP interface allocated to the target cache unit to transmit the spliced data according to the splitting dimension and the splicing number obtained by the first determination unit 73;
[0151] A configuration unit 75, configured to configure the FPGA according to the number of interfaces obtained by the second determination unit 74 to execute the target fully homomorphic encryption application.
[0152] Optionally, as an embodiment, the second determination unit 74 includes:
[0153] A first allocation subunit, configured to divide the splitting dimension by the splicing number and round up to obtain a first number of HP interfaces allocated to the target cache unit;
[0154] A determination subunit, configured to determine the number of interfaces based on the first number obtained by the first allocation subunit.
[0155] Further, the determination subunit is specifically configured to:
[0156] In a given first pipeline stage, acquire the estimated latency time for each homomorphic basic operation to access the target cache unit in parallel;
[0157] If the estimated latency time does not exceed a preset upper limit, use the first number as the number of interfaces in the first pipeline stage;
[0158] If the estimated delay time exceeds the preset upper limit, the first number of read operations on the target cache unit and the second number of write operations in the first pipeline stage in each homomorphic basic operation are respectively counted, the larger value is selected from the first number and the second number, and the larger value is multiplied by the first quantity to obtain the interface quantity of the first pipeline stage.
[0159] Optionally, as an embodiment, the homomorphic ciphertext data is the polynomial coefficients of a polynomial, and the splicing data includes several consecutive polynomial coefficients of at least one polynomial.
[0160] Optionally, as an embodiment, the apparatus further includes:
[0161] A third acquisition unit, configured to acquire the correspondence between the polynomials of each modulus in the homomorphic ciphertext data and the HP interfaces;
[0162] A third determination unit, configured to determine the target splicing method of the splicing data according to the correspondence acquired by the third acquisition unit.
[0163] Further, the correspondence indicates that among multiple HP interfaces, one interface is specifically responsible for the transmission of the polynomial of the first modulus; the target splicing method of the splicing data includes:
[0164] Splicing the first number of polynomial coefficients of the residue number system (RNS) polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data, and transmitting it through the first HP interface; the sum of the first number and the second number is equal to the splicing number.
[0165] Further, the correspondence indicates that among multiple HP interfaces, two interfaces jointly are responsible for the transmission of the polynomial of the first modulus; the target splicing method of the splicing data includes:
[0166] Splicing the splicing number of polynomial coefficients of the residue number system (RNS) polynomial of the first modulus of the first ciphertext polynomial into a splicing data, and transmitting it through the first HP interface;
[0167] Splicing the splicing number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a splicing data, and transmitting it through the second HP interface.
[0168] Further, the correspondence indicates that one interface is responsible for both the transmission of the polynomial of the first modulus and the transmission of the polynomial of the second modulus; the target splicing method of the splicing data includes:
[0169] Concatenate the first number of polynomial coefficients of the residue number system (RNS) polynomial of the first modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial into a concatenated data, and transmit it through the first HP interface within the first time period; the sum of the first number and the second number is equal to the concatenation number;
[0170] Concatenate the first number of polynomial coefficients of the RNS polynomial of the second modulus of the first ciphertext polynomial and the second number of polynomial coefficients of the RNS polynomial of the second modulus of the second ciphertext polynomial into a concatenated data, and transmit it through the first HP interface within the second time period.
[0171] Optionally, as an embodiment, the device further includes:
[0172] A fourth determination unit, configured to determine the number of storage arrays of the on-chip storage allocated to the target fully homomorphic encryption application according to the partitioning dimensions of each cache unit, the data quantity of each cache unit, the data bit width of the homomorphic ciphertext data, the data depth and data width of a storage array in the on-chip storage, on the premise of minimizing the calculation error, maximizing the security level, and minimizing the application delay.
[0173] Further, the device further includes:
[0174] A fifth determination unit, configured to determine the number of DSP units allocated to the target fully homomorphic encryption application according to the number of digital signal processing (DSP) units respectively used by each homomorphic basic operation under the data bit width and the parallelism of each operation, on the premise of minimizing the calculation error, maximizing the security level, and minimizing the application delay.
[0175] Optionally, as an embodiment, the device further includes:
[0176] A fourth acquisition unit, configured to acquire the usage information corresponding to each cache unit; the usage information is used to indicate the type of cache data stored in the cache unit;
[0177] A writing unit, configured to write the cache data of the cache unit from the on-chip storage unit to the off-chip storage unit if the throughput of any cache unit exceeds a preset ratio of the memory capacity and the usage information of the cache unit indicates that the stored cache data is sequentially accessible.
[0178] Further, the usage information includes at least one of the following:
[0179] For storing ciphertext input data, for storing ciphertext data obtained after an automorphism operation, for storing data of an inverse number-theoretic transform in a modular addition operation, for storing data of a number-theoretic transform in a modular addition operation, for storing data of a number-theoretic transform in a modular subtraction operation, for storing output data of an inner product operation of the last modulus, for storing output data of an inner product operation of other moduli, and for storing rotation factors of a number-theoretic transform or an inverse number-theoretic transform.
[0180] Through the device provided by the embodiments of this specification, first, the first acquisition unit 71 acquires encryption parameter information of a target fully homomorphic encryption application, including a plurality of homomorphic basic operations involved and the data bit width of homomorphic ciphertext data to be processed; then, the second acquisition unit 72 acquires specification information of the FPGA, including cache parameters of each cache unit and the interface bit width of the HP interface; next, for any target cache unit, the first determination unit 73 determines the maximum parallelism of each operation in several homomorphic basic operations that use this cache unit in the pipeline call mode as the segmentation dimension of this cache unit; then, the second determination unit 74 determines the number of splicings for a single HP interface to transmit spliced data according to the data bit width and the interface bit width, and determines the number of interface of the HP interface allocated to the target cache unit to transmit spliced data according to the segmentation dimension and the number of splicings; finally, the configuration unit 75 configures the FPGA according to the number of interfaces to execute the target fully homomorphic encryption application. As can be seen from the above, in the embodiments of this specification, for the case where the interface bit width of the HP interface is greater than the data bit width of homomorphic ciphertext data, a plurality of homomorphic ciphertext data with a small bit width are spliced into a large-bit-width spliced data, and the spliced data is transmitted through the HP interface, so that when the data stored off-chip is transmitted to on-chip, the transmission bandwidth can be better utilized and the transmission efficiency can be improved. Based on this transmission method, the allocation of the HP interface is more reasonable, and a Pareto configuration solution for a given FPGA and a given application can be obtained. Based on this configuration, the performance of the fully homomorphic encryption application can be effectively improved.
[0181] According to an embodiment of another aspect, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in conjunction with Figure 2 what is described.
[0182] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in conjunction with Figure 2 what is described is implemented.
[0183] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0184] The specific embodiments described above have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for configuring a field programmable gate array (FPGA), comprising: Obtain encryption parameter information of the target fully homomorphic encryption application, including multiple homomorphic basic operations involved and the data bit width of the homomorphic ciphertext data to be processed; Obtain FPGA specification information, including cache parameters of each cache unit and interface width of high-performance HP interface; For any target cache unit, the maximum parallelism of each operation among several homomorphic basic operations using the cache unit in the pipeline call mode is determined as the segmentation dimension of the cache unit; Determine the number of splicings of the spliced data transmitted by a single HP interface according to the data bit width and the interface bit width, and determine the number of interfaces of the HP interface allocated to the target cache unit for transmitting the spliced data according to the segmentation dimension and the number of splicings; According to the number of interfaces, the FPGA is configured to execute the target fully homomorphic encryption application.
2. The method of claim 1, wherein: The determining the number of interfaces of the HP interface allocated to the target cache unit for transmitting the spliced data includes: Divide the segmentation dimension by the number of splicing and round up to an integer, to obtain a first number of HP interfaces allocated to the target cache unit; Based on the first number, the number of interfaces is determined.
3. The method of claim 2, wherein: The determining the number of interfaces based on the first number includes: At a given first pipeline stage, obtaining an estimated delay time of each homomorphic basic operation accessing the target cache unit in parallel; If the estimated delay time does not exceed a preset upper limit, using the first number as the number of interfaces in the first pipeline stage; If the estimated delay time exceeds the preset upper limit, the first number of read operations and the second number of write operations on the target cache unit in the first pipeline stage in each homomorphic basic operation are counted respectively, and the larger value is selected from the first number and the second number, and the larger value is multiplied by the first number as the number of interfaces of the first pipeline stage.
4. The method of claim 1, wherein: The homomorphic ciphertext data are polynomial coefficients of a polynomial, and the concatenated data include several consecutive polynomial coefficients of at least one polynomial.
5. The method of claim 1, wherein: The method further comprises: Obtain the corresponding relationship between the polynomial of each modulus in the homomorphic ciphertext data and the HP interface; According to the corresponding relationship, a target splicing mode of the splicing data is determined.
6. The method of claim 5, wherein: The corresponding relationship indicates that, among the multiple HP interfaces, one interface is specifically responsible for the transmission of the polynomial of the first modulus; The target splicing method of the splicing data includes: A first number of polynomial coefficients of a remainder system RNS polynomial of a first modulus of a first ciphertext polynomial and a second number of polynomial coefficients of a RNS polynomial of a first modulus of a second ciphertext polynomial are spliced into a spliced data, which is transmitted through a first HP interface; the sum of the first number and the second number is equal to the spliced number.
7. The method of claim 5, wherein: The corresponding relationship indicates that two interfaces among the multiple HP interfaces are jointly responsible for the transmission of the polynomial of the first modulus; the target splicing mode of the spliced data includes: Splicing the spliced number of polynomial coefficients of the remainder system RNS polynomial of the first modulus of the first ciphertext polynomial into a spliced data, and transmitting it through the first HP interface; The concatenated number of polynomial coefficients of the RNS polynomial of the first modulus of the second ciphertext polynomial are concatenated into a concatenated data, which is transmitted through the second HP interface.
8. The method of claim 5, wherein: The corresponding relationship indicates that one interface is responsible for both the transmission of the polynomial of the first modulus and the transmission of the polynomial of the second modulus; The target splicing method of the splicing data includes: splicing a first number of polynomial coefficients of a remainder system RNS polynomial of a first modulus of a first ciphertext polynomial and a second number of polynomial coefficients of a RNS polynomial of a first modulus of a second ciphertext polynomial into a spliced data, and transmitting the spliced data through a first HP interface within a first time period; the sum of the first number and the second number is equal to the spliced number; A first number of polynomial coefficients of the RNS polynomial of the second modulus of the first ciphertext polynomial and a second number of polynomial coefficients of the remainder system RNS polynomial of the second modulus of the second ciphertext polynomial are concatenated into a concatenated data, which is transmitted through the first HP interface within a second time period.
9. The method of claim 1, wherein: The method further comprises: Under the premise of minimizing computational errors, maximizing security levels, and minimizing application delays, the number of on-chip storage arrays allocated to the target fully homomorphic encryption application is determined based on the segmentation dimensions of each cache unit, the amount of data in each cache unit, the data bit width of the homomorphic ciphertext data, and the data depth and data width of a storage array stored on the chip.
10. The method of claim 9, wherein: The method further comprises: Under the premise of minimizing computational errors, maximizing security levels, and minimizing application delays, the number of DSP units allocated to the target fully homomorphic encryption application is determined based on the number of digital signal processing DSP units used by each homomorphic basic operation under the data bit width and the parallelism of each operation.
11. The method of claim 1, wherein: The method further comprises: Obtaining usage information corresponding to each cache unit; the usage information is used to indicate the type of cache data stored in the cache unit; If the throughput of any cache unit exceeds a preset ratio of the memory capacity, and the usage information of the cache unit indicates that it stores cache data that can be accessed sequentially, the cache data of the cache unit is written from the on-chip storage unit to the off-chip storage unit.
12. The method of claim 11, wherein: The usage information includes at least one of the following: Used to store ciphertext input data, used to store ciphertext data obtained after automorphism operation, used to store data of inverse number theory transformation in modular addition operation, used to store data of number theory transformation in modular addition operation, used to store data of number theory transformation in modular subtraction operation, used to store output data of inner product operation of the last modulus, used to store output data of inner product operation of other modulus, used to store rotation factors of number theory transformation or inverse number theory transformation.
13. A configuration device for a field programmable gate array (FPGA), comprising: A first acquisition unit is used to acquire encryption parameter information of a target fully homomorphic encryption application, including multiple homomorphic basic operations involved and a data bit width of homomorphic ciphertext data to be processed; A second acquisition unit is used to acquire specification information of the FPGA, including cache parameters of each cache unit and an interface width of a high-performance HP interface; A first determining unit is used to determine, for any target cache unit, the maximum parallelism of each operation among a plurality of homomorphic basic operations using the cache unit in a pipeline calling mode as a segmentation dimension of the cache unit; A second determining unit is used to determine the number of splicing data transmitted by a single HP interface according to the data bit width obtained by the first obtaining unit and the interface bit width obtained by the second obtaining unit, and determine the number of interfaces of the HP interface allocated to the target cache unit for transmitting the splicing data according to the segmentation dimension and the number of splicing obtained by the first determining unit; A configuration unit is used to configure the FPGA according to the number of interfaces obtained by the second determination unit to execute the target fully homomorphic encryption application.
14. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 12.
15. A computing device, comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Data processing method and device based on interface configuration, equipment and medium
CN113885959A
Fully homomorphic encryption neural network reasoning acceleration method and system based on resource reuse
CN116048811A
Processing method and acceleration chip for polynomial data
CN117931130A
Asynchronous clock data rate matching method and device
CN117998144A