Hardware acceleration system based on fully homomorphic encryption and implementation method thereof

By optimizing the fully homomorphic encryption hardware acceleration system through Ethernet transmission and unified computing architecture, the problems of insufficient data transmission efficiency and resource reuse rate are solved, and more efficient computing and more comprehensive functional coverage are achieved.

CN120744995APending Publication Date: 2025-10-03SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510845178.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing fully homomorphic encryption hardware acceleration systems have deficiencies in data transmission efficiency and resource reuse rate, resulting in low computing efficiency, incomplete functions, and difficulty in meeting actual application needs.

Method used

Ethernet is used for data transmission to build a unified computing architecture, including the host, on-chip cache module, computing array, replacement module and data reordering module, to optimize on-chip and off-chip storage management, support a variety of basic homomorphic operations, and have functional scalability.

Benefits of technology

It improves the computing efficiency, resource reuse rate and functional scalability of fully homomorphic encryption, and achieves more efficient data transmission and more complete functional coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744995A_ABST
    Figure CN120744995A_ABST
Patent Text Reader

Abstract

The invention discloses a hardware acceleration system based on fully homomorphic encryption and an implementation method thereof.The hardware acceleration system comprises an upper computer and a hardware accelerator, and the hardware accelerator comprises a host, an off-chip memory, an on-chip cache module, a computing array, a replacement module, a data rearrangement module and a control module; the to-be-processed data, the evaluation key and the twiddle factor are transmitted to the off-chip memory through the Ethernet, and compared with the mode of adopting a URAT serial port transmission mode or directly storing the data in the off-chip memory, the transmission efficiency is higher, and the system function is more complete; a unified calculation array and an optimized on-chip cache module are constructed, so that calculation units are fully fused and reused, on-chip and off-chip storage management is optimized, and the data scheduling efficiency is improved while the on-chip cache requirement is reduced; six types of basic homomorphic operations including homomorphic multiplication, homomorphic rotation and the like can be supported, the function coverage is more comprehensive, and the function expansibility is further achieved on the basis. The method can be widely applied to the technical field of homomorphic encryption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of homomorphic encryption technology, and in particular to a hardware acceleration system based on fully homomorphic encryption and an implementation method thereof. Background Art

[0002] Cloud computing provides convenient, on-demand access to a shared pool of configurable computing resources (such as networks, servers, storage, and applications) over the internet. These resources can be quickly provisioned and released with minimal management effort and interaction with the service provider. However, while cloud computing offers user convenience, it also raises concerns about the security of private data. Homomorphic encryption is a promising solution for protecting sensitive information in cloud computing. Based on mathematical challenges and computational complexity theory, it offers high security and can perform computations directly on encrypted data (ciphertext).

[0003] The concept of homomorphic encryption was first proposed by Rivest and Adleman in 1977, and various fully homomorphic encryption algorithms have subsequently emerged. Currently, the most representative schemes include BGV, BFV, and CKKS, which can be used in a variety of fields, including encrypted keyword retrieval and machine learning. While homomorphic encryption provides strong privacy protection, its implementation is very expensive: computations using homomorphic encryption are several orders of magnitude slower than those on unencrypted data. This is because fully homomorphic encryption schemes are computationally much more complex than traditional encryption algorithms, involve extremely intensive memory accesses, increasing the storage and computational burden, and involve numerous parallel computations, resulting in inefficient software implementations and making them difficult to meet practical application requirements in most scenarios. Therefore, it is necessary to design hardware acceleration systems for homomorphic encryption to improve computational efficiency and make fully homomorphic encryption schemes more efficient in practical deployments.

[0004] When designing a homomorphic encryption hardware acceleration system-level solution, it can often be divided into two main parts: one is the data transmission method between the client that provides computing data and the server of the hardware acceleration circuit, and the other is the hardware acceleration circuit itself.

[0005] Regarding the interaction between the client and server, Sujoy et al. directly stored data in the hardware acceleration circuit's off-chip memory, failing to consider the client's transfer of computational data to the server's off-chip memory. Xie et al. generated ciphertext data and computational keys on the PC, then transferred them via the UART serial port and stored them in off-chip memory. While this further improved the homomorphic encryption hardware acceleration system, the efficiency of data transfer via the UART serial port was limited. When computing large parameter sets ranging from tens to hundreds of MB, data transfer takes up a significant portion of the time, significantly reducing the efficiency of the homomorphic encryption hardware acceleration system.

[0006] In the hardware acceleration circuit, Riazi et al. designed a configurable PE unit for the hardware accelerator's computational module and customized a fixed pipeline architecture for key switching operations, supporting ciphertext-ciphertext and ciphertext-plaintext homomorphic multiplication. However, while this fixed pipeline architecture achieves high computational throughput, it consumes significant resources and results in low resource reuse. Su et al. proposed a hardware accelerator with a reconfigurable multi-core array architecture and a unified computational model. All computational operations are performed by the same PE array, rather than allocating computational resources for individual operations, thereby improving resource reuse. However, the designed hardware acceleration circuit primarily implements homomorphic addition and homomorphic multiplication. Other complex basic operations, such as homomorphic rotation, remain unimplemented, leaving the functionality to be improved. Furthermore, the computational keys and rotation factors required for the computational process are stored in on-chip memory. This on-chip memory overhead affects the choice of key switching algorithm, halving the multiplication depth that can be performed, and limiting its applicability to larger parameter sets.

[0007] The above problems need to be solved urgently. Summary of the Invention

[0008] In order to solve the above technical problems, the purpose of the present invention is: the present invention proposes a hardware acceleration system based on fully homomorphic encryption and its implementation method, which adopts an efficient data transmission method and a unified computing architecture to realize hardware acceleration of homomorphic encryption, thereby improving the computing efficiency, resource reuse rate and functional scalability of fully homomorphic encryption.

[0009] The first technical solution adopted by the present invention is:

[0010] A hardware acceleration system based on fully homomorphic encryption includes a host computer and a hardware accelerator. The hardware accelerator includes a host, an off-chip memory, an on-chip cache module, a computing array, a permutation module, a data rearrangement module, and a control module, wherein:

[0011] The host computer is used to send the data to be processed, the evaluation key and the rotation factor to the off-chip memory for storage via Ethernet, and obtain the operation result returned by the off-chip memory via Ethernet;

[0012] The host is configured to write the data to be processed, the evaluation key, the starting address of the rotation factor in the off-chip memory, and the operation configuration information into the on-chip cache module;

[0013] The control module is configured to read the operation configuration information from the on-chip cache module, schedule the memory access logic of the on-chip cache module according to the operation configuration information, and control the operation logic of the computing array, the permutation module, and the data rearrangement module;

[0014] The on-chip cache module is used to obtain the data to be processed, the evaluation key, and the rotation factor from the off-chip memory according to the starting address for caching, route the cached data to the computing array and the permutation module as needed, and return the calculation result to the off-chip memory after the calculation task is completed;

[0015] The calculation array is used to calculate the input data and transmit the output result to the data rearrangement module;

[0016] The replacement module is used to perform position replacement on different point values ​​of input data based on automorphic mapping, and transmit the output result to the data rearrangement module;

[0017] The data rearrangement module is used to rearrange and dynamically filter the output results of the computing array, obtain the calculation results and return them to the on-chip cache module for caching, and dynamically filter the output results of the replacement module, obtain the calculation results and return them to the on-chip cache module for caching.

[0018] Furthermore, the host computer includes a data endian conversion module, a data sending module and a data receiving module. The data endian conversion module is used to convert the big endian data of the host computer and the little endian data of the hardware accelerator into each other. The data sending module is used to send the little endian data output by the data endian conversion module to the off-chip memory. The data receiving module is used to receive the little endian data returned by the off-chip memory and send it to the data endian conversion module.

[0019] Furthermore, the on-chip cache module includes a parameter cache area and a data cache area. The parameter cache area is used to cache the operation configuration information, and the operation configuration information includes pre-calculated constants, moduli and algorithm parameters. The data cache area is used to cache the data to be processed, the evaluation key, the rotation factor and the operation result.

[0020] Furthermore, the parameter buffer area includes a constant buffer area, a modulus buffer area and an information register stack, the constant buffer area is used to cache the pre-calculated constants, the modulus buffer area is used to cache the modulus, the information register stack is used to cache the algorithm parameters, the data buffer area includes an input data buffer area, an intermediate data buffer area, a result data buffer area and an evaluation key and rotation factor buffer area, the input data buffer area is used to cache the data to be processed, the intermediate data buffer area is used to cache intermediate calculation results, the result data buffer area is used to cache final calculation results, and the evaluation key and rotation factor buffer area is used to cache the evaluation key and the rotation factor.

[0021] Furthermore, the computing array includes multiple parallel configurable computing units, each of which includes a modular adder / subtractor, a modular multiplier / divider, a pipeline register and a multiplexer. The control module configures the computing mode of each computing unit according to the computing configuration information to control the computing logic of the computing array.

[0022] Furthermore, the replacement module is composed of a multi-level interconnected structure, and each level of the structure includes a plurality of two-to-one selectors and corresponding pipeline registers.

[0023] Furthermore, the data rearrangement module includes a global bus and multiple register groups, each of the register groups includes multiple registers, the register groups are used to rearrange the data of the output results of the calculation array, and the global bus is used to dynamically filter the data rearrangement results / the output results of the permutation module, and return the calculation results to the on-chip cache module.

[0024] Furthermore, the control module includes a main state machine and a sub-state machine, wherein the sub-state machine is used for macro-process scheduling and top-level task allocation of homomorphic operations, and the sub-state machine is used for computing operation execution and data transmission control.

[0025] The second technical solution adopted by the present invention is:

[0026] A method for implementing a hardware acceleration system based on fully homomorphic encryption, which is implemented by the above-mentioned hardware acceleration system based on fully homomorphic encryption, includes the following steps:

[0027] Sending the data to be processed, the evaluation key, and the rotation factor to the designated address of the off-chip memory for storage by the host computer;

[0028] Writing the data to be processed, the evaluation key, the starting address of the rotation factor in the off-chip memory, and operation configuration information into the on-chip cache module through the host;

[0029] Reading the operation configuration information from the on-chip cache module through a control module, generating a homomorphic computing task according to the operation configuration information, scheduling the memory access logic of the on-chip cache module according to the homomorphic computing task, and controlling the operation logic of the computing array, the permutation module, and the data rearrangement module;

[0030] Obtaining the data to be processed, the evaluation key, and the rotation factor from the off-chip memory according to the starting address by the on-chip cache module for caching, and routing the cached data to the computing array and the permutation module as needed;

[0031] Calculate the input data through the calculation array and transmit the output result to the data rearrangement module;

[0032] Performing position replacement on different point values ​​of input data based on automorphism mapping by the replacement module, and transmitting the output result to the data rearrangement module;

[0033] The data rearrangement module performs data rearrangement and dynamic screening on the output results of the computing array to obtain a computing result and returns it to the on-chip cache module for caching; and the output results of the replacement module are dynamically screened to obtain a computing result and returns it to the on-chip cache module for caching;

[0034] When the homomorphic computing task is completed, the final calculation result is transmitted back to the off-chip memory through the on-chip cache module, and the final calculation result returned by the off-chip memory is obtained through the host computer.

[0035] The beneficial effects of the present invention are as follows: the present invention provides a hardware acceleration system based on fully homomorphic encryption and its implementation method, including a host computer and a hardware accelerator, the hardware accelerator including a host, an off-chip memory, an on-chip cache module, a computing array, a permutation module, a data rearrangement module and a control module, and transmits the data to be processed, the evaluation key and the rotation factor to the off-chip memory via Ethernet. Compared with adopting the URAT serial port transmission method or directly storing them in the off-chip memory, the transmission efficiency is higher and the system function is more complete; a unified computing array and an optimized on-chip cache module are constructed, so that the computing units are fully integrated and reused, and the on-chip and off-chip storage management is optimized, while reducing the on-chip cache demand, the data scheduling efficiency is improved; six basic homomorphic operations including homomorphic multiplication and homomorphic rotation can be supported, and the functional coverage is more comprehensive, and on this basis, it also has functional scalability, and can realize more complex homomorphic operations. The present invention adopts an efficient data transmission method and a unified computing architecture to realize hardware acceleration of homomorphic encryption, thereby improving the computing efficiency, resource reuse rate and functional scalability of fully homomorphic encryption. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A schematic diagram of the structure of a hardware acceleration system based on fully homomorphic encryption provided by an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of data interaction between a host computer and a hardware accelerator provided in an embodiment of the present invention;

[0038] Figure 3 A schematic structural diagram of an on-chip cache module provided by an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of the structure of a computing unit provided in an embodiment of the present invention;

[0040] Figure 5 A flowchart of the steps of a method for implementing a hardware acceleration system based on fully homomorphic encryption provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.

[0042] In the description of the present invention, "a plurality" means more than two. If a first or second is described, it is only used to distinguish technical features and should not be understood as indicating or implying relative importance, implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used in this specification are only for describing specific embodiments and are not intended to limit the present invention.

[0043] Reference Figure 1 The embodiment of the present invention provides a hardware acceleration system based on fully homomorphic encryption, including a host computer and a hardware accelerator. The hardware accelerator includes a host, an off-chip memory, an on-chip cache module, a computing array, a permutation module, a data rearrangement module, and a control module, wherein:

[0044] The host computer is used to send the data to be processed, the evaluation key, and the rotation factor to the off-chip memory for storage via Ethernet, and to obtain the operation results returned by the off-chip memory via Ethernet;

[0045] The host is used to write the starting address of the data to be processed, the evaluation key and the rotation factor in the off-chip memory and the operation configuration information into the on-chip cache module;

[0046] The control module is used to read the operation configuration information from the on-chip cache module, schedule the memory access logic of the on-chip cache module according to the operation configuration information, and control the operation logic of the calculation array, the permutation module and the data rearrangement module;

[0047] The on-chip cache module is used to obtain the data to be processed, the evaluation key and the rotation factor from the off-chip memory according to the starting address for caching, route the cached data to the calculation array and the permutation module as needed, and return the calculation results to the off-chip memory after the calculation task is completed;

[0048] The calculation array is used to calculate the input data and transmit the output results to the data rearrangement module;

[0049] The permutation module is used to permute the positions of different point values ​​of the input data based on the automorphism mapping, and transmit the output results to the data rearrangement module;

[0050] The data rearrangement module is used to rearrange and dynamically filter the output results of the calculation array, obtain the calculation results and return them to the on-chip cache module for caching, and dynamically filter the output results of the replacement module, obtain the calculation results and return them to the on-chip cache module for caching.

[0051] Specifically, the hardware acceleration system of an embodiment of the present invention can configure the architectural parallelism p, the maximum polynomial degree N, and the maximum multiplication depth L during compilation, where p and N are both powers of 2. The architecture includes a host computer and a hardware accelerator. The hardware accelerator includes a host, an off-chip memory, an on-chip cache module, a computing array, a permutation module, a data rearrangement module, and a control module. Specifically, the host computer and the accelerator establish a TCP Ethernet connection, and the host computer sends all the data required for homomorphic computing to the specified address of the off-chip memory on the accelerator side; the host starts the accelerator by writing the configuration information such as the modulus, pre-calculated constants, algorithm parameters, ciphertext, evaluation key and rotation factor in the starting address of the off-chip memory, as well as the activation information, into the parameter cache area of ​​the on-chip cache module; the control module determines the homomorphic computing task according to the configuration information, schedules the loading of data and configures the computing array, so as to complete the corresponding computing function in an orderly manner; the data in the on-chip cache module is routed to the computing array or permutation module on demand to obtain the calculation result; the above calculation result is dynamically filtered by the data rearrangement module or the global bus and written back to the data cache area of ​​the on-chip cache module; after the calculation task is completed, the calculation result is returned from the on-chip cache module to the off-chip memory, and after the data return task is completed, the host is fed back that the current homomorphic computing task has been completed; the calculation result in the off-chip memory is returned to the host via TCP Ethernet.

[0052] As a further optional implementation, the host computer includes a data endian conversion module, a data sending module and a data receiving module. The data endian conversion module is used to convert the big-endian data of the host computer and the little-endian data of the hardware accelerator. The data sending module is used to send the little-endian data output by the data endian conversion module to the off-chip memory. The data receiving module is used to receive the little-endian data returned by the off-chip memory and send it to the data endian conversion module.

[0053] Specifically, the host computer is the interface between the client (providing ciphertext / plaintext, evaluation key, rotation factor, etc.) and the server (hardware accelerator part), and is used to efficiently transmit the data required for homomorphic encryption calculations to the hardware accelerator part.

[0054] The host computer (TCP Ethernet receiving and transmitting) consists of three parts: a data endian conversion module, a data sending module, and a data receiving module. The data endian conversion module is used to convert the endianness of the host computer's big-endian data (high bit first, low bit last) and the little-endian data (low bit first, high bit last) of the accelerator's off-chip memory, thereby meeting the TCP Ethernet transmission requirements and the off-chip memory storage mode requirements. The data sending module is used to send the data output by the endian conversion module to the accelerator's off-chip memory, and the data receiving module is used to receive data transmitted back to the host computer from the accelerator's off-chip memory. The TCP host computer serves as the interface between the client providing computational data and the server performing homomorphic computation. It is connected to the accelerator's off-chip memory via Ethernet to implement data exchange for homomorphic computation.

[0055] like Figure 2 Figure 1 shows a schematic diagram of data interaction between a host computer and a hardware accelerator according to an embodiment of the present invention. The host computer first establishes a TCP Ethernet connection with the host computer, then sends ciphertext / plaintext data, closing the connection after all data has been sent. During this connection, the off-chip memory receives data, and upon detecting that the host computer connection is closed, the host computer modifies the starting address for data reception. The host computer then establishes a second TCP Ethernet connection, sends evaluation key data, and closes the connection after all data has been sent. During this second connection, the off-chip memory receives data from a new starting address, and upon detecting that the host computer connection is closed, the host computer again modifies the starting address for data reception. The host computer then establishes a third TCP Ethernet connection, sends twiddle factor data, and closes the connection after all data has been sent. During the third connection, the off-chip memory receives data from yet another new starting address, and upon detecting that the host computer connection is closed, the host computer modifies the starting address for data reception back to the original one. Finally, the host computer establishes a final TCP Ethernet connection, waits for data reception, and closes the connection after all data has been received. After all homomorphic computing tasks have been completed and the host computer has established a TCP Ethernet connection, the host computer sends data back to the host computer. It should be noted that the above description is only one transmission method; other methods are also possible. For example, a connection can be established only once, with the host computer sending the ciphertext data, evaluation key, and rotation factors all at once. The host can then count the amount of data to determine the starting address of each data type in the off-chip memory.

[0056] Reference Figure 3 As an optional implementation, the on-chip cache module includes a parameter cache area and a data cache area. The parameter cache area is used to cache operation configuration information, and the operation configuration information includes pre-calculated constants, moduli, and algorithm parameters. The data cache area is used to cache data to be processed, evaluation keys, rotation factors, and operation results.

[0057] Specifically, the on-chip cache module is divided into a parameter cache and a data cache. The parameter cache is used to write information such as pre-computed constants, moduli, and algorithm parameter configurations. The data cache is responsible for caching the ciphertext / plaintext required for homomorphic computation input, intermediate results / final ciphertext calculation results during the computation process, evaluation keys, and rotation factors. The computation array is connected to the off-chip memory through the on-chip cache module, which is used to facilitate data exchange between the computation array and the off-chip memory.

[0058] Reference Figure 3 As an optional implementation, the parameter buffer includes a constant buffer, a modulus buffer, and an information register stack. The constant buffer is used to cache pre-calculated constants, the modulus buffer is used to cache moduli, and the information register stack is used to cache algorithm parameters. The data buffer includes an input data buffer, an intermediate data buffer, a result data buffer, and an evaluation key and rotation factor buffer. The input data buffer is used to cache data to be processed, the intermediate data buffer is used to cache intermediate calculation results, the result data buffer is used to cache final calculation results, and the evaluation key and rotation factor buffer is used to cache evaluation keys and rotation factors.

[0059] Specifically, the parameter buffer is used by the host to write pre-calculated constants, moduli, and algorithm parameter configuration information into the constant buffer, modulus buffer, and information register stack respectively. The data buffer is responsible for caching the ciphertext / plaintext required for the input homomorphic calculation, the intermediate results / final ciphertext calculation results during the calculation process, partial evaluation keys, and partial rotation factors into the input data cache, intermediate data cache, result data cache, evaluation key, and rotation factor cache respectively. In the data cache. The number of groups in each cache part is p. The computing array is connected to the off-chip memory through the on-chip cache module, and the on-chip cache module is used to realize data interaction between the computing array and the off-chip memory.

[0060] The modulus buffer is used to cache the expanded moduli {p0,···,pα-1} and the input ciphertext moduli {q0,···,qL}, with a total capacity of MODnum. The constant buffer is used to cache constant terms involved in the key switching and rescaling sub-operations during homomorphic computation. Each modulus has (MODnum+3) constant terms, and the buffer depth is (MODnum+3)×MODnum. The information register file contains multiple registers for caching the starting addresses of various data in off-chip memory, a register for caching real-time algorithm parameters, and a register for controlling the accelerator state.

[0061] The input data cache and result data cache support caching of up to 2×MODnum and 2×(L+1) remainder polynomials, respectively, and both support concurrent read and write of 2p data at a time. The intermediate data cache is primarily used to cache intermediate computation results, but can also be used to cache an input ciphertext / plaintext polynomial. Its capacity is N, and it supports concurrent read and write of 2p data at a time. The evaluation key and twiddle factor cache has a capacity of 4N+32p, supports concurrent read and write of 2p data, and can simultaneously accommodate two remainder polynomials and two sets of twiddle factors.

[0062] Reference Figure 4 As an optional implementation, the computing array includes multiple parallel configurable computing units, each of which includes a modular adder / subtractor, a modular multiplier / divider, a pipeline register, and a multiplexer. The control module configures the computing mode of each computing unit according to the computing configuration information to control the computing logic of the computing array.

[0063] Specifically, the computational array includes p parallel, configurable, and reconfigurable computational units. Each computational unit includes a modular adder / subtractor, a modular multiplier / divider, multiple pipeline registers, and a multiplexer. It supports GS and CT butterfly operations, thereby implementing modular addition / subtraction, modular multiplication / modular multiplication-addition, and NTT / INTT operations.

[0064] by Figure 4 Taking the calculation unit shown in the figure as an example, when the calculation mode is configured to 0, when data enters from inputs a1, b1, and c, respectively, outputs 1 and 2 are calculated as a1+b1*c and a1-b1*c, respectively, enabling modular addition / subtraction, modular multiplication, and modular multiplication-addition operations. Specifically, when input c is a twiddle factor, an NTT operation is implemented. When the calculation mode is configured to 1, when data enters from inputs a2, b2, and c, respectively, outputs 1 and 2 are calculated as (a2+b2) / 2 and (b2-a2)*c / 2, respectively. Specifically, when input c is a twiddle factor, an INTT operation is implemented.

[0065] As a further optional implementation, the replacement module is composed of a multi-level interconnected structure, and each level of the structure includes multiple two-to-one selectors and corresponding pipeline registers.

[0066] Specifically, it consists of a multi-level interconnect structure, each containing multiple binary selectors and corresponding pipeline registers. The permutation module is specifically designed to permute the positions of different point values ​​when implementing automorphic mapping. Its input is connected to the on-chip cache module, and its output is connected to the global bus of the data rearrangement module. Data is read from the data buffer of the on-chip cache, permuted, and then written to the other buffers of the data buffer of the on-chip cache via the global bus.

[0067] In an embodiment of the present invention, the permutation module can be designed to have a size of p × log2 p, consisting of a log2 p-level interconnect structure, with each level containing p MUX2s and associated pipeline registers. The permutation module supports parallel execution of p point value permutations per cycle.

[0068] As a further optional implementation, the data rearrangement module includes a global bus and multiple register groups, each register group includes multiple registers, the register group is used to rearrange the output results of the calculation array, and the global bus is used to dynamically filter the data rearrangement results / output results of the permutation module and return the calculation results to the on-chip cache module.

[0069] Specifically, the data rearrangement module consists of a global bus and multiple register groups, each of which contains multiple registers. The computation array is connected to the on-chip cache module through the data rearrangement module. The data rearrangement module is used to organize and rearrange the output results of the computation array, and then filter them through the global bus to generate data write-back data for the on-chip cache module and transmit them to the various cache areas of the on-chip cache module. The data rearrangement module can also replace the module results and feed them into the global bus, thereby realizing the filtering and write-back of the calculation results.

[0070] As a further optional implementation, the control module includes a main state machine and a sub-state machine, the sub-state machine is used for macro process scheduling and top-level task allocation of homomorphic operations, and the sub-state machine is used for computing operation execution and data transmission control.

[0071] Specifically, the on-chip cache module, computing array, data rearrangement module, and permutation module are all connected to the control module. The control module is used to read the configuration information sent by the host from the parameter buffer area of ​​the on-chip cache area. Based on this configuration information, it controls the operation logic of the computing array, data rearrangement module, and permutation module for the current homomorphic computing task, and schedules the memory access logic of the on-chip cache module.

[0072] The control module can adopt a two-level scheduling control mechanism: the main state machine is responsible for the macro-process scheduling and top-level task allocation of each homomorphic operation; the sub-state machine implements fine-grained task scheduling for each main state, including executing computing operations and data transmission control.

[0073] Specifically, the main state machine can control the scheduling of macro tasks such as ciphertext loading to off-chip memory, ciphertext-ciphertext modular multiplication, key switching, rescaling, automorphic mapping, and ciphertext result data returned to off-chip memory; the sub-state machine can perform specific ciphertext loading, partial evaluation key and rotation factor loading, polynomial point value modular multiplication and addition calculations, scalar-polynomial point value modular multiplication calculations, polynomial NTT / INTT calculations, fast basis transformation calculations, polynomial modular reduction calculations, polynomial point value modular subtraction and multiplication calculations, polynomial point value modular addition calculations, polynomial automorphism mapping calculations, and specific result ciphertext return tasks.

[0074] The above is an explanation of the structure and working principle of the digital converter of the embodiment of the present invention. It can be recognized that the embodiment of the present invention transmits the data to be processed, the evaluation key and the rotation factor to the off-chip memory via Ethernet. Compared with the URAT serial port transmission method or direct storage in the off-chip memory, the transmission efficiency is higher and the system function is more complete; a unified computing array and an optimized on-chip cache module are constructed to fully integrate and reuse the computing units, and optimize the on-chip and off-chip storage management, thereby reducing the on-chip cache requirements and improving the data scheduling efficiency; it can support six types of basic homomorphic operations including homomorphic multiplication and homomorphic rotation, with more comprehensive functional coverage, and on this basis, it also has functional scalability, which can realize more complex homomorphic operations. The present invention adopts an efficient data transmission method and a unified computing architecture to realize hardware acceleration of homomorphic encryption, thereby improving the computing efficiency, resource reuse rate and functional scalability of full homomorphic encryption.

[0075] Reference Figure 5 , an embodiment of the present invention provides a method for implementing a hardware acceleration system based on fully homomorphic encryption, which is implemented by the above-mentioned hardware acceleration system based on fully homomorphic encryption, including the following steps:

[0076] S101, sending the data to be processed, the evaluation key, and the rotation factor to a designated address of an off-chip memory for storage via a host computer;

[0077] S102, writing the data to be processed, the evaluation key and the starting address of the rotation factor in the off-chip memory and the operation configuration information to the on-chip cache module through the host;

[0078] S103, reading operation configuration information from the on-chip cache module through the control module, generating a homomorphic computing task according to the operation configuration information, scheduling the memory access logic of the on-chip cache module according to the homomorphic computing task, and controlling the operation logic of the computing array, the permutation module, and the data rearrangement module;

[0079] S104, obtaining the data to be processed, the evaluation key, and the rotation factor from the off-chip memory according to the starting address through the on-chip cache module, and routing the cached data to the calculation array and the permutation module as needed;

[0080] S105, calculating the input data through the calculation array, and transmitting the output result to the data rearrangement module;

[0081] S106, performing position replacement on different point values ​​of the input data based on the automorphism mapping by the replacement module, and transmitting the output result to the data rearrangement module;

[0082] S107, performing data rearrangement and dynamic screening on the output results of the computing array by the data rearrangement module to obtain a computing result and returning it to the on-chip cache module for caching; and dynamically screening the output results of the replacement module to obtain a computing result and returning it to the on-chip cache module for caching;

[0083] S108. When the homomorphic computing task is completed, the final calculation result is transmitted back to the off-chip memory through the on-chip cache module, and the final calculation result returned by the off-chip memory is obtained through the host computer.

[0084] Specifically, a TCP Ethernet connection is established between the host computer and the accelerator, and the host computer sends the ciphertext, evaluation key, and rotation factor to the specified address of the off-chip memory on the accelerator side; the host starts the accelerator by writing configuration information such as the modulus, pre-calculated constants, algorithm parameters, ciphertext, evaluation key, and rotation factor in the starting address of the off-chip memory, as well as activation information, into the parameter cache area of ​​the on-chip cache module; the control module determines the homomorphic computing task based on the configuration information, issues the corresponding data loading task, and loads the input ciphertext / plaintext polynomial from the off-chip memory to the data cache area of ​​the on-chip cache module; the data in the on-chip cache module is routed to the computing array or permutation module on demand to obtain the calculation result; the above calculation result is dynamically filtered by the data rearrangement module or the global bus and then written back to the data cache area of ​​the on-chip cache module; after the calculation task is completed, the calculation result is returned from the on-chip cache module to the off-chip memory, and after the data return task is completed, feedback is fed back to the host that the current homomorphic computing task has been completed; the calculation result in the off-chip memory is returned to the host computer via TCP Ethernet.

[0085] It can be understood that the contents of the above system embodiments are applicable to the present method embodiments, the functions specifically implemented by the present method embodiments are the same as those of the above system embodiments, and the beneficial effects achieved are also the same as those achieved by the above system embodiments.

[0086] It should be appreciated that embodiments of the present invention can be implemented or practiced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The above methods can be implemented in a computer program using standard programming techniques—including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner—according to the methods and figures described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed application-specific integrated circuit for this purpose.

[0087] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer programs described above include a plurality of instructions that may be executed by one or more processors.

[0088] Furthermore, the above methods can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques described herein, the present invention also includes the computer itself.

[0089] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.

[0090] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods are possible.

Claims

1. A hardware acceleration system based on fully homomorphic encryption, characterized in that: It includes a host computer and a hardware accelerator, wherein the hardware accelerator includes a host, an off-chip memory, an on-chip cache module, a computing array, a replacement module, a data rearrangement module, and a control module, wherein: The host computer is used to send the data to be processed, the evaluation key and the rotation factor to the off-chip memory for storage via Ethernet, and obtain the operation result returned by the off-chip memory via Ethernet; The host is configured to write the data to be processed, the evaluation key, the starting address of the rotation factor in the off-chip memory, and the operation configuration information into the on-chip cache module; The control module is configured to read the operation configuration information from the on-chip cache module, schedule the memory access logic of the on-chip cache module according to the operation configuration information, and control the operation logic of the computing array, the permutation module, and the data rearrangement module; The on-chip cache module is used to obtain the data to be processed, the evaluation key, and the rotation factor from the off-chip memory according to the starting address for caching, route the cached data to the computing array and the permutation module as needed, and return the calculation result to the off-chip memory after the calculation task is completed; The calculation array is used to calculate the input data and transmit the output result to the data rearrangement module; The replacement module is used to perform position replacement on different point values ​​of input data based on automorphic mapping, and transmit the output result to the data rearrangement module; The data rearrangement module is used to rearrange and dynamically filter the output results of the computing array, obtain the calculation results and return them to the on-chip cache module for caching, and dynamically filter the output results of the replacement module, obtain the calculation results and return them to the on-chip cache module for caching.

2. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The host computer includes a data endian conversion module, a data sending module and a data receiving module. The data endian conversion module is used to convert the big-endian data of the host computer and the little-endian data of the hardware accelerator. The data sending module is used to send the little-endian data output by the data endian conversion module to the off-chip memory. The data receiving module is used to receive the little-endian data returned by the off-chip memory and send it to the data endian conversion module.

3. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The on-chip cache module includes a parameter cache area and a data cache area. The parameter cache area is used to cache the operation configuration information, and the operation configuration information includes pre-calculated constants, moduli and algorithm parameters. The data cache area is used to cache the data to be processed, the evaluation key, the rotation factor and the operation result.

4. The hardware acceleration system based on fully homomorphic encryption according to claim 3, characterized in that: The parameter buffer area includes a constant buffer area, a modulus buffer area and an information register stack. The constant buffer area is used to cache the pre-calculated constants, the modulus buffer area is used to cache the modulus, and the information register stack is used to cache the algorithm parameters. The data buffer area includes an input data buffer area, an intermediate data buffer area, a result data buffer area, and an evaluation key and rotation factor buffer area. The input data buffer area is used to cache the data to be processed, the intermediate data buffer area is used to cache intermediate calculation results, the result data buffer area is used to cache final calculation results, and the evaluation key and rotation factor buffer area is used to cache the evaluation key and the rotation factor.

5. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The computing array includes multiple parallel configurable computing units, each of which includes a modular adder / subtractor, a modular multiplier / divider, a pipeline register and a multiplexer. The control module configures the computing mode of each computing unit according to the computing configuration information to control the computing logic of the computing array.

6. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The replacement module is composed of a multi-level interconnected structure, and each level of the structure includes a plurality of two-to-one selectors and corresponding pipeline registers.

7. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The data rearrangement module includes a global bus and multiple register groups, each of which includes multiple registers. The register groups are used to rearrange the output results of the calculation array. The global bus is used to dynamically filter the data rearrangement results / output results of the permutation module and return the calculation results to the on-chip cache module.

8. The hardware acceleration system based on fully homomorphic encryption according to claim 1, characterized in that: The control module includes a main state machine and a sub-state machine, wherein the sub-state machine is used for macro process scheduling and top-level task allocation of homomorphic operations, and the sub-state machine is used for computing operation execution and data transmission control.

9. A method for implementing a hardware acceleration system based on fully homomorphic encryption, for implementing the hardware acceleration system based on fully homomorphic encryption according to any one of claims 1 to 8, characterized in that: The following steps are involved: Sending the data to be processed, the evaluation key, and the rotation factor to the designated address of the off-chip memory for storage by the host computer; Writing the data to be processed, the evaluation key, the starting address of the rotation factor in the off-chip memory, and operation configuration information into the on-chip cache module through the host; Reading the operation configuration information from the on-chip cache module through a control module, generating a homomorphic computing task according to the operation configuration information, scheduling the memory access logic of the on-chip cache module according to the homomorphic computing task, and controlling the operation logic of the computing array, the permutation module, and the data rearrangement module; Obtaining the data to be processed, the evaluation key, and the rotation factor from the off-chip memory according to the starting address by the on-chip cache module for caching, and routing the cached data to the computing array and the permutation module as needed; Calculate the input data through the calculation array and transmit the output result to the data rearrangement module; Performing position replacement on different point values ​​of input data based on automorphism mapping by the replacement module, and transmitting the output result to the data rearrangement module; The data rearrangement module performs data rearrangement and dynamic screening on the output results of the computing array to obtain a computing result and returns it to the on-chip cache module for caching; and the output results of the replacement module are dynamically screened to obtain a computing result and returns it to the on-chip cache module for caching; When the homomorphic computing task is completed, the final calculation result is transmitted back to the off-chip memory through the on-chip cache module, and the final calculation result returned by the off-chip memory is obtained through the host computer.

Citation Information

Cited By

  • Shift register array-based TFHE fully homomorphic operation bootstrap acceleration method

    CN122293305A