A multi-modal scalable high-performance homomorphic encryption system

Through software and hardware collaborative design, the memory scheduling and computing modules of the homomorphic encryption system are optimized, which solves the scheduling inefficiency problem of the multimodal homomorphic encryption system, achieves efficient ciphertext data processing and computing performance improvement, and adapts to application needs of different scales.

CN119814276BActive Publication Date: 2025-10-17SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510026429.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-10-17
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing homomorphic encryption systems suffer from inefficient scheduling in their hardware architecture design, especially in computing clusters of different sizes and multimodal homomorphic encryption application scenarios, which limits the reading and writing efficiency and computing efficiency of ciphertext data.

Method used

By adopting the hardware-software co-design method, we can achieve efficient management and parallel processing of multimodal homomorphic encryption algorithms by optimizing the scheduling strategy at the software level and introducing high-bandwidth memory, fast number theory transformation computing units and scalable computing modules at the hardware level.

Benefits of technology

It significantly improves the computing efficiency and performance of homomorphic encryption systems, solves the performance bottleneck caused by inefficient scheduling in traditional architectures, supports application scenarios of different scales, and improves the flexibility and compatibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814276B_ABST
    Figure CN119814276B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal scalable high-performance homomorphic encryption system, and relates to the technical field of computer systems, which comprises an upper-layer server software and a driving architecture in communication connection and a multi-modal homomorphic encryption hardware architecture realized based on FPGA; the upper-layer server software and the driving architecture comprise a multi-modal homomorphic encryption algorithm software library, a software memory scheduler, a modulo calculation module software driver and a floating-point calculation module software driver; the multi-modal homomorphic encryption hardware architecture realized based on FPGA comprises a communication data controller, a plurality of modulo calculation modules, a floating-point calculation module, a data read-write bus and a high-bandwidth memory; each modulo calculation module comprises a fast number theory transformation calculation unit. Through the design scheme of software and hardware cooperation, the application solves the problem of low scheduling efficiency of different-scale ciphertext calculation tasks under multi-modal homomorphic encryption in the design of hardware architecture, and significantly improves the calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer systems, and in particular to a multi-modal scalable high-performance homomorphic encryption system. BACKGROUND

[0002] Homomorphic encryption is an encryption technology that allows direct computation operations to be performed on ciphertext without decryption. Its main advantage is that it can process data while ensuring data privacy. This feature makes homomorphic encryption highly concerned in many application scenarios, such as data analysis in cloud computing, privacy protection of medical data, and secure computation of financial data. However, homomorphic encryption also has some disadvantages, including high computational complexity, large performance overhead, and the challenge of scheduling multiple homomorphic encryption schemes (multi-modal homomorphic encryption) in practical applications. SUMMARY

[0003] The main purpose of the embodiments of the present application is to propose a multi-modal scalable high-performance homomorphic encryption system to improve the operation efficiency of homomorphic encryption.

[0004] To achieve the above purpose, the embodiments of the present application propose a multi-modal scalable high-performance homomorphic encryption system, which comprises an upper-layer server software and driving architecture and a multi-modal homomorphic encryption hardware architecture realized based on FPGA;

[0005] The upper-layer server software and driving architecture and the multi-modal homomorphic encryption hardware architecture realized based on FPGA are in communication connection;

[0006] The upper-layer server software and driving architecture comprises a multi-modal homomorphic encryption algorithm software library, a software memory scheduler, a modulo calculation module software driver, and a floating-point calculation module software driver;

[0007] The multi-modal homomorphic encryption hardware architecture realized based on FPGA comprises a communication data controller, a plurality of modulo calculation modules, a floating-point calculation module, a data read-write bus, and a high-bandwidth memory; each modulo calculation module comprises a fast number-theoretic transform calculation unit;

[0008] The communication data controller is in communication connection with the upper-layer server software and driving architecture, and is also in communication connection with each modulo calculation module and the floating-point calculation module; the data read-write bus is in communication connection with the communication data controller, each modulo calculation module, the floating-point calculation module, and the high-bandwidth memory.

[0009] In some embodiments, the multi-modal homomorphic encryption algorithm software library is used to manage and schedule multi-modal homomorphic encryption algorithms;

[0010] The software memory scheduler is used for managing and scheduling memory resources;

[0011] The modulo calculation module software driver is used for controlling and driving each of the modulo calculation modules;

[0012] The floating-point calculation module software driver is used for controlling and driving the floating-point calculation module.

[0013] In some embodiments, the communication data controller is used for managing interaction data between the upper-layer server software and driver architecture and the FPGA- implemented multi-modal homomorphic encryption hardware architecture;

[0014] Each of the modulo calculation modules is used for modulo operation;

[0015] The floating-point calculation module is used for floating-point number operation;

[0016] The data read-write bus is used for transmitting data;

[0017] The high-bandwidth memory is used for storing ciphertext data and intermediate calculation results.

[0018] In some embodiments, the upper-layer server software and driver architecture are communicatively connected with the FPGA- implemented multi-modal homomorphic encryption hardware architecture through a PCIe data communication interface;

[0019] The communication data controller adopts a PCIe data controller;

[0020] The PCIe data controller is communicatively connected with each of the modulo calculation modules and the floating-point calculation module through a calculation module control signaling data path.

[0021] In some embodiments, each of the modulo calculation modules and the floating-point calculation module includes a memory array composed of a plurality of SRAM memory blocks, a data multiplexer, a memory block read-write address generator, a configurable parameter register, and a memory data read-write path;

[0022] Each of the SRAM memory blocks is used for storing data;

[0023] The data multiplexer includes a data write path and a data read path, and is used for reading and writing distributed data;

[0024] The memory block read-write address generator is used for generating read-write addresses of each of the SRAM memory blocks;

[0025] The configurable parameter register is used for storing configuration parameters, which are configured by the upper-layer server software and the driving architecture through a calculation module control signaling data channel.

[0026] The memory data read-write channel is connected with the high-bandwidth memory, and is used for reading and writing data.

[0027] In some embodiments, each SRAM memory block is used for storing intermediate data.

[0028] Each SRAM memory block includes a plurality of SRAM memory units, and each SRAM memory block includes a 1024-bit wide read-write interface.

[0029] Each SRAM memory unit has a storage capacity of 4096 words of data, and each word of data is 64 bits; and a single SRAM memory block has a storage capacity of 65536 words of data.

[0030] In some embodiments, the memory array is provided with a data read-write channel A and a data read-write channel B.

[0031] The data read-write channel A is connected to each modulo calculation module, and the data read-write channel B is connected to the floating-point calculation module.

[0032] In some embodiments, the fast number theoretic transform calculation unit is used for executing a homomorphic encryption algorithm through a plurality of parallel operations of the memory array.

[0033] In some embodiments, the upper-layer server software and the driving architecture are used for merging a plurality of physical blocks in the high-bandwidth memory into a virtual block, thereby obtaining a plurality of virtual blocks; and each virtual block corresponds to a modulo calculation module.

[0034] Each virtual block is internally divided into a plurality of virtual sub-blocks, and each virtual sub-block is used for storing polynomial coefficients of different RNS bases.

[0035] In some embodiments, the multi-modal homomorphic encryption algorithm software library includes a CKKS homomorphic encryption algorithm, a BFV homomorphic encryption algorithm and a BGV homomorphic encryption algorithm.

[0036] The embodiments of the present application at least have the following beneficial effects:

[0037] The system of the application comprises an upper server software and driving architecture and a multi-modal homomorphic encryption hardware architecture realized based on FPGA; the upper server software and driving architecture comprises a multi-modal homomorphic encryption algorithm software library, a software memory scheduler, a modulo calculation module software driver and a floating point calculation module software driver; the multi-modal homomorphic encryption hardware architecture realized based on FPGA comprises a communication data controller, a plurality of modulo calculation modules, a floating point calculation module, a data read-write bus and a high-bandwidth memory; each modulo calculation module comprises a fast number theoretic transform calculation unit. Through the design scheme of software and hardware cooperation, the application solves the problem of low efficiency in scheduling of different scale ciphertext calculation tasks under multi-modal homomorphic encryption in the design of hardware architecture, and significantly improves the calculation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0039] Figure 1 An example structure diagram of a multi-modal scalable high-performance homomorphic encryption system provided by the embodiments of the application is shown in the figure.

[0040] Figure 2 A high-speed SRAM memory block structure design schematic diagram provided by the embodiments of the application is shown in the figure.

[0041] Figure 3 A data storage and transmission architecture schematic diagram based on a high-speed SRAM memory block provided by the embodiments of the application is shown in the figure.

[0042] Figure 4 A software and hardware cooperative memory distribution example diagram provided by the embodiments of the application is shown in the figure.

[0043] Figure 5 A pipeline scheduling flowchart of NTT calculation in a modulo calculation unit provided by the embodiments of the application is shown in the figure. DETAILED DESCRIPTION

[0044] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0045] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "when" or "in response to determining".

[0046] The terms "at least one", "multiple", "each", "any" and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0048] Before the embodiments of the present application are described in detail, first, some related technologies that can be involved in the embodiments of the present application are described as follows:

[0049] With the development of homomorphic encryption technology, the processing capability of ciphertext data has been significantly improved. However, the existing homomorphic encryption system still faces many challenges in the implementation of hardware architecture, especially in different scale computing clusters and different modal homomorphic encryption application scenarios. High-performance memory scheduling strategies are key factors affecting the overall performance of the system.

[0050] The traditional ciphertext data read-write strategy, due to the lack of support for multi-modal homomorphic encryption and optimization of ciphertext data of different scales, cannot fully utilize the memory bandwidth, limiting the read-write efficiency of ciphertext data in the system. This limitation further affects the computing efficiency of the ciphertext, becoming a bottleneck for system performance optimization.

[0051] In order to improve the system performance, how to effectively schedule and parallelize the ciphertext data in the memory becomes a problem to be solved. By optimizing the memory scheduling strategy and designing the hardware computing module to be scalable, the data read-write efficiency can be improved and the overall computing capacity of the system can be enhanced. This not only helps to improve the performance of the encryption system, but also enables the deployment of hardware platforms according to different application scenarios.

[0052] In order to solve the problem of low scheduling efficiency of homomorphic encryption ciphertext in the hardware architecture design in the prior art, the present application adopts a software and hardware collaborative design methodology, which integrates scheduling optimization at the software level and efficient computation at the hardware level to comprehensively improve system performance.

[0053] The upper layer software implements a multi-modal homomorphic encryption algorithm library, which calls the underlying hardware resources through the hardware calling interface function in the algorithm library to realize the acceleration of the multi-modal homomorphic encryption upper application; the upper layer software optimizes the scheduling strategy of the ciphertext data by scheduling algorithm programs, adopts an RNS (Residue Number System) based memory scheduling algorithm, decomposes large integers into small moduli for calculation, allocates ciphertext storage space and computing tasks according to the available computing resources of the hardware resources, thereby improving the parallelism and efficiency of data processing. The memory bandwidth utilization and data processing efficiency are improved.

[0054] The underlying hardware integrates an NTT (Number Theoretic Transform) computing unit for efficiently processing large integer polynomial multiplication operations required by homomorphic encryption, improving computing speed and accuracy; the underlying hardware introduces HBM (High Bandwidth Memory) memory and high-speed on-chip SRAM (Static Random-Access Memory) memory, which utilizes the multi-channel high bandwidth of HBM and the high read-write characteristics of SRAM to provide high-throughput data transmission capability for the system, and in combination with the optimized scheduling algorithm, fully utilizes the memory bandwidth to realize efficient parallel processing of ciphertext data. The underlying memory scheduling circuit adopts a dynamic scheduling strategy to adjust the memory read-write order in real time according to the size of the ciphertext data, ensuring the continuity and efficiency of data transmission; the underlying hardware introduces a high-performance double-precision floating-point processing module to provide hardware support for the underlying operators of the multi-modal homomorphic encryption scheme. The computing module is scalable and can allocate different numbers of computing units according to different application scenarios, fully compatible with homomorphic encryption applications in different scenarios.

[0055] Through the above technical means and measures, the scheduling efficiency and computing performance of the homomorphic encryption hardware architecture are significantly improved, and the performance bottleneck problem caused by low scheduling efficiency in the traditional architecture is solved.

[0056] The embodiment of the application provides a multi-modal scalable high-performance homomorphic encryption system, which comprises an upper-layer server software and driving architecture and a multi-modal homomorphic encryption hardware architecture realized based on FPGA;

[0057] The upper-layer server software and driving architecture and the multi-modal homomorphic encryption hardware architecture realized based on FPGA are in communication connection;

[0058] The upper-layer server software and driving architecture comprises a multi-modal homomorphic encryption algorithm software library, a software memory scheduler, a modulo calculation module software driver and a floating-point calculation module software driver;

[0059] The multi-modal homomorphic encryption hardware architecture realized based on FPGA comprises a communication data controller, a plurality of modulo calculation modules, a floating-point calculation module, a data read-write bus and a high-bandwidth memory; each modulo calculation module comprises a fast number-theoretic transform calculation unit;

[0060] The communication data controller is in communication connection with the upper-layer server software and driving architecture, and is also in communication connection with each modulo calculation module and the floating-point calculation module respectively; the data read-write bus is in communication connection with the communication data controller, each modulo calculation module, the floating-point calculation module and the high-bandwidth memory respectively.

[0061] Further, the multi-modal homomorphic encryption algorithm software library is used for managing and scheduling multi-modal homomorphic encryption algorithms;

[0062] The software memory scheduler is used for managing and scheduling memory resources;

[0063] The modulo calculation module software driver is used for controlling and driving each modulo calculation module;

[0064] The floating-point calculation module software driver is used for controlling and driving the floating-point calculation module.

[0065] Further, the communication data controller is used for managing interactive data between the upper-layer server software and driving architecture and the multi-modal homomorphic encryption hardware architecture realized based on FPGA;

[0066] Each modulo calculation module is used for modulo operation;

[0067] The floating-point calculation module is used for floating-point number operation;

[0068] The data read-write bus is used for transmitting data;

[0069] The high-bandwidth memory is used for storing ciphertext data and intermediate calculation results.

[0070] Further, the upper-layer server software and driver architecture is in communication connection with the FPGA-based multi-modal homomorphic encryption hardware architecture through a PCIe data communication interface;

[0071] The communication data controller adopts a PCIe data controller;

[0072] The PCIe data controller is in communication connection with each of the modulo calculation module and the floating-point calculation module through a calculation module control signaling data path.

[0073] Further, each of the modulo calculation module and the floating-point calculation module includes a memory array composed of a plurality of SRAM memory blocks, a data multiplexer, a memory block read-write address generator, a configurable parameter register, and a memory data read-write path;

[0074] Each of the SRAM memory blocks is configured to store data.

[0075] The data multiplexer includes a data write path and a data read path, and is configured to read and write distributed data.

[0076] The memory block read-write address generator is configured to generate read-write addresses of each of the SRAM memory blocks.

[0077] The configurable parameter register is configured to store configuration parameters, which are configured by the upper-layer server software and driver architecture through a calculation module control signaling data path.

[0078] The memory data read-write path is connected with the high-bandwidth memory, and is configured to read and write data.

[0079] As a more specific embodiment, each of the SRAM memory blocks is configured to store intermediate data.

[0080] Each of the SRAM memory blocks includes a plurality of SRAM memory cells; and each of the SRAM memory blocks includes a 1024-bit wide read-write interface.

[0081] Each of the SRAM memory cells has a storage capacity of 4096 word data, each word data being 64 bits; and a single SRAM memory block has a storage capacity of 65536 word data.

[0082] As a more specific embodiment, the memory array is provided with a data read-write channel A and a data read-write channel B.

[0083] The data read-write channel A is connected to each of the modulo calculation modules, and the data read-write channel B is connected to the floating point calculation module.

[0084] As a more specific embodiment, the fast number theory transform calculation unit is used to execute a homomorphic encryption algorithm through a plurality of parallel operations of the memory array.

[0085] Further, the upper-layer server software and driver architecture are used to combine a plurality of physical blocks in the high-bandwidth memory into a virtual block, thereby obtaining a plurality of virtual blocks; each of the virtual blocks corresponds to one of the modulo calculation modules.

[0086] Each of the virtual blocks is internally divided into a plurality of virtual sub-blocks, and each of the virtual sub-blocks is used to store polynomial coefficients of different RNS bases.

[0087] Further, the multi-modal homomorphic encryption algorithm software library includes CKKS homomorphic encryption algorithms, BFV homomorphic encryption algorithms, and BGV homomorphic encryption algorithms.

[0088] Next, the scheme of the embodiments of the present application will be described in detail with reference to specific application examples.

[0089] Referring to Figure 1 , Figure 1 is a structure diagram of a multi-modal homomorphic encryption calculation system implemented based on FPGA (Field Programmable Gate Array) provided by the embodiments of the present application. The system includes upper-layer server software and driver architecture and multi-modal homomorphic encryption hardware architecture implemented based on FPGA. The upper-layer server software and driver architecture are connected to the multi-modal homomorphic encryption hardware architecture implemented based on FPGA through a PCIe (Peripheral Component Interconnect Express) data communication interface.

[0090] The upper-layer server software and driver architecture include a multi-modal homomorphic encryption algorithm software library, a software memory scheduler, and modulo and floating point calculation module software drivers.

[0091] Specifically, the multi-modal homomorphic encryption algorithm software library is used to manage and schedule multi-modal homomorphic encryption algorithms. The software memory scheduler is responsible for managing and scheduling system memory resources. The modulo and floating point calculation module software drivers are used to control and drive corresponding hardware modules. The PCIe data communication interface is used for high-speed data transmission between the upper-layer server and the FPGA hardware.

[0092] Specifically, the system supports multiple homomorphic encryption algorithms, mainly including CKKS (Cheon-Kim-Kim-Song), BFV (Brakerski / Fan-Vercauteren) and BGV (Brakerski-Gentry-Vaikuntanathan) mainstream homomorphic encryption algorithms. The CKKS algorithm is mainly used for approximate number calculation and supports floating-point number operation; the BFV algorithm is suitable for integer operation and can keep the size of the ciphertext unchanged during the process of ciphertext calculation, having high calculation efficiency; the BGV algorithm supports circuit evaluation of any depth and is suitable for complex encryption calculation tasks. The multi-modal homomorphic encryption algorithm software library manages these different algorithms, so that the system can select the most suitable encryption scheme according to the specific application requirements, thereby realizing efficient and flexible homomorphic encryption calculation.

[0093] The multi-modal homomorphic encryption hardware architecture implemented based on the FPGA includes a PCIe data controller, a calculation module control signaling data path, a modulo calculation module, a high-performance double-precision floating-point calculation module, a data read-write bus and a high-bandwidth HBM memory.

[0094] The PCIe data controller is responsible for managing the data transmission of the PCIe interface. The calculation module control signaling data path is used to coordinate and schedule the work between the various calculation modules. The modulo calculation module includes multiple modulo calculation units, including an NTT calculation unit. The high-performance double-precision floating-point calculation module is used to perform high-precision floating-point number operation. The data read-write bus connects the various calculation modules and the memory, realizing efficient data transmission. The high-bandwidth HBM memory is used to store ciphertext data and intermediate calculation results.

[0095] Referring to Figure 2 , Figure 2 is a high-speed SRAM memory block structure design schematic diagram provided by an embodiment of the present application. The design is mainly used for intermediate data storage in the modulo / floating-point calculation module.

[0096] Specifically, the high-speed SRAM memory block is composed of multiple SRAM memory units. Each SRAM memory unit can store 4096 word data, and each word data is 64 bits. Multiple SRAM memory units are combined into a memory block, and a single memory block can store a total of 65536 word data. This design can effectively improve the data storage density and access speed, and can be compatible with different sizes of ciphertext data, providing good algorithm compatibility for the upper application.

[0097] Each memory block provides a total of 1024 bit wide read-write interface. This wide interface design can significantly improve the data transmission efficiency and meet the needs of large-scale data processing in homomorphic encryption calculation.

[0098] Referring to Figure 3 , Figure 3 is a data storage and transmission architecture based on a high-speed SRAM memory block provided by an embodiment of the present application. The architecture is mainly used to optimize data flow and storage management in the modulo / floating point calculation module.

[0099] Specifically, the architecture implemented based on the SRAM memory includes the following main components:

[0100] SRAM memory block: composed of multiple high-speed on-chip SRAM memories, each high-speed memory array is composed of 4 high-speed on-chip SRAM memory blocks. This design provides large-capacity, high-speed on-chip data storage capability.

[0101] Data multiplexer: responsible for data read / write distribution, including data write-in path and data read-out path. Through the data multiplexer, the system can realize efficient read / write and transmission of data.

[0102] Memory block read / write address generator: used to generate read / write addresses of the memory block, to ensure that data can be correctly stored and accessed.

[0103] Configurable parameter register: stores system configuration parameters, which can be configured by the upper server through the calculation module control signaling data path. This design increases the flexibility and configurability of the system.

[0104] HBM memory data read / write path: connected to the HBM memory, used for high-speed read / write of large-capacity data.

[0105] As the core design of the present application, the memory array is provided with data read / write channel A and data read / write channel B, which are connected to the modulo calculation module (including NTT calculation unit) and the floating point calculation module respectively. This design enables the system to support integer operations and floating point operations simultaneously, thereby meeting the needs of different homomorphic encryption algorithms.

[0106] Specifically, the data stored in the SRAM memory array of the modulo calculation module can be quickly transmitted to the modulo calculation circuit through data read / write channel A or B for modulus multiplication, modulus addition, etc., and then written back to the SRAM memory array through data write-back channel A or B; the data stored in the SRAM memory array of the floating point calculation module can be transmitted to the floating point calculation circuit through data read / write channel A or B for addition, multiplication, etc., and then written back to the SRAM memory array through data write-back channel A or B. This direct data path design significantly reduces the delay of data transmission and improves the efficiency of integer operations. This bidirectional data flow design ensures that the calculation results can be stored in time, preparing for subsequent calculations or data transmission.

[0107] The architecture achieves efficient data storage and transmission through the design of multi-level storage and parallel data paths. The SRAM memory block provides fast on-chip data access, while the HBM memory provides large-capacity data storage. This design fully considers the needs of large-scale data processing in homomorphic encryption computation and can effectively improve the overall performance of the system.

[0108] As an improvement of the above scheme, the number and capacity of the SRAM memory block and the configuration of the HBM memory can be adjusted according to specific computing needs, thereby optimizing system performance and resource utilization.

[0109] Specifically, this data storage and transmission architecture can support efficient execution of multiple homomorphic encryption algorithms such as CKKS, BFV, and BGV. For example, when performing ciphertext computation based on large integer polynomials, the SRAM memory block can be used to store intermediate calculation results, and the data multiplexer can be used to quickly access and update data, thereby speeding up the operation process of the ciphertext polynomial. The HBM memory can be used to store large-scale ciphertext data, and the high-bandwidth interface can be used to quickly call the required data.

[0110] Referring to Figure 4 , Figure 4 This application shows an example of a soft and hardware cooperative memory distribution. This design is mainly used to optimize the modulo operation in homomorphic encryption algorithms, especially for efficient parallel processing of polynomials represented in RNS residue number system.

[0111] Specifically, the memory distribution scheme includes the following main components and features:

[0112] Modulo calculation module: Figure 4 As shown in the figure, there are 8 modulo calculation modules (modulo calculation module 1 to modulo calculation module 8). These modules can perform modulo operations in parallel, significantly improving the calculation efficiency.

[0113] HBM memory: The physical memory space of the HBM memory is divided into multiple virtual blocks, and each virtual block corresponds to a modulo calculation module. The capacity of each physical block is 256MB, and the secondary allocation operation of the physical block is defined by software. This division method allows each modulo calculation module to independently and in parallel access its corresponding memory space, avoiding read-write blocking caused by multiple calculation modules accessing a similar memory space at the same time.

[0114] Memory distribution strategy: The upper software combines multiple physical blocks into a virtual block, which increases the flexibility of memory management. Each virtual block is further divided to store polynomial coefficients of different RNS bases.

[0115] Polynomial storage mode: The lower half of the figure shows the distribution of polynomials represented by the RNS in the HBM memory. Each virtual block stores a set of related polynomial coefficients. This distribution takes into account the characteristics of the RNS modulus, allowing the modulus calculation unit to efficiently access the required data in parallel.

[0116] Parameter description:

[0117] k: represents the number of modulus calculation modules;

[0118] n: represents the number of RNS moduli;

[0119] m = n / k: represents the number of moduli processed by each modulus calculation module;

[0120] The main advantages of this design are: multiple modulus calculation modules can operate simultaneously, greatly improving the calculation efficiency; each modulus calculation module has a dedicated HBM memory area, reducing memory access conflicts and improving data reading speed; the upper software can dynamically adjust memory allocation according to specific requirements, adapting to different homomorphic encryption algorithms and parameter settings; the memory distribution mode fully considers the characteristics of the RNS, allowing related polynomial coefficients to be quickly accessed and processed.

[0121] In implementation, for example, when executing a multi-modal homomorphic encryption algorithm, the system can distribute the corresponding polynomial coefficients into different HBM virtual blocks according to the selected RNS basis. Each modulus calculation module can simultaneously process the data in its corresponding virtual block, performing modulus multiplication, modulus addition, etc. This parallel processing method can significantly improve the efficiency of large integer operations in homomorphic encryption.

[0122] Reference Figure 5 , Figure 5 This figure shows the pipeline scheduling process of NTT calculation in the modulus calculation unit in the embodiments of the present application. This design is mainly used to optimize the NTT operation in the homomorphic encryption algorithm, and improves the calculation efficiency through pipeline technology.

[0123] Specifically, this NTT calculation process utilizes multiple parallel SRAM arrays (such as SRAM arrays 1, 2, 3, 4, etc. in the figure) to achieve efficient pipeline calculation. The entire process is divided into multiple consecutive stages, each stage is performed simultaneously in different SRAM arrays. For example, SRAM arrays 1 and 2 are responsible for reading polynomials under rotation factors ω0 and q0, respectively, and then performing NTT calculation internally. At the same time, SRAM arrays 3 and 4 have already started preparing for the next round of NTT calculation, reading polynomials under rotation factors ω1 and q1, respectively. This design allows data reading, NTT calculation, and result writing back to be performed in parallel, greatly improving the calculation efficiency.

[0124] The key advantages of this pipeline design lie in its efficient parallel processing and optimized memory access. The simultaneous operation of multiple SRAM arrays significantly increases data processing speed, while dividing the NTT calculation process into multiple stages enables continuous data processing and reduces idle time. In particular, the system can perform data read and write operations in parallel while performing NTT calculations, effectively reducing overall computation time. Furthermore, this design offers excellent flexibility, allowing it to adapt to NTT calculations of varying scales by adjusting the number and allocation of SRAM arrays.

[0125] In practical applications, this pipelined NTT calculation method can effectively support various homomorphic encryption algorithms. For example, when performing polynomial multiplication in a multimodal homomorphic encryption algorithm, while one set of data is undergoing NTT transformation, the next set of data is already being prepared and loaded, significantly reducing the total execution time of the algorithm.

[0126] This pipelined NTT computing solution fully leverages the parallel computing capabilities of the FPGA and the high-speed access characteristics of SRAM, providing powerful hardware acceleration support for homomorphic encryption algorithms and enabling efficient, large-scale encryption calculations. This optimization is crucial for improving the overall performance of homomorphic encryption systems, enabling them to better meet the performance requirements of real-world applications.

[0127] based on Figures 1 to 5 Based on the information displayed, the overall workflow of this system can be described as follows:

[0128] The upper-layer server sends encrypted data and computing tasks to the FPGA hardware via the PCIe interface. After receiving the data, the PCIe data controller distributes it to the corresponding computing module based on the task type. The computing module control signaling data path coordinates the work of each computing module, ensuring the coordinated operation of the entire system.

[0129] During data processing, the system relies primarily on two core computing modules: the modular computing module and the high-performance double-precision floating-point computing module. The modular computing module plays a key role in homomorphic encryption algorithms such as CKKS, BGV, and BFV. It is primarily used to perform modular operations on polynomials, such as modular multiplication and modular addition. These operations are the fundamental computational steps in homomorphic encryption algorithms.

[0130] The main role of the high-performance double-precision floating-point computation module is to complete the RNS polynomial basis transformation operation, also known as base conversion. This operation plays a crucial role in the implementation of the ciphertext multiplication calculation of multi-modal homomorphic encryption. Base conversion allows efficient conversion of ciphertext data between different RNS bases, which is necessary for managing the modulus growth of ciphertext and maintaining computational precision.

[0131] By combining these two computation modules, the system can efficiently perform the complex calculations required by various homomorphic encryption algorithms. The modulo computation module handles basic polynomial operations, while the high-performance double-precision floating-point computation module handles base conversion operations that require high precision. Both modules work together to enable the system to support multiple homomorphic encryption schemes and flexibly switch between different encryption modes. This design not only improves the computational efficiency of the system but also enhances its flexibility and adaptability in handling various homomorphic encryption tasks.

[0132] Fast data transmission between modules is achieved through an efficient data read-write bus. At the same time, the system adopts an innovative memory management strategy, using HBM memory to store a large number of intermediate results and final results. As shown in Figure 4 , the HBM is divided into multiple virtual blocks by the upper memory scheduler, with each block corresponding to a modulo computation module. This design greatly improves data access efficiency.

[0133] Specifically, the modulo computation module internally adopts a pipeline NTT computation scheme as shown in Figure 5 . Multiple SRAM arrays work in parallel to achieve efficient parallel processing of data reading, NTT computation, and result writing back, significantly improving system performance.

[0134] In addition, the design of the system also considers flexibility and scalability. The modulo computation module and the high-performance double-precision floating-point computation module can contain different numbers of modulo computation units according to demand to adapt to different scales of computation tasks.

[0135] After the computation is completed, the results are returned to the upper server through the PCIe interface. Throughout the process, the system fully utilizes the parallel computing capabilities of FPGA, combined with optimized memory access strategies and pipeline design, significantly improving the efficiency of homomorphic encryption operations. This hardware acceleration scheme makes multi-modal homomorphic encryption more feasible in practical applications of different computation scales, maintaining good flexibility and scalability, and allowing for functional updates and performance optimization according to demand.

[0136] In summary, the embodiments of the present application can include the following beneficial effects:

[0137] The embodiments of the present application realize higher efficiency by combining the optimization design of software and hardware. The scheduling strategy of the ciphertext data is optimized on the software level, ensuring that the memory bandwidth is fully utilized, thereby improving the overall efficiency of data processing.

[0138] The embodiments of the present application integrate a fast number theory transformation circuit, making the large integer polynomial operation more efficient. By accelerating the homomorphic encryption calculation process, the calculation speed and accuracy are improved, and the polynomial multiplication processing time is significantly shortened.

[0139] The embodiments of the present application adopt a full-pipelined computing circuit design, significantly improving the throughput of ciphertext calculation. Through the pipeline technology, the system can quickly process multiple data streams, greatly improving the calculation efficiency and processing speed.

[0140] The embodiments of the present application perform memory scheduling according to the principle of the residue number system, decompose large integers into small modulus for calculation, and fully utilize hardware resources through scheduling algorithms, thereby improving parallel processing capability. This method effectively enhances the processing efficiency of the system and improves the throughput of ciphertext calculation.

[0141] The embodiments of the present application introduce high-bandwidth memory, which provides high-throughput data transmission capability through its multi-channel characteristics, solves the problem of insufficient memory bandwidth, and supports efficient scheduling and processing of large-scale ciphertext data.

[0142] The embodiments of the present application adjust the memory read-write order in real time according to the size of the ciphertext data through a dynamic scheduling strategy. This ensures the continuity and efficiency of the data stream, optimizes the utilization of system resources, and improves the overall performance.

[0143] The embodiments of the present application introduce a high-performance double-precision floating-point processing module to provide hardware support for the underlying operators of the multi-modal homomorphic encryption scheme, enabling the system to support different types of homomorphic encryption schemes simultaneously, thereby improving the compatibility of the system.

[0144] The embodiments of the present application design the computing module to be scalable, enabling the proposed architecture to support homomorphic encryption application scenarios under different computing scales, thereby optimizing the cost of deploying the homomorphic encryption hardware platform.

[0145] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0146] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures, or combine certain steps, or different steps.

[0147] The terms "first", "second", "third", "fourth" and the like in the description of the present application and in the claims, if any, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so construed can be interchanged, such that the embodiments of the present application described herein can operate in other sequences than illustrated or described herein. Further, the terms "comprise", "comprising", "include", "including", and the like, are meant to encompass non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise, include, or the like, a list of steps or elements, can include other steps or elements not expressly listed or inherent to such processes, methods, articles, or apparatuses.

[0148] It should be understood that, in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0149] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not limited to the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A multimodal, scalable, high-performance homomorphic encryption system, characterized by: The system includes upper-layer server software and driver architecture and multi-modal homomorphic encryption hardware architecture based on FPGA; The upper server software is in communication with the driver architecture and the multimodal homomorphic encryption hardware architecture implemented based on FPGA; The upper-layer server software and driver architecture includes a multimodal homomorphic encryption algorithm software library, a software memory scheduler, a modulo calculation module software driver, and a floating-point calculation module software driver; The multimodal homomorphic encryption hardware architecture implemented based on FPGA includes a communication data controller, multiple modulus calculation modules, floating-point calculation modules, a data read-write bus and a high-bandwidth memory; each of the modulus calculation modules includes a fast number theory transformation calculation unit; The communication data controller and the upper-layer server software are in communication with the driver architecture, and the communication data controller is also in communication with each of the modulo calculation modules and the floating-point calculation modules; the data read and write bus is in communication with the communication data controller, each of the modulo calculation modules, the floating-point calculation module, and the high-bandwidth memory. Each of the modulo calculation module and the floating-point calculation module includes a memory array composed of multiple SRAM memory blocks, a data multiplexer, a memory block read and write address generator, a configurable parameter register and a memory data read and write path; Wherein, each of the SRAM memory blocks is used to store data; The data multiplexer includes a data write path and a data read path, and the data multiplexer is used to read and write distribution data; The memory block read and write address generator is used to generate the read and write addresses of each of the SRAM memory blocks; The configurable parameter register is used to store configuration parameters, which are configured by the upper-layer server software and the driver architecture through the computing module control signaling data path; The memory data read and write path is connected to the high bandwidth memory, and the memory data read and write path is used to read and write data; The upper-layer server software and driver architecture are used to merge several physical blocks in the high-bandwidth memory into a virtual block, thereby obtaining a plurality of virtual blocks; each virtual block corresponds to one modulus calculation module; Each virtual block is divided into a plurality of virtual sub-blocks, and each virtual sub-block is used to store polynomial coefficients of different RNS bases.

2. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: The multimodal homomorphic encryption algorithm software library is used to manage and schedule the multimodal homomorphic encryption algorithm; The software memory scheduler is used to manage and schedule memory resources; The modulo calculation module software driver is used to control and drive each of the modulo calculation modules; The floating-point calculation module software driver is used to control and drive the floating-point calculation module.

3. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: The communication data controller is used to manage the interactive data between the upper server software and the driver architecture and the multimodal homomorphic encryption hardware architecture implemented based on FPGA; Each of the modulo calculation modules is used for modulo calculation; The floating-point calculation module is used for floating-point operations; The data read and write bus is used to transmit data; The high bandwidth memory is used to store ciphertext data and intermediate calculation results.

4. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: The upper server software and driver architecture are connected to the multimodal homomorphic encryption hardware architecture implemented based on FPGA through a PCIe data communication interface; The communication data controller adopts a PCIe data controller; The PCIe data controller is respectively connected to each of the modulo calculation modules and the floating-point calculation modules through the calculation module control signaling data path.

5. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: Each of the SRAM memory blocks is used to store intermediate data; Each of the SRAM memory blocks includes a plurality of SRAM memory cells; each of the SRAM memory blocks includes a read / write interface with a width of 1024 bits; The storage capacity of each of the SRAM memory cells is 4096 words of data, and each word of data is 64 bits; the storage capacity of a single SRAM memory block is 65536 words of data.

6. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: The memory array is provided with a data read and write channel A and a data read and write channel B; The data read and write channel A is connected to each of the modulo calculation modules, and the data read and write channel B is connected to the floating-point calculation module.

7. A multimodal, scalable, high-performance homomorphic encryption system according to claim 1, characterized in that: The fast number theory transformation calculation unit is used to execute a homomorphic encryption algorithm through a plurality of parallel operating memory arrays.

8. A multimodal, scalable, high-performance homomorphic encryption system according to any one of claims 1 to 7, characterized in that: The multimodal homomorphic encryption algorithm software library includes the CKKS homomorphic encryption algorithm, the BFV homomorphic encryption algorithm, and the BGV homomorphic encryption algorithm.

Citation Information

Patent Citations

  • Homomorphic encryption system and homomorphic encryption execution method based on reconfigurable technology

    CN113660076A

  • Hardware accelerator of fully homomorphic encryption algorithm, homomorphic encryption method and electronic equipment

    CN116488788A