A privacy protection medical image homomorphic inference method based on FPGA number theory transformation acceleration
Patent Information
- Application Number
- CN202610833150.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于提供一种基于FPGA数论变换加速的隐私保护医学影像同态推理方法,旨在解决现有技术存在硬件参数适配灵活性差、并行计算与数据流水线效率低以及数论变换计算延迟高的技术问题
[0056](1)本发明通过编译时参数化机制,以硬件描述语言参数定义语句集中声明多项式度数N、数据位宽B和并行BU数量P,片上BRAM阵列深度、地址映射关系及并行数据通道宽度均由顶层参数自动派生,切换多项式规模或并行度时仅需修改顶层参数并重新综合,无需修改子模块逻辑代码,避免了为不同医学影像同态推理场景设计多套硬件代码,缩短开发周期,适配不同医学影像分辨率和安全等级对多项式规模的差异化需求。
Smart Images

Figure CN122845731A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration, belonging to the field of homomorphic encryption hardware acceleration technology. Background Technology
[0002] Fully homomorphic encryption (FHE) allows arbitrary computations to be performed directly on encrypted data and is a key technology in the field of data privacy protection. Among many FHE schemes, the CKKS (Cheon-Kim-Kim-Song) scheme has become the mainstream choice for privacy-preserving machine learning due to its efficient support for real number approximation operations. In privacy-preserving medical image diagnosis scenarios, medical institutions need to send patients' CT, MRI, fundus, and other medical image data to the cloud for AI-assisted diagnosis. However, medical images involve patients' personal privacy, and directly transmitting plaintext data faces serious privacy leakage risks. After encrypting medical image data using the CKKS scheme, the cloud can directly perform convolutional neural network (CNN) inference operations in the encrypted domain, completing diagnostic tasks such as image classification and lesion detection without touching the plaintext, achieving privacy-preserving inference that is "usable but invisible to the data".
[0003] In CNN inference of encrypted medical images based on the CKKS scheme, the operations between the weights of convolutional and fully connected layers and the encrypted image pixels are ultimately decomposed into a large number of polynomial multiplication operations. The core computational step of polynomial multiplication is the number-theoretic transformation (NTT). Each ciphertext multiplication requires multiple forward and inverse NTT operations. During the inference process of a single medical image through multiple CNN layers, such operations need to be repeatedly performed on a large number of ciphertexts. The computational latency of NTT has become the main bottleneck restricting the throughput of encrypted medical image inference.
[0004] However, existing NTT hardware acceleration solutions still have many unresolved issues. Most NTT hardware solutions are designed for small-scale polynomial designs in post-quantum cryptography, lacking adaptability to large-scale polynomials in encrypted medical image inference. Existing FPGA NTT accelerators are mostly hard-coded designs with fixed scale and parallelism; switching polynomial degrees or the number of parallel cores requires redesigning the logic code, failing to adapt to the varying polynomial scale requirements of different medical image resolutions and security levels. Furthermore, current CKKS-based FPGA acceleration solutions mostly employ High-Level Synthesis (HLS) design methods, which, while efficient, have limitations in fine-grained circuit control and timing optimization. SystemVerilog hardware description methods based on Register Transfer Level (RTL) can achieve more precise pipeline control. Summary of the Invention
[0005] The purpose of this invention is to provide a privacy-preserving homomorphic inference method for medical images based on FPGA-accelerated number-theory transformations, aiming to solve the technical problems of poor hardware parameter adaptation flexibility, low efficiency of parallel computing and data pipeline, and high latency of number-theory transformation calculations in existing technologies.
[0006] To achieve the above objectives, the technical solution of the present invention is: a privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration. This method is aimed at the inference scenario of convolutional neural networks for encrypted medical images based on the CKKS fully homomorphic encryption scheme. The pixel data of the encrypted medical image is encoded by CKKS to obtain polynomial coefficients as input data of the accelerator. The forward NTT transformation result is used for subsequent homomorphic inference operations such as convolutional layer weight multiplication and fully connected layer matrix multiplication in the ciphertext domain. This method uses a hardware description language to centrally manage compile-time configurations, enabling the depth, address mapping, and parallel data channel width of the on-chip Block Random Access Memory Array (BRAM Array) to be automatically derived based on the polynomial degree and the number of parallel cores. This allows for adaptation to homomorphic inference scenarios corresponding to different medical image resolutions and security levels without modifying submodule logic. Simultaneously, efficient pairing of multiple polynomial coefficients is achieved through register buffer chains and multiplexers. After precise alignment with twiddle factors and pre-computed factors on the pipeline, these coefficients are fed in parallel into a butterfly computation core, completing modular multiplication, Barrett reduction, and modular addition / subtraction operations in a single clock cycle. The results are then written back in place, maximizing the parallel throughput of forward number theory transformations. Finally, a high-speed streaming transfer of polynomial coefficients, twiddle factors, pre-computed factors, and transformation results is achieved through an Advanced eXtensible Interface Stream (AXI-Stream), providing high-throughput, low-latency acceleration support for homomorphic inference tasks in encrypted medical images. The method includes the following steps:
[0007] Step 1: Encrypt the encrypted medical image data using the CKKS fully homomorphic encryption scheme to obtain ciphertext polynomial coefficients. Use the ciphertext polynomial coefficients as accelerator input and determine the accelerator intermediate parameters based on the set compile-time hardware basic parameters.
[0008] Step 2: Based on the intermediate parameters of the accelerator, the polynomial coefficients, twist factors at all levels, and pre-calculated factors are loaded from the external processor to the on-chip memory array through the advanced scalable streaming interface. After loading is completed, the control of the memory array interface is released, and the accelerator enters the forward number theory transformation calculation state.
[0009] Step 3: In the forward number theory transformation calculation state, the accelerator reads multiple polynomial coefficients, twiddle factors, and pre-calculated factors from the on-chip memory array each clock cycle. After being stored in register buffer chains and paired by multiplexers, these are sent to each parallel butterfly calculation core for butterfly operations. After multiple levels of butterfly operations, the forward number theory transformation is completed. Each butterfly calculation core performs one butterfly operation in a single clock cycle. The butterfly operations include multiplication, modular reduction and subtraction, and modular addition and subtraction operations. The accelerator writes the calculation results back to the coefficient memory array in the original read address.
[0010] Step 4: After the forward number theory transformation is completed, the forward number theory transformation result is read from the coefficient memory array through the advanced scalable streaming interface and transmitted to the external processor. The encrypted medical image homomorphic inference operation is then performed based on the forward number theory transformation result.
[0011] Optionally, step 1 specifically includes:
[0012] Step 1.1: Declare basic hardware parameters using the parameter definition statement set of the hardware description language, wherein the basic hardware parameters include polynomial degree N, data bit width B, and number of parallel butterfly computing cores P;
[0013] Step 1.2: Based on the declared basic hardware parameters, each submodule automatically obtains the basic hardware parameters and derives the accelerator intermediate parameters through parameter passing between module levels. The accelerator intermediate parameters include the coefficient memory array depth. Rotation factor memory array depth and pre-calculated factor memory array depth Address mapping relationships and parallel data channel width;
[0014] Step 1.3: When switching the polynomial degree N or the number of parallel butterfly computation cores P, only the top-level parameter definition needs to be modified, without modifying the submodule logic code.
[0015] Optionally, step 2 specifically includes:
[0016] Step 2.1: The external processor accesses memory via a direct memory array (DMI) using a bit-width... A high-level, scalable streaming interface data bus transmits data packets.
[0017] When sending the rotation factor data packet, a set of P rotation factors and a set of P corresponding pre-calculated factors are filled into the lower half and the higher half of the data bus, respectively, so that a set of rotation factors and their pre-calculated factors are transmitted in parallel each clock cycle, for a total of 2P rotation factors and pre-calculated factor data.
[0018] Specifically, when sending polynomial coefficient data packets, 2P polynomial coefficients are filled in parallel to all channels of the data bus, so that 2P polynomial coefficients are transmitted in parallel each clock cycle.
[0019] Step 2.2: The accelerator identifies whether it is currently in the rotation factor loading stage or the coefficient loading stage based on the commands issued by the external processor through the control channel;
[0020] In the rotation factor loading stage, P rotation factors and P pre-calculated factors are extracted from the lower half and the higher half of the data bus, respectively, and written into the rotation factor memory array and the pre-calculated factor memory array in address order.
[0021] In the coefficient loading stage, 2P polynomial coefficients are separated from the data bus and written into the coefficient memory array in address order.
[0022] The write addresses of the rotation factor memory array, the pre-calculated factor memory array, and the coefficient memory array are generated through a depth-address mapping relationship.
[0023] Step 2.3: After all the rotation factors, pre-calculated factors and polynomial coefficients have been written, the accelerator's internal state machine ends the loading process, releases the interface control over the rotation factor memory array, pre-calculated factor memory array and coefficient memory array, and enters the calculation state of the forward number theory transformation.
[0024] Optionally, step 3 specifically includes:
[0025] Step 3.1: After the accelerator enters the forward number theory transformation calculation state, it reads 2P polynomial coefficients in parallel from the coefficient memory array each clock cycle according to the address sequence determined by the butterfly series and butterfly group index. The read 2P polynomial coefficients are sent to the register buffer chain and multiplexer for pairing. The first-stage register group R1 receives 2P parallel polynomial coefficients from the coefficient memory array interface. The second-stage register group includes two sets of registers, R2 and R3. R2 conditionally registers the polynomial coefficients of R1, and R3 conditionally registers the polynomial coefficients from the coefficient memory array interface. The polynomial coefficients, based on the parity condition of the lower-order bits of the address sequence, are latched by R2 and R3, determining whether to latch the current polynomial coefficients or the polynomial coefficients in the condition register R1, or the polynomial coefficients from the coefficient memory array interface. The multiplexer, according to the current butterfly operation level, selects to pair polynomial coefficients from both sets of registers R2 and R3, or to cross-pair polynomial coefficients from a single set of registers R2. After the polynomial coefficients are stored in the register buffer chain and paired by the multiplexer, the multiplexer ultimately outputs 2P polynomial coefficients as P operand pairs required by P parallel butterfly computation cores. and , where subscript The index of the butterfly group to be transformed;
[0026] During polynomial coefficient buffering and pairing, P twitch factors are read in parallel from the twitch factor memory array, and P corresponding pre-computed factors are read in parallel from the pre-computed factor memory array. The read twitch factors and pre-computed factors are paired with operands. and Maintain pipeline alignment so that P operand pairs and P twitch factors and P pre-calculated factors arrive at the input of P butterfly computing cores on the same clock edge;
[0027] Step 3.2: Each butterfly computation core receives the paired operand pairs, twiddle factors, and pre-calculated factors. The operand pairs are... and The rotation factor is denoted as The modulus is denoted as , wherein the rotation factor For model Below The original unit root The expression for an integer power is:
[0028]
[0029] Among them, the index Determined by the current butterfly level and the index within the butterfly group. This represents the modulo operation;
[0030] Each rotation factor Corresponding to a pre-calculated factor The expression is:
[0031]
[0032] in, This indicates the floor function;
[0033] During butterfly operations, the Barrett reduction method is used for modulo reduction. First, the twiddle factor is calculated. operands The modular multiplication value is expressed as:
[0034]
[0035] in, Impairment value;
[0036] The Barrett auxiliary product is calculated using the following expression:
[0037]
[0038] in, An approximate quotient estimate is used to generate the Barrett reduction process, taking an auxiliary product. of high The approximate quotient is obtained. The expression is:
[0039]
[0040] Calculate relative to The approximate quotient product result is expressed as:
[0041]
[0042] in, This represents the estimated modulus multiple obtained by recovering the modulus from the approximate quotient;
[0043] Calculate Barrett reduction candidate values The expression is:
[0044]
[0045] like ,but ,otherwise , This represents the result of Barrett reduction in butterfly operations;
[0046] The butterfly operation output and channel results are as follows:
[0047]
[0048] The difference channel result is:
[0049]
[0050] in, For channel output, For differential channel output;
[0051] Step 3.3: Output A′ from the obtained P butterfly computing cores and channels. k AND / Difference channel output B′ k In-situ write-back coefficient memory array, write-back address and read operand A k and B k The addresses are the same; the results of each butterfly operation are the same. and The original operands in the coefficient memory array are covered sequentially. and ,through After iterative reading and writing back of the butterfly traversal, the coefficient memory array stores the final forward number theory transformation result.
[0052] Optionally, step 4 specifically includes:
[0053] Step 4.1: After the forward number theory transformation is completed, the accelerator notifies the external processor through an interrupt signal and starts the result output process. The N forward number theory transformation results in the coefficient memory array are transmitted in ascending order of address and in parallel 2P-way data transmission at a time through the advanced scalable streaming interface data channel to the pre-configured receive buffer of the external processor.
[0054] Step 4.2: After all N results have been transmitted, the accelerator sends a transmission end flag through the Advanced Scalable Streaming Interface data channel, and then returns to the idle state to wait for the next computation task; after receiving the forward number theory transformation result, the external processor completes the encrypted medical image homomorphic inference operation based on the forward number theory transformation result.
[0055] The beneficial effects of this invention are:
[0056] (1) This invention uses a compile-time parameterization mechanism to declare the polynomial degree N, data bit width B, and number of parallel BUs P in the hardware description language parameter definition statement set. The on-chip BRAM array depth, address mapping relationship, and parallel data channel width are all automatically derived from the top-level parameters. When switching polynomial scale or parallelism, only the top-level parameters need to be modified and resynthesized. There is no need to modify the submodule logic code, which avoids designing multiple sets of hardware code for different medical image homomorphic inference scenarios, shortens the development cycle, and adapts to the differentiated requirements of different medical image resolutions and security levels for polynomial scale.
[0057] (2) The present invention is based on the register transfer stage design of FPGA. Through multi-level register buffer chain and multiplexer, the conditional register and cross-arrangement pairing of multi-way polynomial coefficients are realized. P operand pairs, P twitch factors and P pre-calculated factors are precisely aligned on the same clock edge and sent to each parallel BU in parallel. Fixed delay is used to replace dynamic data dependency detection to ensure the pipeline runs continuously without blocking. Each butterfly computing core completes modular multiplication, Barrett modular reduction and modular addition and subtraction operations in a single clock cycle and writes the results back in place, maximizing the parallel throughput of forward NTT and meeting the requirements of low latency and high throughput for homomorphic inference of encrypted medical images.
[0058] (3) The present invention realizes the high-speed streaming transfer of polynomial coefficient loading and positive NTT result output through the AXI-Stream interface. 2P polynomial coefficients or a set of P twitch factors and their corresponding pre-calculated factors are transmitted in parallel in each clock cycle. The data supply and calculation result recovery efficiency is high, and the end-to-end positive NTT calculation delay is reduced. Attached Figure Description
[0059] Figure 1 This is a block diagram of the overall system architecture of the present invention;
[0060] Figure 2 This is a diagram of the internal microarchitecture of the butterfly computing core of the present invention. Detailed Implementation
[0061] To make the uses, technical solutions, and advantages of this invention clearer and easier to understand, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0062] Example 1: As Figure 1 As shown, a privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration includes the following steps:
[0063] Step 1: Encrypt the encrypted medical image data using the CKKS fully homomorphic encryption scheme to obtain ciphertext polynomial coefficients. Use the ciphertext polynomial coefficients as accelerator input and determine the accelerator intermediate parameters based on the set compile-time hardware basic parameters.
[0064] Step 1.1: Declare basic hardware parameters using the parameter definition statement set of the hardware description language, wherein the basic hardware parameters include polynomial degree N, data bit width B, and number of parallel butterfly computing cores P;
[0065] Step 1.2: Based on the declared basic hardware parameters, each submodule automatically obtains the basic hardware parameters and derives the accelerator intermediate parameters through parameter passing between module levels. The accelerator intermediate parameters include the coefficient memory array depth. Rotation factor memory array depth and pre-calculated factor memory array depth Address mapping relationships and parallel data channel width;
[0066] Step 1.3: When switching the polynomial degree N or the number of parallel butterfly computation cores P, only the top-level parameter definition needs to be modified and the implementation resynthesized; no modification to the submodule logic code is required.
[0067] Step 2: Based on the intermediate parameters of the accelerator, the polynomial coefficients, twist factors at all levels, and pre-calculated factors are loaded from the external processor to the on-chip memory array through the advanced scalable streaming interface. After loading is completed, the control of the memory array interface is released, and the accelerator enters the forward number theory transformation calculation state.
[0068] Step 2.1: The external processor accesses memory via a direct memory array (DMI) using a bit-width... A high-level, scalable streaming interface data bus transmits data packets.
[0069] When sending the rotation factor data packet, a set of P rotation factors and a set of P corresponding pre-calculated factors are filled into the lower half and the higher half of the data bus, respectively, so that a set of rotation factors and their pre-calculated factors are transmitted in parallel each clock cycle, for a total of 2P rotation factors and pre-calculated factor data.
[0070] Specifically, when sending polynomial coefficient data packets, 2P polynomial coefficients are filled in parallel to all channels of the data bus, so that 2P polynomial coefficients are transmitted in parallel each clock cycle.
[0071] Step 2.2: The accelerator identifies whether it is currently in the rotation factor loading stage or the coefficient loading stage based on the commands issued by the external processor through the control channel;
[0072] In the rotation factor loading stage, P rotation factors and P pre-calculated factors are extracted from the lower half and the higher half of the data bus, respectively, and written into the rotation factor memory array and the pre-calculated factor memory array in address order.
[0073] In the coefficient loading stage, 2P polynomial coefficients are separated from the data bus and written into the coefficient memory array in address order.
[0074] The write addresses of the rotation factor memory array, the pre-calculated factor memory array, and the coefficient memory array are generated through the depth and address mapping relationship determined in step 1.
[0075] Step 2.3: After all the rotation factors, pre-calculated factors and polynomial coefficients have been written, the accelerator's internal state machine ends the loading process, releases the interface control over the rotation factor memory array, pre-calculated factor memory array and coefficient memory array, and enters the calculation state of the forward number theory transformation.
[0076] Step 3: As Figure 2 As shown, in the forward number theory transformation calculation state, the accelerator reads multiple polynomial coefficients, twiddle factors, and pre-calculated factors from the on-chip memory array each clock cycle. After being stored in register buffer chains and paired by multiplexers, these are sent to each parallel butterfly unit (BU) for butterfly operations. After multiple levels of butterfly operations, the forward number theory transformation is completed. Each butterfly unit performs one butterfly operation in a single clock cycle. The butterfly operations include multiplication, modular reduction and subtraction, and modular addition and subtraction operations. The accelerator writes the calculation results back to the coefficient memory array in the original read address.
[0077] Step 3.1: After the accelerator enters the forward number theory transformation calculation state, it reads 2P polynomial coefficients in parallel from the coefficient memory array each clock cycle according to the address sequence determined by the butterfly series and butterfly group index. The read 2P polynomial coefficients are sent to the register buffer chain and multiplexer for pairing. The first-stage register group R1 receives 2P parallel polynomial coefficients from the coefficient memory array interface. The second-stage register group includes two sets of registers, R2 and R3. R2 conditionally registers the polynomial coefficients of R1, and R3 conditionally registers the polynomial coefficients from the coefficient memory array interface. The polynomial coefficients, based on the parity condition of the lower-order bits of the address sequence, are latched by R2 and R3, determining whether to latch the current polynomial coefficients or the polynomial coefficients in the condition register R1, or the polynomial coefficients from the coefficient memory array interface. The multiplexer, according to the current butterfly operation level, selects to pair polynomial coefficients from both sets of registers R2 and R3, or to cross-pair polynomial coefficients from a single set of registers R2. After the polynomial coefficients are stored in the register buffer chain and paired by the multiplexer, the multiplexer ultimately outputs 2P polynomial coefficients as P operand pairs required by P parallel butterfly computation cores. and , where subscript The index of the butterfly group to be transformed;
[0078] During polynomial coefficient buffering and pairing, P twitch factors are read in parallel from the twitch factor memory array, and P corresponding pre-computed factors are read in parallel from the pre-computed factor memory array. The read twitch factors and pre-computed factors are paired with operands. and Maintain pipeline alignment so that P operand pairs and P twitch factors and P pre-calculated factors arrive at the input of P butterfly computing cores on the same clock edge;
[0079] Step 3.2: Each butterfly computation core receives the paired operand pairs, twiddle factors, and pre-calculated factors. The operand pairs are... and The rotation factor is denoted as The modulus is denoted as , wherein the rotation factor For model Below The original unit root The expression for an integer power is:
[0080]
[0081] Among them, the index Determined by the current butterfly level and the index within the butterfly group. This represents the modulo operation;
[0082] Each rotation factor Corresponding to a pre-calculated factor The expression is:
[0083]
[0084] in, This indicates the floor function;
[0085] During butterfly operations, the Barrett reduction method is used for modulo reduction. First, the twiddle factor is calculated. operands The modular multiplication value is expressed as:
[0086]
[0087] in, Impairment value;
[0088] The Barrett auxiliary product is calculated using the following expression:
[0089]
[0090] in, An approximate quotient estimate is used to generate the Barrett reduction process, taking an auxiliary product. of high The approximate quotient is obtained. The expression is:
[0091]
[0092] Calculate relative to The approximate quotient product result is expressed as:
[0093]
[0094] in, This represents the estimated modulus multiple obtained by recovering the modulus from the approximate quotient;
[0095] Calculate Barrett reduction candidate values The expression is:
[0096]
[0097] like ,but ,otherwise , This represents the result of the Barrett reduction in the butterfly operation; in this embodiment, the high-order bits are taken during the Barrett reduction process. Floor down operation and It can be converted into a shift operation;
[0098] The butterfly operation output and channel results are as follows:
[0099]
[0100] The difference channel result is:
[0101]
[0102] in, For channel output, For differential channel output;
[0103] Optionally, in this embodiment, butterfly operations only require multiplication, shifting, and modulo addition and subtraction operations, and do not require a general-purpose divider;
[0104] Step 3.3: Output A′ from the obtained P butterfly computing cores and channels. k AND / Difference channel output B′ k Write back the coefficient memory array in situ, and write back the address and operand A read in step 3.1. k and B k The addresses are the same; the results of each butterfly operation are the same. and The original operands in the coefficient memory array are covered sequentially. and ,through After iterative reading and writing back of the butterfly traversal, the coefficient memory array stores the final forward number theory transformation result.
[0105] Step 4: After the forward number theory transformation is completed, the forward number theory transformation result is read from the coefficient memory array through the advanced scalable streaming interface and transmitted to the external processor. The encrypted medical image homomorphic inference operation is then performed based on the forward number theory transformation result.
[0106] Step 4.1: After the forward number theory transformation is completed, the accelerator notifies the external processor through an interrupt signal and starts the result output process. The N forward number theory transformation results in the coefficient memory array are transmitted in ascending order of address and in parallel 2P-way data transmission at a time through the advanced scalable streaming interface data channel to the pre-configured receive buffer of the external processor.
[0107] Step 4.2: After all N results have been transmitted, the accelerator sends a transmission end flag through the Advanced Scalable Streaming Interface data channel, and then returns to the idle state to wait for the next computation task; after receiving the forward number theory transformation result, the external processor completes the encrypted medical image homomorphic inference operation based on the forward number theory transformation result.
[0108] Optionally, in this embodiment, after obtaining the forward number theory transformation result, homomorphic inference operations of encrypted medical images, such as convolutional layer weight multiplication and fully connected layer matrix multiplication in the encrypted domain, are performed to complete a complete encrypted medical image homomorphic inference number theory transformation acceleration process.
[0109] Based on the detailed implementation description, the effectiveness of the technical solution of the present invention will be illustrated below through a specific example and experiment.
[0110] Specifically, in the scenario of homomorphic inference for privacy-preserving medical images, pixel data from patient CT, MRI, and other medical images are encoded using CKKS and represented as polynomial coefficients, which serve as input data to the accelerator. After the accelerator performs a forward NTT transformation, the transformation result is returned to an external processor for subsequent operations such as convolutional layer weight multiplication and fully connected layer matrix multiplication in the encrypted domain, thereby achieving cloud-based CNN inference for encrypted medical images. Further explanation is provided below with reference to specific parameters.
[0111] Specifically, taking the CKKS scenario with polynomial degree n=8192, RNS module number of 7, and 30 bits in homomorphic inference of privacy-preserving medical images as an example, the implementation process is as follows:
[0112] S1: Compile-time parameter settings. Key parameters are set through top-level macro definitions: N=8192, data width B=32, and number of parallel bus units P=8. Each submodule automatically obtains parameter values through a unified parameter import mechanism.
[0113] S2: Hardware Initialization. After system power-on, the loading module receives external data from the PS-side DMA via the AXI-Stream interface: 8192 polynomial coefficients are mapped and written to the coefficient BRAM array according to predetermined addresses; rotation factor... (Total 8192) and pre-calculated factors (A total of 8192), which are written to the corresponding twiddle factor BRAM and pre-calculated factor BRAM respectively. After loading is complete, control of the BRAM is released, and the system enters an idle state to wait for the calculation to start.
[0114] S3: NTT forward transition execution. The external controller sends a forward NTT start signal, and the forward transition control state machine generates control signals according to the following timing sequence:
[0115] Level 0 (stage=0): Butterfly step size Number of groups Each core processes 512 butterfly pairs, and each core performs a radix-2 butterfly operation on operands with an address spacing of 1. The current stage is completed in 512 clock cycles.
[0116] Level 1 (stage=1): Butterfly step size Number of groups Each core processes 512 butterfly pairs with an address spacing of 2, and the process takes 512 cycles.
[0117] And so on up to level 12 (stage=12): butterfly step length Number of groups Each core processes 512 butterfly pairs with an address spacing of 4,096, and the process takes 512 cycles.
[0118] total One effective computation clock cycle, with a forward NTT delay of ≈60.51μs at a clock speed of 110MHz.
[0119] S4: Result Reading. After the forward NTT is completed, the system sends a completion signal. The external controller reads 8192 calculation results from the coefficient BRAM via the AXI-Stream interface.
[0120] S5: The system was synthesized, placed, routed, and verified on the xczu9eg-ffvb1156-2-i (Zynq UltraScale+ MPSoC) FPGA platform. Device resource utilization and power consumption data are summarized in Table 1. Resources include look-up tables (LUTs), flip-flops (FFs), digital signal processing units (DSPs), and block RAM arrays (BRAMs).
[0121] Table 1 FPGA Accelerator Resource Consumption
[0122]
[0123] Furthermore, under the configuration of n=8192 and P=8, the performance comparison between the hardware and the Intel Core i5-10300H@2.5GHz single-threaded SEAL software is shown in Table 2. INTT is the reverse process of NTT, and the process is completely consistent with NTT.
[0124] Table 2 Performance Comparison of n=8192 (110MHz)
[0125]
[0126] Timing convergence status: Clock frequency 110MHz (period 9.000ns), global timing convergence, no timing violations, no functional errors.
[0127] As shown in Tables 1 and 2, the system completed placement and routing verification on the xczu9eg-ffvb1156-2-i FPGA platform. At a 110MHz clock speed (9.000ns period), the latency of a single NTT with n=8192 elements was 66.56μs. Resource consumption was 160 DSPs (6.35%), 48426 LUTs (17.67%), 32177 FFs (5.87%), and 43 BRAMs (4.71%). The total power consumption of the PL logic was 1.242W, and the timing was fully converged. The architecture adapts to different polynomial scales and parallelisms through compile-time parameterization without requiring modification to the hardware logic.
[0128] In summary, this invention addresses the inference scenario of convolutional neural networks for encrypted medical images based on the CKKS fully homomorphic encryption scheme. It employs the Cooley-Tukey butterfly algorithm with time decimation to decompose the N-point single-modulus number theory transformation into multi-level butterfly operations, which are executed in blocks by a configurable number of parallel butterfly computation cores. The memory subsystem utilizes a dual-port BRAM array in conjunction with a multi-level register buffer chain to achieve conflict-free in-situ memory access.
[0129] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration, characterized in that, The method includes the following steps: Step 1: Encrypt the encrypted medical image data using the CKKS fully homomorphic encryption scheme to obtain ciphertext polynomial coefficients. Use the ciphertext polynomial coefficients as accelerator input and determine the accelerator intermediate parameters based on the set compile-time hardware basic parameters. Step 2: Based on the intermediate parameters of the accelerator, the polynomial coefficients, twist factors at all levels, and pre-calculated factors are loaded from the external processor to the on-chip memory array through the advanced scalable streaming interface. After loading is completed, the control of the memory array interface is released, and the accelerator enters the forward number theory transformation calculation state. Step 3: In the forward number theory transformation calculation state, the accelerator reads multiple polynomial coefficients, twiddle factors, and pre-calculated factors from the on-chip memory array each clock cycle. After being stored in register buffer chains and paired by multiplexers, these are sent to each parallel butterfly calculation core for butterfly operations. After multiple levels of butterfly operations, the forward number theory transformation is completed. Each butterfly calculation core performs one butterfly operation in a single clock cycle. The butterfly operations include multiplication, modular reduction and subtraction, and modular addition and subtraction operations. The accelerator writes the calculation results back to the coefficient memory array in the original read address. Step 4: After the forward number theory transformation is completed, the forward number theory transformation result is read from the coefficient memory array through the advanced scalable streaming interface and transmitted to the external processor. The encrypted medical image homomorphic inference operation is then performed based on the forward number theory transformation result.
2. The privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration as described in claim 1, characterized in that, Step 1 specifically involves: Step 1.1: Declare basic hardware parameters using the parameter definition statement set of the hardware description language, wherein the basic hardware parameters include polynomial degree N, data bit width B, and number of parallel butterfly computing cores P; Step 1.2: Based on the declared basic hardware parameters, each submodule automatically obtains the basic hardware parameters and derives the accelerator intermediate parameters through parameter passing between module levels. The accelerator intermediate parameters include the coefficient memory array depth. Rotation factor memory array depth and pre-calculated factor memory array depth Address mapping relationships and parallel data channel width; Step 1.3: When switching the polynomial degree N or the number of parallel butterfly computation cores P, only the top-level parameter definition needs to be modified, without modifying the submodule logic code.
3. The privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration as described in claim 2, characterized in that, Step 2 specifically involves: Step 2.1: The external processor accesses memory via a direct memory array (DMI) using a bit-width... A high-level, scalable streaming interface data bus transmits data packets. When sending the rotation factor data packet, a set of P rotation factors and a set of P corresponding pre-calculated factors are filled into the lower half and the higher half of the data bus, respectively, so that a set of rotation factors and their pre-calculated factors are transmitted in parallel each clock cycle, for a total of 2P rotation factors and pre-calculated factor data. Specifically, when sending polynomial coefficient data packets, 2P polynomial coefficients are filled in parallel to all channels of the data bus, so that 2P polynomial coefficients are transmitted in parallel each clock cycle. Step 2.2: The accelerator identifies whether it is currently in the rotation factor loading stage or the coefficient loading stage based on the commands issued by the external processor through the control channel; In the rotation factor loading stage, P rotation factors and P pre-calculated factors are extracted from the lower half and the higher half of the data bus, respectively, and written into the rotation factor memory array and the pre-calculated factor memory array in address order. In the coefficient loading stage, 2P polynomial coefficients are separated from the data bus and written into the coefficient memory array in address order. The write addresses of the rotation factor memory array, the pre-calculated factor memory array, and the coefficient memory array are generated through a depth-address mapping relationship. Step 2.3: After all the rotation factors, pre-calculated factors and polynomial coefficients have been written, the accelerator's internal state machine ends the loading process, releases the interface control over the rotation factor memory array, pre-calculated factor memory array and coefficient memory array, and enters the calculation state of the forward number theory transformation.
4. The privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration as described in claim 3, characterized in that, Step 3 specifically involves: Step 3.1: After the accelerator enters the forward number theory transformation calculation state, it reads 2P polynomial coefficients in parallel from the coefficient memory array each clock cycle according to the address sequence determined by the butterfly series and butterfly group index. The read 2P polynomial coefficients are sent to the register buffer chain and multiplexer for pairing. The first-stage register group R1 receives 2P parallel polynomial coefficients from the coefficient memory array interface. The second-stage register group includes two sets of registers, R2 and R3. R2 conditionally registers the polynomial coefficients of R1, and R3 conditionally registers the polynomial coefficients from the coefficient memory array interface. The polynomial coefficients, based on the parity condition of the lower-order bits of the address sequence, are latched by R2 and R3, determining whether to latch the current polynomial coefficients or the polynomial coefficients in the condition register R1, or the polynomial coefficients from the coefficient memory array interface. The multiplexer, according to the current butterfly operation level, selects to pair polynomial coefficients from both sets of registers R2 and R3, or to cross-pair polynomial coefficients from a single set of registers R2. After the polynomial coefficients are stored in the register buffer chain and paired by the multiplexer, the multiplexer ultimately outputs 2P polynomial coefficients as P operand pairs required by P parallel butterfly computation cores. and , where subscript The index of the butterfly group to be transformed; During polynomial coefficient buffering and pairing, P twitch factors are read in parallel from the twitch factor memory array, and P corresponding pre-computed factors are read in parallel from the pre-computed factor memory array. The read twitch factors and pre-computed factors are paired with operands. and Maintain pipeline alignment so that P operand pairs and P twitch factors and P pre-calculated factors arrive at the input of P butterfly computing cores on the same clock edge; Step 3.2: Each butterfly computation core receives the paired operand pairs, twiddle factors, and pre-calculated factors. The operand pairs are... and The rotation factor is denoted as The modulus is denoted as , wherein the rotation factor For model Below The original unit root The expression for an integer power is: ; Among them, the index Determined by the current butterfly level and the index within the butterfly group. This represents the modulo operation; Each rotation factor Corresponding to a pre-calculated factor The expression is: ; in, This indicates the floor function; During butterfly operations, the Barrett reduction method is used for modulo reduction. First, the twiddle factor is calculated. operands The modular multiplication value is expressed as: ; in, Impairment value; The Barrett auxiliary product is calculated using the following expression: ; in, An approximate quotient estimate is used to generate the Barrett reduction process, taking an auxiliary product. of high The approximate quotient is obtained. The expression is: ; Calculate relative to The approximate quotient product result is expressed as: ; in, This represents the estimated modulus multiple obtained by recovering the modulus from the approximate quotient; Calculate Barrett reduction candidate values The expression is: ; like ,but ,otherwise , This represents the result of Barrett reduction in butterfly operations; The butterfly operation output and channel results are as follows: ; The difference channel result is: ; in, For channel output, For differential channel output; Step 3.3: Output A′ from the obtained P butterfly computing cores and channels. k AND / Difference channel output B′ k In-situ write-back coefficient memory array, write-back address and read operand A k and B k The addresses are the same; the results of each butterfly operation are the same. and The original operands in the coefficient memory array are covered sequentially. and ,through After iterative reading and writing back of the butterfly traversal, the coefficient memory array stores the final forward number theory transformation result.
5. The privacy-preserving homomorphic inference method for medical images based on FPGA number theory transformation acceleration as described in claim 4, characterized in that, Step 4 specifically involves: Step 4.1: After the forward number theory transformation is completed, the accelerator notifies the external processor through an interrupt signal and starts the result output process. The N forward number theory transformation results in the coefficient memory array are transmitted in ascending order of address and in parallel 2P-way data transmission at a time through the advanced scalable streaming interface data channel to the pre-configured receive buffer of the external processor. Step 4.2: After all N results have been transmitted, the accelerator sends a transmission end flag through the Advanced Scalable Streaming Interface data channel, and then returns to the idle state to wait for the next computation task; After receiving the results of the forward number theory transformation, the external processor performs homomorphic inference operations on encrypted medical images based on the results of the forward number theory transformation.