Acceleration unit, related device and method

By introducing number theory transformation units and arithmetic logic units into homomorphic encryption hardware and using scheduler to allocate instructions, the problems of poor performance and poor versatility of existing hardware are solved, and efficient algorithm compatibility and scalability are achieved.

CN114816334BActive Publication Date: 2025-06-03ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110067073.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-19
Publication Date
2025-06-03
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

The existing homomorphic encryption hardware implementations have poor global performance, poor generality and insufficient scalability, making it difficult to be compatible with changes in different algorithms.

Method used

An acceleration unit including a number theory transformation unit and an arithmetic logic unit is designed. Homomorphic encryption instructions are allocated to the corresponding unit for execution through a scheduler, realizing the decomposition and execution of number theory transformation and arithmetic logic operations.

Benefits of technology

It improves the global performance and versatility of homomorphic encryption hardware, enhances compatibility and scalability for different algorithms, and reduces the replacement cost of hardware structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816334B_ABST
    Figure CN114816334B_ABST
Patent Text Reader

Abstract

The present disclosure provides an acceleration unit, related devices, and methods. The acceleration unit includes: one or more number-theoretic transform units for performing number-theoretic transforms in a homomorphic encryption process; one or more arithmetic logic units for performing arithmetic operations in a homomorphic encryption process; and a scheduler for allocating operations in a homomorphic encryption instruction to be executed to at least one of the one or more number-theoretic transform units and the one or more arithmetic logic units. Embodiments of the present disclosure improve the versatility, global performance, and scalability of homomorphic encryption hardware deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of chips, and more particularly, to an acceleration unit, related devices and methods. Background Art

[0002] Big data provides great opportunities for machine learning and model analysis. However, due to privacy concerns, data often cannot be shared, forming data silos. For example, a platform needs to collect the behavior preference data of users on user terminals for big data analysis and more reasonable resource allocation, but users do not want to expose their privacy. Privacy computing based on homomorphic encryption has emerged as the times require. It aims to break data silos and perform computational modeling using data without revealing data privacy. Homomorphic encryption refers to such an encryption function that the addition and multiplication operations of plaintext on a ring are encrypted and then encrypted, and the result is equivalent to performing the corresponding operations on the ciphertext after encryption. Due to this good property, a third party can be entrusted to process the data without leaking information. An encryption function with homomorphic properties refers to an encryption function in which two plaintexts a and b satisfy Dec(En(a)⊙En(b)) = a⊕b, where En is the encryption operation, Dec is the decryption operation, and ⊙ and ⊕ correspond to the operations in the plaintext and ciphertext domains respectively. When ⊕ represents addition, the encryption is called additive homomorphic encryption; when ⊕ represents multiplication, the encryption is called multiplicative homomorphic encryption.

[0003] Currently, homomorphic encryption can be implemented on hardware through a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), etc. Whether it is a CPU, a GPU, or an FPGA, their hardware is designed for local algorithms of homomorphic encryption, which is detrimental to global performance, and the hardware is designed for dedicated algorithms of homomorphic encryption, with poor generality. Once the algorithm of homomorphic encryption changes, the originally planned hardware methods such as CPU, GPU, or FPGA may no longer be applicable and need to be migrated to another hardware method. Summary of the Invention

[0004] In view of this, the present disclosure aims to provide a hardware implementation method for homomorphic encryption that is general, has good global performance, and strong scalability.

[0005] According to one aspect of the present disclosure, an acceleration unit is provided, including:

[0006] One or more number theory transformation units for performing number theory transformation in the homomorphic encryption process;

[0007] One or more arithmetic logic units for performing arithmetic operations in the homomorphic encryption process;

[0008] A scheduler for allocating the operations in the to-be-executed homomorphic encryption instruction to at least one of the one or more number-theoretic transform units and the one or more arithmetic logic units.

[0009] Optionally, the acceleration unit further includes:

[0010] An instruction buffer for receiving control signals, where the control signals include the to-be-executed homomorphic encryption instruction;

[0011] An instruction fetch unit for fetching the to-be-executed homomorphic encryption instruction from the instruction buffer;

[0012] An instruction decoding unit for decoding the to-be-executed homomorphic encryption instruction fetched by the instruction fetch unit and sending the decoded to-be-executed homomorphic encryption instruction to the scheduler.

[0013] Optionally, the control signals further include an access memory address; the acceleration unit further includes: a memory interface for data transmission with the memory; a direct memory access unit for receiving the access memory address sent by the instruction buffer and instructing the memory interface to fetch the data required by the to-be-executed homomorphic encryption instruction according to the access memory address.

[0014] Optionally, the scheduler divides the to-be-executed homomorphic encryption instruction into at least one of the following operations: modular multiplication operation, modular addition operation, number-theoretic transform, inverse number-theoretic transform, modular swap, key exchange, rescaling.

[0015] Optionally, for the modular multiplication operation or modular addition operation, the scheduler allocates it to at least one of the one or more arithmetic logic units.

[0016] Optionally, for the number-theoretic transform, the scheduler allocates it to at least one of the one or more number-theoretic transform units.

[0017] Optionally, for the inverse number-theoretic transform, the scheduler decomposes the inverse number-theoretic transform into a combination of a number-theoretic transform and a modular multiplication operation, allocates the decomposed number-theoretic transform to at least one of the one or more number-theoretic transform units, and allocates the decomposed modular multiplication operation to at least one of the one or more arithmetic logic units.

[0018] Optionally, for the modular swap, the scheduler decomposes the modular swap into a combination of modular addition and modular multiplication and allocates it to at least one of the one or more arithmetic logic units.

[0019] Optionally, for the key exchange, the scheduler decomposes the key exchange into number-theoretic transforms, inverse number-theoretic transforms, modular multiplications, and modular exchanges, and allocates them to at least one of the one or more number-theoretic transform units or the one or more arithmetic logic units according to the decomposition results.

[0020] Optionally, for the rescaling, the scheduler decomposes the rescaling into number-theoretic transforms, inverse number-theoretic transforms, and modular exchanges, and allocates them to at least one of the one or more number-theoretic transform units or the one or more arithmetic logic units according to the decomposition results.

[0021] Optionally, at least one of the one or more number-theoretic transform units includes:

[0022] A first polynomial coefficient storage subunit;

[0023] A second polynomial coefficient storage subunit;

[0024] A rotation factor storage subunit for storing rotation factors in the number-theoretic transform;

[0025] A butterfly processing subunit for performing a first butterfly process and a second butterfly process, where the first butterfly process includes: reading a first pair of polynomial coefficients from the first polynomial coefficient storage subunit, obtaining the rotation factor corresponding to the first pair of polynomial coefficients from the rotation factor storage subunit, obtaining a second pair of polynomial coefficients based on the first pair of polynomial coefficients and the rotation factor, and writing them into the second polynomial coefficient storage subunit; the second butterfly process includes: reading a third pair of polynomial coefficients from the second polynomial coefficient storage subunit, obtaining the rotation factor corresponding to the third pair of polynomial coefficients from the rotation factor storage subunit, obtaining a fourth pair of polynomial coefficients based on the third pair of polynomial coefficients and the rotation factor, and writing them into the first polynomial coefficient storage subunit.

[0026] Optionally, at least one of the one or more number-theoretic transform units further includes: a control unit for controlling the operations of the first polynomial coefficient storage subunit, the second polynomial coefficient storage subunit, the rotation factor storage subunit, and the butterfly processing subunit.

[0027] Optionally, the first polynomial coefficient storage subunit and the second polynomial coefficient storage subunit each include a plurality of memory banks, each memory bank having a corresponding index, the first pair of polynomial coefficients and the third pair of polynomial coefficients come from the memory bank with the same index, and the second pair of polynomial coefficients and the fourth pair of polynomial coefficients come from the memory bank with the same index.

[0028] Optionally, there are M memory banks in the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit, and M / 2 butterfly processing subunits. For a single butterfly processing subunit, the indices of the pair of memory banks from which the first polynomial coefficient pair comes are the same as the indices of the pair of memory banks from which the third polynomial coefficient pair comes, and the indices of the two memory banks in this pair of memory banks differ by M / 2; the indices of the pair of memory banks from which the second polynomial coefficient pair comes are the same as the indices of the pair of memory banks from which the fourth polynomial coefficient pair comes, and the indices of the two memory banks in this pair of memory banks are adjacent.

[0029] Optionally, after the first butterfly processing or the second butterfly processing is performed log 2 M times, the control unit transposes the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit, and then the butterfly processing subunit performs the first butterfly processing or the second butterfly processing of log 2 M, where the transposition includes: taking out the polynomial coefficients queued with the same serial number in each memory bank of the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit before transposition, and placing them in a memory bank after transposition in the order of the memory bank indices.

[0030] Optionally, the butterfly processing subunit includes a first multiplexer, a second multiplexer, a third multiplexer, a fourth multiplexer, a first adder, a second adder, a first subtractor, a second subtractor, and a first multiplier.

[0031] Optionally, the first coefficient in the first polynomial coefficient pair or the third polynomial coefficient pair is input to the first input end of the first multiplexer, and after being added to the second coefficient in the first polynomial coefficient pair or the third polynomial coefficient pair by the first adder, it is input to the second input end of the first multiplexer. The first selection signal of the first multiplexer selects one of the first input end and the second input end to the output end; the second coefficient is input to the second input end of the third multiplexer, and after being subtracted from the first coefficient by the first subtractor, it is input to the first input end of the third multiplexer. The third selection signal of the third multiplexer selects one of the first input end and the second input end to the output end, and after being multiplied by the corresponding rotation factor by the first multiplier, a product signal is obtained.

[0032] Optionally, the signal output from the output terminal of the first multiplexer is input to the first input terminal of the second multiplexer, and is added to the product signal by a second adder and then input to the second input terminal of the second multiplexer. One of the first input terminal and the second input terminal is selected by the second selection signal of the second multiplexer and output to the output terminal as one of the second polynomial coefficient pair or the fourth polynomial coefficient pair; the product signal is input to the second input terminal of the fourth multiplexer, and is subtracted from the signal output from the first multiplexer by a second subtractor and then input to the first input terminal of the fourth multiplexer. One of the first input terminal and the second input terminal is selected by the fourth selection signal of the fourth multiplexer and output to the output terminal as the other coefficient of the second polynomial coefficient pair or the fourth polynomial coefficient pair.

[0033] Optionally, the arithmetic logic unit includes a modulo adder, a modulo multiplier, a fifth multiplexer, a sixth multiplexer, and a seventh multiplexer. Among them, the first input signal is input to the first input terminals of the fifth multiplexer and the sixth multiplexer, the output of the modulo adder is input to the second input terminal of the sixth multiplexer, the output of the modulo multiplier is input to the second input terminal of the fifth multiplexer. One of the first input terminal and the second input terminal of the fifth multiplexer is selected by the fifth selection signal of the fifth multiplexer and output to the output terminal. One of the first input terminal and the second input terminal of the sixth multiplexer is selected by the sixth selection signal of the sixth multiplexer and output to the output terminal. The output terminal of the fifth multiplexer is connected to the first input terminal of the modulo adder, the second input signal is input to the second input terminal of the modulo adder, the output terminal of the sixth multiplexer is connected to the first input terminal of the modulo multiplier, the third input signal is input to the second input terminal of the modulo multiplier, the output terminal of the modulo adder is connected to the first input terminal of the seventh multiplexer, and the output terminal of the modulo multiplier is connected to the second input terminal of the seventh multiplexer. One of the first input terminal and the second input terminal of the seventh multiplexer is selected by the seventh selection signal of the seventh multiplexer and output to the output terminal.

[0034] Optionally, by setting the fifth selection signal to select the first input terminal and setting the seventh selection signal to select the first input terminal, the arithmetic logic unit is used for modulo addition operation.

[0035] Optionally, by setting the sixth selection signal to select the first input terminal and setting the seventh selection signal to select the second input terminal, the arithmetic logic unit is used for modulo multiplication operation.

[0036] Optionally, by setting the fifth strobe signal to strobe the second input terminal, setting the sixth strobe signal to strobe the first input terminal, and setting the seventh strobe signal to strobe the first input terminal, the arithmetic logic unit is used for modulo multiplication followed by modulo addition operations.

[0037] Optionally, by setting the fifth strobe signal to strobe the first input terminal, setting the sixth strobe signal to strobe the second input terminal, and setting the seventh strobe signal to strobe the second input terminal, the arithmetic logic unit is used for modulo addition followed by modulo multiplication operations.

[0038] According to one aspect of the present disclosure, there is provided a computing device, including:

[0039] A memory for storing homomorphic encryption instructions to be executed;

[0040] The acceleration unit as described above;

[0041] A processing unit for loading the homomorphic encryption instructions to be executed and distributing the homomorphic encryption instructions to be executed to the acceleration unit for execution.

[0042] According to one aspect of the present disclosure, there is provided a system-on-chip including the acceleration unit as described above.

[0043] According to one aspect of the present disclosure, there is provided a data center including the computing device as described above.

[0044] According to one aspect of the present disclosure, there is provided a homomorphic encryption method, including:

[0045] Receiving homomorphic encryption instructions to be executed;

[0046] Decomposing the homomorphic encryption instructions to be executed into operations;

[0047] Allocating the number-theoretic transforms included in the operations to at least one of one or more number-theoretic transform units for execution, and allocating the arithmetic operations included in the operations to at least one of one or more arithmetic logic units for execution.

[0048] Through the analysis and decomposition of various algorithms of homomorphic encryption, it is determined that various operations (including modular addition, modular multiplication, number-theoretic transform / inverse number-theoretic transform, key exchange, modulus exchange, rescaling, etc.) into which homomorphic encryption is decomposed can ultimately be decomposed into number-theoretic transform and arithmetic logic (modular addition, modular multiplication, and their combinations). Therefore, a number-theoretic transform unit is introduced to perform number-theoretic transform, and an arithmetic logic unit is introduced to perform arithmetic logic. Through the scheduling of a scheduler, several number-theoretic transform units and several arithmetic logic units can independently execute different tasks, or can be combined into a pipeline to sequentially execute different stages of the same task. Therefore, this architecture can efficiently be compatible with different types of algorithms. Compared with the prior art in which hardware is designed for local algorithms, it improves the global performance. Compared with the prior art in which hardware is designed for dedicated algorithms, it has strong scalability and versatility. Once the algorithm of homomorphic encryption changes, since the new algorithm can still be implemented through various combinations of number-theoretic transform and arithmetic logic, the hardware structure still does not need to be changed. Therefore, the embodiments of the present disclosure propose a hardware implementation method for homomorphic encryption that is general, has good global performance, and has strong scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Through the description of the embodiments of the present disclosure with reference to the following drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0050] Figure 1 is a structural diagram of a data center applied in an embodiment of the present disclosure;

[0051] Figure 2 is an internal structural diagram of a server in the data center of an embodiment of the present disclosure;

[0052] Figure 3 is an internal structural diagram of a processing unit and an acceleration unit inside the server according to an embodiment of the present disclosure;

[0053] Figure 4 is Figure 3 an internal structural diagram of a number-theoretic transform unit in;

[0054] Figure 5 is Figure 4 an internal structural diagram of a butterfly processing subunit in;

[0055] Figure 6 is Figure 3 an internal structural diagram of an arithmetic logic unit in;

[0056] Figure 7 is Figure 6 a table of the functions exercised by the arithmetic logic unit in the case where the control terminals a, b, and c are combinations of different inputs of;

[0057] Figure 8Schematic diagram of how a butterfly processing subunit fetches and stores data between a first polynomial coefficient storage subunit and a second polynomial coefficient storage subunit according to an embodiment of the present disclosure;

[0058] Figure 9 Flowchart of a homomorphic encryption method according to an embodiment of the present disclosure. Detailed implementation manners

[0059] The present disclosure is described below based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. Additionally, the drawings are not necessarily drawn to scale.

[0060] The following terms are used herein.

[0061] Privacy computing: Big data provides great opportunities for machine learning and model analysis. However, due to privacy concerns, data often cannot be shared, resulting in data silos. For example, a platform needs to collect the behavioral preference data of users on user terminals for big data analysis and more reasonable resource allocation, but users do not want to expose their privacy. Privacy computing emerges as the times require. Privacy computing refers to using data for computational modeling without revealing data privacy.

[0062] Homomorphic encryption: It refers to an encryption function such that the result of encrypting the plaintext after performing addition and multiplication operations on the ring is equivalent to performing the corresponding operations on the ciphertext after encryption. Due to this good property, a third party can be entrusted to process the data without revealing information. An encryption function with homomorphic properties is an encryption function where two plaintexts a and b satisfy Dec(En(a)⊙En(b)) = a⊕b, where En is the encryption operation, Dec is the decryption operation, and ⊙, correspond to the operations on the plaintext and ciphertext domains respectively. When ⊕ represents addition, the encryption is called additive homomorphic encryption; when represents multiplication, the encryption is called multiplicative homomorphic encryption.

[0063] Acceleration Unit: A processing unit designed to improve the data processing speed in some specialized fields (such as processing images, processing various operations of deep learning networks, etc.) where traditional processing units are not efficient. The acceleration unit is also known as an artificial intelligence (AI) processing unit, including a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and dedicated intelligent acceleration hardware (such as a neural network processor NPU, a hardware accelerator, etc.). The acceleration unit adopted in the embodiments of the present disclosure is designed for the field of homomorphic encryption, improving the generality and globality of homomorphic encryption.

[0064] Processing Unit: A unit that performs traditional processing (processing other than complex operations such as image processing and fully connected operations in various deep learning networks) in the servers of a data center. In addition, the processing unit also undertakes the scheduling function for the acceleration unit and itself, and allocates the tasks that need to be undertaken to the acceleration unit and itself. The processing unit can adopt various forms such as a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.

[0065] Number Theoretic Transform: A fast algorithm for calculating convolution. The algorithm itself is similar to the fast Fourier transform, but different from the fast Fourier transform where the rotation factor is a complex number, the rotation factor in the number theoretic transform is an integer modulo p. In other words, the fast Fourier transform is a transform defined on the complex plane, while the number theoretic transform is a transform defined on the polynomial ring. It is widely used in homomorphic encryption. In homomorphic encryption, the plaintext or ciphertext is represented in the form of polynomial coefficients. The coefficients of a polynomial are connected end to end to form a polynomial ring. During encryption, the plaintext is first processed into polynomial coefficients, and the encryption process is to perform arithmetic operations on each coefficient to change it into other values, and the decryption process is to change it back to the original value. Various processes on the ciphertext are also achieved by performing various transformations on the ciphertext polynomial coefficients. Whether it is the transformation of the plaintext coefficients or the ciphertext coefficients, an important transformation is the number theoretic transform.

[0066] Inverse Number Theoretic Transform: The reverse process of the above number theoretic transform. The inverse number theoretic transform can be regarded as a combination of the number theoretic transform and modular multiplication.

[0067] Arithmetic operations in the encryption process: Generally refers to modular operations, mainly modular addition and modular multiplication operations, as well as various combinations of modular addition and modular multiplication. The modulus is the transliteration of "Mod", and modular operations are widely used in program writing. The meaning of Mod is to find the remainder. Modular operations have extensive applications in both number theory and programming. From the discrimination of odd and even numbers to the discrimination of prime numbers, from modular exponentiation operations to the method of finding the greatest common divisor, from the Chinese Remainder Theorem problem to the Caesar cipher problem, the figure of modular operations can be found everywhere. Modular addition is an operation like this: (a + b) % p, and its result is the remainder when the arithmetic sum of a + b is divided by p, which is used as the coefficient of the transformed polynomial. That is to say, (a + b) = kp + r, where k, p, and r are positive integers. Modular multiplication is an operation like this: (a * b) % p, and its result is the remainder when the arithmetic multiplication of a * b is divided by p, which is used as the coefficient of the transformed polynomial. That is to say, (a * b) = kp + r, where k, p, and r are positive integers.

[0068] Modulus Switch: In the technology of homomorphic encryption, the ciphertext exists in the form of polynomial coefficients, and these coefficients are the residue classes modulo another number. The purpose of modulus switch is to change the modulus corresponding to these coefficients to another modulus. For example, originally these coefficients are the residue classes modulo 5, and after modulus switch, they become the residue classes modulo 3. Therefore, during modulus switch, the polynomial coefficients will change. Modulus switch can be decomposed into a combination of the above basic operations of modular addition and modular multiplication.

[0069] Key Switch: In the technology of homomorphic encryption, the polynomial coefficients of the ciphertext are obtained by encrypting the plaintext with a specific key. Key switch can process the ciphertext and change its key to another key. This processing process requires an additional key for replacement, which is called the key-switch key. After key switch, the polynomial coefficients of the ciphertext are transformed into other polynomial coefficients of the ciphertext. Key switch can be decomposed into a combination of number-theoretic transform, inverse number-theoretic transform, modular multiplication, and modulus switch operations.

[0070] Rescale: Given a number a (0 <= a < M) modulo M and a scaling factor r, rescale means dividing a by r, and the range of the scaled number a' = a / r becomes 0 <= a < M / r. Through rescaling, the above-mentioned polynomial coefficients of the ciphertext all become 1 / r of the original. Rescale can be decomposed into a combination of number-theoretic transform, inverse number-theoretic transform, and modulus switch.

[0071] Scheduling: On the one hand, sort the operations in the instruction according to the operation order in the instruction. On the other hand, according to the load of the units to be scheduled (in the embodiments of the present disclosure, they are various number-theoretic transform units and various arithmetic logic units), allocate these operations to the appropriate units for execution. This entire process is called scheduling.

[0072] Butterfly processing: The number-theoretic transform can be regarded as a process of repeatedly processing the polynomial coefficients in the polynomial ring according to the same rule. For example, the polynomial coefficients at two fixed positions in the polynomial ring are taken out, processed with a rotation factor, and a pair of new polynomial coefficients are obtained and put back into the other two fixed positions in the polynomial ring. For example, the first polynomial coefficient and the fifth polynomial coefficient in the polynomial ring are taken out, processed with a rotation factor, and then put back into the positions of the first and the second polynomial coefficients in the polynomial ring. When all the polynomial coefficients are taken out, processed with the rotation factor, and then put back into the positions of the other polynomial coefficients in the polynomial ring, the above process is repeated, that is, the first polynomial coefficient and the fifth polynomial coefficient in the new polynomial ring are taken out, processed with the rotation factor, and then put back into the positions of the first and the second polynomial coefficients in the polynomial ring. In order to repeatedly execute this operation, butterfly processing is created. That is, two storage sub-units of the polynomial ring are set. When the polynomial coefficients are taken out from a predetermined position of the first storage sub-unit, they are processed with the rotation factor and put into another predetermined position of the second storage sub-unit. Conversely, when the polynomial coefficients are taken out from a predetermined position of the second storage sub-unit, they are processed with the rotation factor and put into another predetermined position of the first storage sub-unit. In this way, the process of repeatedly processing with the rotation factor in the number-theoretic transform is simplified, and this repeated process is regarded as an inversion operation of the same process for different storage sub-units. Each operation of taking out polynomial coefficients, processing them, and writing the results into different polynomial coefficient positions is called butterfly processing.

[0073] Memory bank: A storage module, and each storage module sequentially accommodates a storage queue. For example, the queue formed by the polynomial coefficients in the embodiments of the present disclosure. As Figure 8 In the example shown, each of the first polynomial coefficient storage sub-unit 2351 or the second polynomial coefficient storage sub-unit 2352 includes a number of memory banks 23511, and each memory bank stores a queue of 8 polynomial coefficients. When the ciphertext is represented as a polynomial of 64 polynomial coefficients, the 64 polynomial coefficients can be distributed in 8 memory banks, and each memory bank stores 8 polynomial coefficients.

[0074] Transpose: The rows of the array become the columns of the array, and the columns of the array become the rows of the array. When the plaintext or ciphertext is represented as the polynomial coefficients of a polynomial, the polynomial coefficients are stored in a plurality of memory banks as described above, and each memory bank stores a part of the polynomial coefficients. In this context, transpose means that the polynomial coefficients at the same position (the same index) of each memory bank become a new column (memory bank), so that each new column (each memory bank) contains one polynomial coefficient from each of the original memory banks, that is, the rows and columns of the array of polynomial coefficients stored in each memory bank are interchanged.

[0075] Multiplexer: A device that connects multiple input terminals and, in response to a strobe signal, connects one of the multiple input terminals to the output terminal for output.

[0076] System-on-a-chip (SoC) refers to the technology of integrating a complete system on a single chip and grouping all or part of the necessary electronic circuits. A complete system generally includes a central processing unit (CPU) or acceleration unit, memory, and peripheral circuits, etc. SoC has developed in parallel with other technologies, such as silicon-on-insulator (SOI), which can provide enhanced clock frequencies, thereby reducing the power consumption of microchips.

[0077] Data center

[0078] A data center is a specific network of devices for global collaboration, used to transfer, accelerate, display, compute, and store data information on the Internet network infrastructure. In future development, data centers will also become assets for enterprise competition. With the widespread application of data centers, artificial intelligence and the like are increasingly applied to data centers. And deep learning, as an important technology of artificial intelligence, has been widely applied to big data analysis and computing in data centers.

[0079] In traditional large data centers, the network structure is usually as Figure 1 shown, that is, the hierarchical inter-networking model. This model consists of the following parts:

[0080] Server 140: Each server 140 is a processing and storage entity in the data center, and the processing and storage of a large amount of data in the data center are completed by these servers 140.

[0081] Access switch 130: The access switch 130 is a switch used to allow the server 140 to access the data center. One access switch 130 accesses multiple servers 140. The access switch 130 is usually located at the top of the rack, so they are also called Top of Rack switches, and they physically connect to the servers.

[0082] Aggregation switch 120: Each aggregation switch 120 connects multiple access switches 130 and provides other services at the same time, such as firewalls, intrusion detection, network analysis, etc.

[0083] Core switch 110: The core switch 110 provides high-speed forwarding for packets entering and leaving the data center and provides connectivity for the aggregation switch 120. The network of the entire data center is divided into an L3 layer routing network and an L2 layer routing network, and the core switch 110 usually provides a flexible L3 layer routing network for the network of the entire data center.

[0084] Under normal circumstances, the aggregation switch 120 is the demarcation point of the L2 and L3 layer routing networks. Below the aggregation switch 120 is the L2 network, and above is the L3 network. Each group of aggregation switches manages a Point of Delivery (POD), and each POD has an independent VLAN network. When a server migrates within a POD, it does not need to modify the IP address and default gateway because one POD corresponds to one L2 broadcast domain.

[0085] The Spanning Tree Protocol (STP) is usually used between the aggregation switch 120 and the access switch 130. STP enables only one aggregation layer switch 120 to be available for a VLAN network, and other aggregation switches 120 are used only when a failure occurs. That is to say, at the level of the aggregation switch 120, horizontal expansion cannot be achieved because even if multiple aggregation switches 120 are added, only one is still working.

[0086] The embodiments of the present disclosure can be applied to scenarios such as privacy-preserving multiparty computing, privacy-preserving machine learning, and on-device prediction. When applied to privacy-preserving multiparty computing, the computing is performed by Figure 1 one of the servers 140 therein. Multiparty computing requires data from multiple parties, including plaintext data of some parties and ciphertext data of other parties. These plaintext data or ciphertext data may come from Figure 1 other servers 140 therein respectively. The server 140 performing the computing is connected to the server 140 where the plaintext data or ciphertext data is located through the above-mentioned access switch 130, aggregation switch 120, core switch 110, etc., obtains the plaintext data or ciphertext data from the server 140 where the data is located, and performs homomorphic encryption operations using the following hardware structure of the embodiments of the present disclosure.

[0087] Server

[0088] The server 140 is the actual processing device in the data center. Figure 2A structural block diagram inside a server 140 is shown. The server 140 includes a memory 210, a processing unit cluster 270, and an acceleration unit cluster 280 connected by a bus. The processing unit cluster 270 includes multiple processing units 220. The acceleration unit cluster 280 includes multiple acceleration units 230. The acceleration unit 230 is a processing unit designed to improve the data processing speed in specialized application fields. The acceleration unit is an artificial intelligence (AI) processing unit, including a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and dedicated intelligent acceleration hardware (e.g., a neural network processor NPU, a hardware accelerator, etc.).

[0089] In the architecture design of the traditional processing unit 220, a large part of the space in the architecture is occupied by the control unit and the memory, while the space occupied by the computing unit is insufficient. Therefore, it is very effective in logical control, but inefficient in large-scale parallel computing. Therefore, various specialized acceleration units 230 have been developed to perform more effective processing for improving the operation speed for different functions and different fields of computing. In the embodiments of the present disclosure, the acceleration unit 230 is a hardware accelerator designed for homomorphic encryption operations. The data and intermediate results in this kind of operation are closely related throughout the calculation process and are often used. With the existing processing unit architecture, since the memory capacity inside the core of the processing unit is very small, a large amount of external memory needs to be frequently accessed, resulting in inefficient processing. By using this dedicated acceleration unit, since it has structures such as a large number of internal buffers, it avoids frequent access to the external memory of the core, greatly improving the processing efficiency and computing performance. In addition, when designing the acceleration unit 230 in the embodiments of the present disclosure, the generality and overallness for various homomorphic encryption operations are considered, which will be described in detail below.

[0090] The acceleration unit 230 is to be scheduled by the processing unit 220. The memory 210 stores homomorphic encryption instructions. When needed, these instructions are Figure 2 deployed by one of the processing units 220 in the

[0091] Internal structure of the processing unit and the acceleration unit

[0092] The following is combined with Figure 3Internal structure diagrams of the processing unit 220 and the acceleration unit 230, specifically explaining how the processing unit 220 schedules the acceleration unit 230 and itself to work.

[0093] As Figure 3 shown, the processing unit 220 includes multiple processor cores 222 and a cache 221 shared by the multiple processor cores 222. Each processor core 222 includes an instruction fetch unit 203, an instruction decoding unit 224, an instruction issue unit 225, and an instruction execution unit 226.

[0094] The instruction fetch unit 223 is used to fetch the instruction to be executed from the memory 210 into the instruction register (which can be one of the register banks 229 shown Figure 3 for storing instructions) and receive the next fetch address or calculate the next fetch address according to the fetch algorithm. The fetch algorithm includes, for example: incrementing or decrementing the address according to the instruction length.

[0095] After fetching the instruction, the processing unit 220 enters the instruction decoding stage. The instruction decoding unit 224 decodes the fetched instruction according to a predetermined instruction format to obtain the operand acquisition information required for the fetched instruction, so as to prepare for the operation of the instruction execution unit 225. The operand acquisition information points to, for example, an immediate number, a register, or other software / hardware that can provide source operands.

[0096] The instruction issue unit 225 is located between the instruction decoding unit 224 and the instruction execution unit 226 and is used for instruction scheduling and control to efficiently allocate each instruction to different instruction execution units 226, making it possible to perform parallel operations on multiple instructions.

[0097] After the instruction issue unit 225 issues the instruction to the instruction execution unit 226, the instruction execution unit 226 starts to execute the instruction. However, if the instruction execution unit 226 determines that the instruction should be executed by the acceleration unit, it forwards it to the corresponding acceleration unit for execution. For example, if the instruction is a homomorphic encryption instruction to be executed, the instruction execution unit 226 no longer executes the instruction but sends the instruction to the acceleration unit 230 through the bus for execution by the acceleration unit 230.

[0098] In the prior art, the acceleration unit 230 for accelerating homomorphic encryption operations can adopt three methods: central processing unit (CPU), graphics processing unit (GPU), and field programmable gate array (FPGA). The solution using the CPU implements relatively complete operators in homomorphic encryption, including key generation, encryption and decryption, key exchange, modulus exchange, etc. In some scenarios, the multi-threading function provided by the CPU is also utilized. The solution using the GPU makes full use of the parallelism of the GPU for processing and accelerates the operators (such as number theoretic transform) that are easy to parallelize in homomorphic encryption. Experiments show that it can be accelerated by several times to dozens of times compared with CPU processing. The solution using the FPGA implements the algorithm of homomorphic encryption on the FPGA and makes full use of the characteristics of the hardware pipeline and high throughput. Their disadvantage is that whether it is the CPU, GPU, or FPGA, they are all designed for local algorithms, which is detrimental to the global performance. The hardware is implemented for proprietary algorithms and has poor generality.

[0099] The embodiment of the present disclosure adopts Figure 3 the structure of the acceleration unit 230 in [reference], which overcomes the above-mentioned disadvantages in the hardware implementation of the prior art, has strong generality, good global performance, and strong scalability.

[0100] As Figure 3 shown, the acceleration unit 230 includes an instruction buffer 231, an instruction fetch unit 232, an instruction decoding unit 233, a scheduler 234, a number theoretic transform unit 235, an arithmetic logic unit 236, a direct memory access unit 237, a memory interface 238, a shared buffer 239, an internal interconnection 240, etc. The number theoretic transform unit 235 can have one or more, and the arithmetic logic unit 236 can also have one or more.

[0101] The instruction execution unit 226 sends not only the homomorphic encryption instruction to be executed to the acceleration unit 230, but also the access memory address of the data required by the homomorphic encryption instruction to be executed stored in the memory 210. The instruction execution unit 226 can synthesize the homomorphic encryption instruction to be executed and the access memory address into a control signal and send it to the acceleration unit 230.

[0102] This control signal first enters the instruction buffer 231 of the acceleration unit 230. On the one hand, the instruction buffer 231 caches the homomorphic encryption instruction to be executed, and on the other hand, it sends the access memory address to the direct memory access unit 237. The direct memory access unit 237 receives the access memory address and instructs the memory interface 238 for data transmission with the memory 210 to fetch the data required by the homomorphic encryption instruction to be executed from the memory 210 according to the access memory address. In this way, when the number theoretic transform unit 235 and the arithmetic logic unit 236 actually execute the instruction, they can directly obtain the data from the direct memory access unit 237.

[0103] The instruction fetch unit 232 fetches the to-be-executed homomorphic encryption instruction from the instruction buffer 231 and sends it to the instruction decoding unit 233. The instruction decoding unit 233 decodes the to-be-executed homomorphic encryption instruction fetched by the instruction fetch unit 232 and sends the decoded to-be-executed homomorphic encryption instruction to the scheduler 234. Note that the instruction decoding unit 224 in the processing unit 220 has already decoded the to-be-executed homomorphic encryption instruction. The reason for re-decoding here is that the instruction set understandable by the acceleration unit 230 may not be the same as the instruction set understandable by the processing unit 220. Therefore, it is necessary to re-decode it into an instruction set understandable by the acceleration unit 230.

[0104] As described above, various operations of homomorphic encryption include modulo addition, modulo multiplication, number-theoretic transform / inverse number-theoretic transform, key exchange, modulo exchange, rescaling, etc. The concepts of modulo addition, modulo multiplication, number-theoretic transform / inverse number-theoretic transform, key exchange, modulo exchange, and rescaling have been described in the above glossary. For modulo addition and modulo multiplication operations, they are arithmetic logic operations that can be executed by at least one of one or more arithmetic logic units 236. For number-theoretic transform, it can be executed by at least one of one or more number-theoretic transform units 235. For inverse number-theoretic transform, it can be regarded as a combination of number-theoretic transform and modulo multiplication operation, where the number-theoretic transform is executed by at least one of one or more number-theoretic transform units 235, and the modulo multiplication operation is executed by at least one of one or more arithmetic logic units 236. For modulo exchange, it can be regarded as a combination of modulo addition and modulo multiplication, and is executed by at least one of one or more arithmetic logic units 236. For key exchange, it can be regarded as a combination of number-theoretic transform, inverse number-theoretic transform, modulo multiplication, and modulo exchange. The inverse number-theoretic transform and modulo exchange can be further decomposed and finally executed by at least one of one or more number-theoretic transform units 235 and one or more arithmetic logic units 236. For rescaling, it can be regarded as a combination of number-theoretic transform, inverse number-theoretic transform, and modulo exchange. The inverse number-theoretic transform and modulo exchange can be further decomposed and finally executed by at least one of one or more number-theoretic transform units 235 and one or more arithmetic logic units 236. As can be seen from the above, all operations included in the homomorphic encryption instruction to be executed can ultimately be decomposed into operations executed by at least one of one or more number-theoretic transform units 235 and one or more arithmetic logic units 236. The scheduler 234 is the unit that distributes each operation in the homomorphic encryption instruction to be executed to at least one of one or more number-theoretic transform units 235 and one or more arithmetic logic units 236. The specific approach is to divide the homomorphic encryption instruction to be executed into a combination of at least one of modulo multiplication operation, modulo addition operation, number-theoretic transform, inverse number-theoretic transform, modulo exchange, key exchange, and rescaling operation, and then decompose the modulo multiplication operation, modulo addition operation, number-theoretic transform, inverse number-theoretic transform, modulo exchange, key exchange, and rescaling operation according to the above rules and hand them over to be executed by at least one of one or more number-theoretic transform units 235 and one or more arithmetic logic units 236. The scheduler 234 stores the above decomposition rules internally.

[0105] Regardless of the specific homomorphic encryption algorithm, it can ultimately be decomposed into different combinations of number-theoretic transforms and arithmetic logic (modular addition, modular multiplication, and their combinations). In the embodiments of the present disclosure, a number-theoretic transform unit 235 is introduced to perform number-theoretic transforms, and an arithmetic logic unit 236 is introduced to perform arithmetic logic. Through the scheduling of the scheduler 234, several number-theoretic transform units 235 and several arithmetic logic units 236 can independently execute different tasks or form a pipeline to sequentially execute different stages of the same task. Therefore, the embodiments of the present disclosure can efficiently be compatible with different types of algorithms, improving global performance, scalability, and generality.

[0106] In addition to the instruction buffer 231 for storing the homomorphic encryption instructions to be executed as described above, the number-theoretic transform unit 235 and the arithmetic logic unit 236 may also include local buffers inside for storing the data required and the intermediate data generated during the number-theoretic transforms performed by the number-theoretic transform unit 235, as well as the data required and the intermediate data generated during the modular addition and modular multiplication operations performed by the arithmetic logic unit 236. When the data required and the intermediate data generated during number-theoretic transforms, modular addition, and modular multiplication cannot be stored in the local buffers, they can be stored in the shared buffer 239 shared by the number-theoretic transform unit 235 and the arithmetic logic unit 236. The shared buffer 239 communicates with each other through the internal interconnection 240 to achieve data sharing. Since the acceleration unit 230 in the embodiments of the present disclosure has a large number of local buffers, shared buffers 239, etc. inside, it avoids frequent access to the memory 210 outside the core, greatly improving the processing efficiency of homomorphic encryption and enhancing the computing performance.

[0107] Internal structure of the number theory transform unit

[0108] A number-theoretic transform can be regarded as a process of repeatedly processing the polynomial coefficients of a polynomial ring according to the same rule. For example, two polynomial coefficients at fixed positions in the polynomial ring are taken out, processed with a rotation factor, and a pair of new polynomial coefficients are obtained and placed back at two other fixed positions in the polynomial ring. For instance, the first polynomial coefficient and the fifth polynomial coefficient in the polynomial ring are taken out, processed with a rotation factor, and then placed back at the first polynomial coefficient position and the second polynomial coefficient position in the polynomial ring. When all the polynomial coefficients are processed in the above manner, a new polynomial ring is generated. Then, the above process is repeated, that is, the first polynomial coefficient and the fifth polynomial coefficient in the new polynomial ring are taken out, processed with a rotation factor, and then placed back at the first polynomial coefficient position and the second polynomial coefficient position in the polynomial ring. To repeatedly execute this operation, a butterfly process is created. As Figure 4As shown, a first polynomial coefficient storage subunit 2351 and a second polynomial coefficient storage subunit 2352 are provided in the number-theoretic transform unit 2354. The first polynomial coefficient storage subunit 2351 stores each polynomial coefficient in the initial polynomial ring. After taking out the polynomial coefficients from their predetermined positions, they are processed with rotation factors and placed in other predetermined positions of the second polynomial coefficient storage subunit 2352. In this way, new polynomial coefficients are obtained. Conversely, after taking out the polynomial coefficients from the predetermined positions of the second polynomial coefficient storage subunit 2352, they are processed with rotation factors and placed in other predetermined positions of the first polynomial coefficient storage subunit 2351. In this way, the swapping process is repeated until the predetermined requirements are met. At this time, the new polynomial coefficients stored in the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352 are the results of the number-theoretic transform. Each operation of taking polynomial coefficients, processing, and writing the results to different polynomial coefficient positions is called a butterfly operation, which is performed by the butterfly processing subunit 2354. The number-theoretic transform unit 235 further includes a rotation factor storage subunit 2353 for storing the rotation factors used in the above butterfly operation, which are set and stored in the rotation factor storage subunit 2353 in advance. The number-theoretic transform unit 235 may further include a control unit 2355 for controlling the operations of the first polynomial coefficient storage subunit 2351, the second polynomial coefficient storage subunit 2352, the rotation factor storage subunit 2353, and the butterfly processing subunit 2354.

[0109] The above butterfly operation can be divided into a first butterfly operation and a second butterfly operation.

[0110] The first butterfly operation includes: reading a first pair of polynomial coefficients from the first polynomial coefficient storage subunit 2351, obtaining the rotation factor corresponding to the first pair of polynomial coefficients from the rotation factor storage subunit 2353, obtaining a second pair of polynomial coefficients based on the first pair of polynomial coefficients and the rotation factor, and writing them into the second polynomial coefficient storage subunit 2352. In the above example, taking out the first polynomial coefficient and the fifth polynomial coefficient in the first polynomial coefficient storage subunit 2351 as the first pair of polynomial coefficients, processing them with rotation factors, and then putting them back into the first polynomial coefficient position and the second polynomial coefficient position in the second polynomial coefficient storage subunit 2352 belongs to the first butterfly operation.

[0111] The second butterfly process includes: reading a third polynomial coefficient pair from the second polynomial coefficient storage subunit 2352, obtaining the rotation factor corresponding to the third polynomial coefficient pair from the rotation factor storage subunit 2353, obtaining a fourth polynomial coefficient pair based on the third polynomial coefficient pair and the rotation factor, and writing it into the first polynomial coefficient storage subunit 2351. In the above example, taking out the first polynomial coefficient and the fifth polynomial coefficient in the second polynomial coefficient storage subunit 2352 as the third polynomial coefficient pair, after being processed by the rotation factor, it becomes the fourth polynomial coefficient pair and is put back to the positions of the first polynomial coefficient and the second polynomial coefficient in the first polynomial coefficient storage subunit 2351, which belongs to the second butterfly process.

[0112] The first polynomial coefficient storage subunit 2351 and the second polynomial coefficient storage subunit 2352 each include a plurality of memory banks 23511, and each memory bank 23511 sequentially accommodates a storage queue. As Figure 8In the example shown, each of the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352 includes 8 memory banks 23511, with indexes B1 - 8 respectively. Each memory bank stores a queue of 8 polynomial coefficients. That is, when the ciphertext represents a polynomial of 64 polynomial coefficients, the 64 polynomial coefficients can be distributed among 8 memory banks, with 8 polynomial coefficients stored in each memory bank. Polynomial coefficients C0 - 7 are placed in memory bank B1; polynomial coefficients C8 - 15 are placed in memory bank B2; polynomial coefficients C16 - 23 are placed in memory bank B3... Polynomial coefficients C56 - 63 are placed in memory bank B8. Generally, the number of polynomial coefficients in the polynomial ring and the number of memory banks can be set to a positive integer power of 2. Assume that each of the memory banks in the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352 has M, such as 8 in the above example, which is equal to 2 to the 3rd power. This is beneficial because after multiple butterfly operations, the order of the polynomial coefficients arranged in each memory bank can return to the initial state, thus ending the butterfly operation. When fetching the first polynomial coefficient pair and the third polynomial coefficient pair, try to make the two coefficients in the coefficient pair come from memory banks that are far apart and maintain an equal separation distance. For example, divide the M memory banks into the first M / 2 memory banks and the last M / 2 memory banks, and take polynomial coefficients from the first memory bank of the first M / 2 memory banks and the first memory bank of the last M / 2 memory banks to form a first polynomial coefficient pair, then take polynomial coefficients from the second memory bank of the first M / 2 memory banks and the second memory bank of the last M / 2 memory banks to form a first polynomial coefficient pair, and so on... In this way, the indexes of the memory banks from which the two polynomial coefficients in the fetched polynomial coefficient pair are taken always differ by M / 2. And the indexes of the pair of memory banks from which the first polynomial coefficient pair comes can be the same as the indexes of the pair of memory banks from which the third polynomial coefficient pair comes. For example, when fetching the first polynomial coefficient pair, take polynomial coefficients from the first and fifth memory banks; when fetching the corresponding third polynomial coefficient pair, also take polynomial coefficients from the first and fifth memory banks. Additionally, the indexes of the pair of memory banks from which the second polynomial coefficient pair comes are the same as the indexes of the pair of memory banks from which the fourth polynomial coefficient pair comes, and the indexes of the two memory banks in this pair are adjacent. For example, after forming the second polynomial coefficient pair, place it in the first and second memory banks; after forming the corresponding fourth polynomial coefficient pair, also place it in the first and second memory banks. Perform the second butterfly operation after the first butterfly operation, then the first butterfly operation again... and so on, until the memory banks in which the polynomial coefficients fetched for the butterfly operation are arranged return to the memory banks where they were initially located.

[0113] In Figure 8For example, the initial sorting of the 8 memory banks in the first polynomial coefficient storage subunit 2351 is B1 - 8. That is, M = 8. There are 4 butterfly processing subunits 2354. The first butterfly processing subunit 2354 fetches the polynomial coefficients in memory bank B1 and memory bank B5 in the first polynomial coefficient storage subunit 2351 as the first polynomial coefficient pair, processes them with the rotation factor to obtain the second polynomial coefficient pair, and then puts them back into memory banks B1 and B2 in the second polynomial coefficient storage subunit 2352. At this time, the polynomial coefficients placed in memory banks B1 and B2 in the second polynomial coefficient storage subunit 2352 are from memory banks B1 and B5 in the first polynomial coefficient storage subunit 2351. Similarly, the second butterfly processing subunit 2354 fetches the polynomial coefficients in memory bank B2 and memory bank B6 in the first polynomial coefficient storage subunit 2351 as the first polynomial coefficient pair, processes them with the rotation factor to obtain the second polynomial coefficient pair, and then puts them back into memory banks B3 and B4 in the second polynomial coefficient storage subunit 2352. At this time, the polynomial coefficients placed in memory banks B3 and B4 in the second polynomial coefficient storage subunit 2352 are from memory banks B2 and B6 in the first polynomial coefficient storage subunit 2351. By analogy, after the first butterfly processing, the polynomial coefficients placed in memory banks B5 and B6 in the second polynomial coefficient storage subunit 2352 are from memory banks B3 and B7 in the first polynomial coefficient storage subunit 2351, and the polynomial coefficients placed in memory banks B7 and B8 in the second polynomial coefficient storage subunit 2352 are from memory banks B4 and B8 in the first polynomial coefficient storage subunit 2351. Therefore, after the first butterfly processing, the contents stored in memory banks B1 - B8 in the second polynomial coefficient storage subunit 2352 are substantially equivalent to those in memory banks B1, B5, B2, B6, B3, B7, B4, B8 in the original first polynomial coefficient storage subunit 2351 respectively.

[0114] In the second butterfly process, the first butterfly process sub-unit 2354 extracts the polynomial coefficients stored in the first and fifth memory banks of the second polynomial coefficient storage sub-unit 2352. In fact, it extracts the polynomial coefficients in the original memory banks B1 and B3, uses them as the third polynomial coefficient pair, processes them with the rotation factor to obtain the fourth polynomial coefficient pair, and then puts them back into the first and second memory banks of the first polynomial coefficient storage sub-unit 2351. At this time, the polynomial coefficients placed in the memory banks B1 and B2 of the first polynomial coefficient storage sub-unit 2351 are substantially the contents of the original memory banks B1 and B3. Similarly, after being processed by the second to fourth butterfly process sub-units 2354, the polynomial coefficients placed in the memory banks B3 and B4 of the first polynomial coefficient storage sub-unit 2351 are substantially the contents of the original memory banks B5 and B7; the polynomial coefficients placed in the memory banks B5 and B6 of the first polynomial coefficient storage sub-unit 2351 are substantially the contents of the original memory banks B2 and B4; the polynomial coefficients placed in the memory banks B7 and B8 of the first polynomial coefficient storage sub-unit 2351 are substantially the contents of the original memory banks B6 and B8. After the second butterfly process, the contents stored in the memory banks B1 - B8 of the first polynomial coefficient storage sub-unit 2351 are substantially equivalent to the original memory banks B1, B3, B5, B7, B2, B4, B6, B8 respectively.

[0115] Then, after another round of the first butterfly process, the contents stored in the memory banks B1 - B8 of the second polynomial coefficient storage sub-unit 2352 are substantially equivalent to the memory banks B1, B2, B3, B4, B5, B6, B7, B8 of the original first polynomial coefficient storage sub-unit 2351, that is, in the same order as the original memory banks. At this time, according to the general butterfly process principle, after the above-mentioned log 2 M times of butterfly processes, the processing can be stopped, and the polynomial coefficients stored in these memory banks become the processing results of the number-theoretic transform.

[0116] It can be seen that in the above process, for a single butterfly processing subunit 2354 (the first, second, third, or fourth butterfly processing subunit 2354), the memory bank indices for fetching the first polynomial coefficient pairs from the first polynomial coefficient storage subunit 2351 in the first butterfly processing (such as B1 and B5 of the first butterfly processing subunit, B2 and B6 of the second butterfly processing subunit, B3 and B7 of the third butterfly processing subunit, B4 and B8 of the fourth butterfly processing subunit) are the same as those for fetching the third polynomial coefficient pairs from the second polynomial coefficient storage subunit 2352 in the second butterfly processing. This way of consistent reading and writing can relieve the pressure of layout and wiring. Additionally, the memory bank indices they are fetched from differ by M / 2 (such as a difference of 4 between 1 and 5, 2 and 6, 3 and 7, 4 and 8). The memory bank indices in the second polynomial coefficient storage subunit 2352 where the single butterfly processing subunit 2354 puts the generated second polynomial coefficient pairs in the first butterfly processing (such as B1 and B2 of the first butterfly processing subunit, B3 and B4 of the second butterfly processing subunit, B5 and B6 of the third butterfly processing subunit, B7 and B8 of the fourth butterfly processing subunit) are the same as those in the first polynomial coefficient storage subunit 2351 where the generated fourth polynomial coefficient pairs are put in the second butterfly processing, to relieve the pressure of layout and wiring. Additionally, the memory bank indices they are put into are adjacent. Only in this way can the order of the memory banks from which the polynomial coefficients stored in each memory bank are derived return to the initial state after several butterfly processings, such as B1, B2, B3, B4, B5, B6, B7, B8 in the above example, and it returns to the initial state after log 2 M butterfly processings, thus reaching the termination condition of number-theoretic transform in the general sense.

[0117] What is mentioned above is only the termination condition of number-theoretic transform in the general sense. After the termination condition of number-theoretic transform in the general sense is reached in the embodiments of the present disclosure, the control unit 2355 transposes the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352, and then the butterfly processing subunit 2354 performs log 2 M times of the first butterfly processing or the second butterfly processing, that is, the above-mentioned termination condition of number-theoretic transform is reached again on the basis of transposition.

[0118] The transposition includes: taking out the polynomial coefficients queued with the same serial number in each memory bank of the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352 before transposition, and placing them in a memory bank 23511 after transposition in the order of the memory bank index. That is, the rows and columns of the array formed by each polynomial coefficient in the first polynomial coefficient storage subunit 2351 or the second polynomial coefficient storage subunit 2352 are reversed. The originally formed columns (memory bank 23511) serve as the rows of the new array (polynomial coefficients with the same serial number in each memory bank 23511), and the originally formed rows (polynomial coefficients with the same serial number in each memory bank 23511) serve as the columns of the new array (memory bank 23511). As Figure 8 shown, the polynomial coefficients C0, C8, C16, C24, C32, C40, C48, C56 queued in the first place in the memory banks B1-8 of the original first polynomial coefficient storage subunit 2351 are placed in the memory bank B1 after transposition in ascending order of the index; the polynomial coefficients C1, C9, C17, C25, C33, C41, C49, C57 queued in the second place in the memory banks B1-8 of the original first polynomial coefficient storage subunit 2351 are placed in the memory bank B2 after transposition in ascending order of the index; the polynomial coefficients C2, C10, C18, C26, C34, C42, C50, C58 queued in the third place in the memory banks B1-8 of the original first polynomial coefficient storage subunit 2351 are placed in the memory bank B3 after transposition in ascending order of the index... The polynomial coefficients C7, C15, C23, C31, C39, C47, C55, C63 queued in the last place in the memory banks B1-8 of the original first polynomial coefficient storage subunit 2351 are placed in the memory bank B8 after transposition in ascending order of the index.

[0119] After the transposition is completed, the butterfly processing subunit 2354 performs the first butterfly processing or the second butterfly processing for log 2 M times. Since this process is exactly the same as the processing process before transposition, for the sake of saving space, it will not be elaborated here.

[0120] The role of transposition in the embodiments of the present disclosure is as follows: Without transposition, the butterfly processing sub-unit 2354 always performs butterfly processing on the polynomial coefficients 235111 in two different memory banks 23511. However, in practice, it is sometimes necessary to perform butterfly processing on different polynomial coefficients 235111 in the same memory bank 23511. Therefore, transposition is introduced. Before transposition, the two polynomial coefficients for which butterfly processing is performed come from different memory banks 23511. After transposition, the two polynomial coefficients for which butterfly processing is performed come from the lower and upper bits of the same original memory bank 23511. For example, there are 8 polynomial coefficients 235111 queued in the memory bank 23511, the first four polynomial coefficients 235111 are the lower bits, and the last four polynomial coefficients are the upper bits. In this way, the polynomial coefficients can be read and written back and forth between the upper and lower bits of the same memory bank, enriching the applicable range of butterfly processing.

[0121] Internal structure of the butterfly processing subunit

[0122] As Figure 5 shown, the butterfly processing sub-unit 2354 according to an embodiment of the present disclosure includes a first multiplexer 23541, a second multiplexer 23542, a third multiplexer 23543, a fourth multiplexer 23544, a first adder 23545, a second adder 23546, a first subtractor 23547, a second subtractor 23548, and a first multiplier 23549. Figure 3 Each multiplexer in has a first input terminal (0 input terminal), a second input terminal (1 input terminal), a control terminal, and an output terminal. The control terminal is connected to the control signal SEL or the negation of SEL. When SEL is set to 1, the second input terminal (1 input terminal) is turned on, and its signal directly enters the output terminal. When SEL is set to 0, the first input terminal (0 input terminal) is turned on, and its signal directly enters the output terminal. By determining whether the signal connected to the control terminal is set to 0 or 1, it is determined which of the first input terminal and the second input terminal is output.

[0123] As Figure 5As shown, the first coefficient 301 in the first polynomial coefficient pair or the third polynomial coefficient pair is input to the first input terminal (0 input terminal) of the first multiplexer 23541, and after being added to the second coefficient 302 in the first polynomial coefficient pair or the third polynomial coefficient pair by the first adder 23545, it is input to the second input terminal (1 input terminal) of the first multiplexer 23541. The first input terminal and the second input terminal are selected by the first selection signal (inverted SEL) of the first multiplexer 23541 to be input to the output terminal through one of them. The second coefficient 302 is input to the second input terminal (1 input terminal) of the third multiplexer 23543, and after being subtracted from the first coefficient 301 by the first subtractor 23547, it is input to the first input terminal (0 input terminal) of the third multiplexer 23543. The first input terminal and the second input terminal are selected by the third selection signal (inverted 11SEL) of the third multiplexer 23543 to be input to the output terminal through one of them, and then multiplied by the corresponding rotation factor through the first multiplier 23549 to obtain a product signal.

[0124] The signal output from the output terminal of the first multiplexer 23541 is input to the first input terminal (0 input terminal) of the second multiplexer 23546, and after being added to the product signal by the second adder 23546, it is input to the second input terminal (1 input terminal) of the second multiplexer 23542. The first input terminal and the second input terminal are selected by the second selection signal (SEL) of the second multiplexer 23542 to be input to the output terminal through one of them, serving as one coefficient 303 in the second polynomial coefficient pair or the fourth polynomial coefficient pair. The product signal is input to the second input terminal (1 input terminal) of the fourth multiplexer 23544, and after being subtracted from the signal output from the first multiplexer by the second subtractor 23548, it is input to the first input terminal (0 input terminal) of the fourth multiplexer 23544. The first input terminal and the second input terminal are selected by the fourth selection signal (SEL) of the fourth multiplexer 23544 to be input to the output terminal through one of them, serving as the other coefficient 304 in the second polynomial coefficient pair or the fourth polynomial coefficient pair.

[0125] In the above structure, if SEL = 1, the non of SEL is 0, the first inputs (0 inputs) of the first multiplexer 23541 and the third multiplexer 23543 are turned on. The signal output by the first multiplexer 23541 is the first coefficient 301, and the signal output by the third multiplexer 23543 is (the second coefficient 302 - the first coefficient 301). After passing through the first multiplier 23549, the product signal = (the second coefficient 302 - the first coefficient 301) × rotation factor. Since SEL = 1, the second inputs (1 inputs) of the second multiplexer 23542 and the fourth multiplexer 23544 are turned on. The signal output by the second multiplexer 23542 is the first coefficient 301 + the product signal = the first coefficient 301 + (the second coefficient 302 - the first coefficient 301) × rotation factor = the first coefficient 301 × (1 - rotation factor) + the second coefficient 302 × rotation factor, and the signal output by the fourth multiplexer 23544 is the product signal = (the second coefficient 302 - the first coefficient 301) × rotation factor. The above formulas of the outputs 303 and 304 exactly match the requirements of the number-theoretic transform, that is, when SEL = 1, the above structure can be used for the number-theoretic transform.

[0126] In the above structure, if SEL = 0, the non of SEL is 1, the second inputs (1 inputs) of the first multiplexer 23541 and the third multiplexer 23543 are turned on. The signal output by the first multiplexer 23541 is (the first coefficient 301 + the second coefficient 302), and the signal output by the third multiplexer 23543 is the second coefficient 302. After passing through the first multiplier 23549, the product signal = the second coefficient 302 × rotation factor. Since the non of SEL is 0, the first inputs (0 inputs) of the second multiplexer 23542 and the fourth multiplexer 23544 are turned on. The signal output by the second multiplexer 23542 is (the first coefficient 301 + the second coefficient 302), and the signal output by the fourth multiplexer 23544 is the product signal - (the first coefficient 301 + the second coefficient 302) = the second coefficient 302 × (rotation factor - 1) - the first coefficient 301. The above formulas of the outputs 303 and 304 exactly match the requirements of the inverse number-theoretic transform, that is, when SEL = 0, the above structure can be used for the inverse number-theoretic transform.

[0127] Through the above embodiments, the number-theoretic transform and the inverse number-theoretic transform are realized using a simple structure, improving the implementation efficiency of the number-theoretic transform and the inverse number-theoretic transform.

[0128] The structure of the above butterfly processing subunit 2354 is only an example, and there can be other structures that can realize the number-theoretic transform and the inverse number-theoretic transform as described above.

[0129] Internal structure of the arithmetic logic unit

[0130] As Figure 6 shown, an arithmetic logic unit 236 of a vector according to the present disclosure includes a modulo adder 2363, a modulo multiplier 2364, a fifth multiplexer 2361, a sixth multiplexer 2362, and a seventh multiplexer 2365. Each of the above multiplexers has a first input terminal (0 input terminal), a second input terminal (1 input terminal), a control terminal, and an output terminal. The control terminal is connected to control signals a, b, or c. When the control signal a, b, or c is set to 1, the corresponding second input terminal (1 input terminal) is turned on, and its signal directly enters the output terminal. When the control signal a, b, or c is set to 0, the first input terminal (0 input terminal) is turned on, and its signal directly enters the output terminal. By determining whether the signal connected to the control terminal is set to 0 or 1, it is determined which one of the first input terminal and the second input terminal is output.

[0131] As Figure 6 shown, a first input signal 305 is input to the first input terminals (0 input terminals) of the fifth multiplexer 2361 and the sixth multiplexer 2362. The output of the modulo adder 263 is input to the second input terminal (1 input terminal) of the sixth multiplexer 2362. The output of the modulo multiplier 2364 is input to the second input terminal (1 input terminal) of the fifth multiplexer 2361. One of the first input terminal (0 input terminal) and the second input terminal (1 input terminal) of the fifth multiplexer 2361 is selected by the fifth selection signal a of the fifth multiplexer 2361 to the output terminal, and one of the first input terminal (0 input terminal) and the second input terminal (1 input terminal) of the sixth multiplexer 2362 is selected by the sixth selection signal b of the sixth multiplexer 2362 to the output terminal. The output terminal of the fifth multiplexer 2361 is connected to the first input terminal of the modulo adder 2363, and the second input signal 306 is input to the second input terminal of the modulo adder 2363. The modulo adder 2363 performs modulo addition on the input signal at the first input terminal and the input signal at the second input terminal, and the obtained result is output from its output terminal. The output terminal of the sixth multiplexer 2362 is connected to the first input terminal of the modulo multiplier 2364, and the third input signal 307 is input to the second input terminal of the modulo multiplier 2364. The output terminal of the modulo adder 2363 is connected to the first input terminal (0 input terminal) of the seventh multiplexer 2365, the output terminal of the modulo multiplier 2364 is connected to the second input terminal (1 input terminal) of the seventh multiplexer 2365, and one of the first input terminal (0 input terminal) and the second input terminal (1 input terminal) of the seventh multiplexer 2365 is selected by the seventh selection signal c of the seventh multiplexer 2365 to the output terminal 308.

[0132] Through the above circuit structure, by setting the fifth, sixth, and seventh selection signals a, b, and c of the fifth, sixth, and seventh multiplexers 2361, 2362, and 2365 differently, Figure 6The arithmetic logic unit 236 can perform operations such as modulo addition, modulo multiplication, and their different combinations.

[0133] As Figure 6 shown, when the fifth strobe signal a is set to 0, that is, its first input terminal (0 input terminal) is strobed, and the seventh strobe signal c is set to 0, that is, its first input terminal (0 input terminal) is strobed, the output 308 of the seventh multiplexer 2365 is equal to the input of its first input terminal (0 input terminal), and this input is the result of modulo addition by the modulo adder 2363. The two inputs of the modulo adder 2363 are respectively the output of the fifth multiplexer 2361 and the second input 306. The output of the fifth multiplexer 2361 is equal to the first input 305 of the first input terminal of the fifth multiplexer 2361 because its fifth strobe signal a is set to 0. In this way, the output of the modulo adder 2363 is equal to the first input 305 + the second input 306, and thus the output of the arithmetic logic unit 236 is also equal to the first input 305 + the second input 306. In this way, by setting the fifth strobe signal a to 0 and setting the seventh strobe signal c to 0, the arithmetic logic unit 236 completes the function of modulo addition, as Figure 7 shown.

[0134] As Figure 6 shown, when the sixth strobe signal b is set to strobe the first input terminal (0 input terminal) and the seventh strobe signal c is set to strobe the second input terminal (1 input terminal), the output 308 of the seventh multiplexer 2365 is equal to the input of its second input terminal (1 input terminal). This input is the result of modulo multiplication by the modulo multiplier 2364. The two inputs of the modulo multiplier 2364 are respectively the output of the sixth multiplexer 2362 and the third input 307. The output of the sixth multiplexer 2362 is equal to the first input 305 of the first input terminal (0 input terminal) of the sixth multiplexer 2362 because its sixth strobe signal b is set to 0. In this way, the output of the modulo multiplier 2364 is equal to the first input 305 × the third input 307, and thus the output of the arithmetic logic unit 236 is also equal to the first input 305 × the third input 307. In this way, by setting the sixth strobe signal b to 0 and setting the seventh strobe signal c to 1, the arithmetic logic unit 236 completes the function of modulo multiplication, as Figure 7 shown.

[0135] As Figure 6As shown, when the fifth strobe signal a is set to strobe the second input terminal (1 input terminal), the sixth strobe signal b is set to strobe the first input terminal (0 input terminal), and the seventh strobe signal c is set to strobe the first input terminal (0 input terminal), the output 308 of the seventh multiplexer 2365 is equal to the input of its first input terminal (0 input terminal). This input is the result of modulo addition by the modulo adder 2363. The two inputs of the modulo adder 2363 are respectively the output of the fifth multiplexer 2361 and the second input 306. Therefore, the output 308 of the arithmetic logic unit 236 = the output of the fifth multiplexer 2361 + the second input 306. Since the fifth strobe signal a is set to strobe the second input terminal (1 input terminal), the output of the fifth multiplexer 2361 is equal to the input of its second input terminal (1 input terminal), which is the output of the multiplier 2364. The output of the multiplier 2364 is equal to the output of the sixth multiplexer 2362 × the third input 307. Since the sixth strobe signal b is set to strobe the first input terminal (0 input terminal), the output of the sixth multiplexer 2362 is equal to the input of its first input terminal (0 input terminal), that is, the first input 305. Therefore, the output of the fifth multiplexer 2361 is equal to the first input 305 × the third input 307. Thus, the output 308 = the first input 305 × the third input 307 + the second input 306. The arithmetic logic unit 236 performs the function of first modulo multiplication and then modulo addition, as Figure 7 shown.

[0136] As Figure 6As shown, when the fifth strobe signal a is set to strobe the first input terminal (0 input terminal), the sixth strobe signal b is set to strobe the second input terminal (1 input terminal), and the seventh strobe signal c is set to strobe the second input terminal (1 input terminal), the output 308 of the seventh multiplexer 2365 is equal to the input of its second input terminal (1 input terminal). This input is the result of the modular multiplication by the multiplier 2364. The two inputs of the multiplier 2364 are respectively the output of the sixth multiplexer 2362 and the third input 307. Therefore, the output 308 of the arithmetic logic unit 236 = the output of the sixth multiplexer 2362 × the third input 307. Since the sixth strobe signal b is set to strobe the second input terminal (1 input terminal), the output of the sixth multiplexer 2362 is equal to the input of its second input terminal (1 input terminal), and this input is the output of the adder 2363. The output of the adder 2363 is equal to the output of the fifth multiplexer 2361 + the second input 306. Since the fifth strobe signal a is set to strobe the first input terminal (0 input terminal), the output of the fifth multiplexer 2361 is equal to the input of its first input terminal (0 input terminal), that is, the first input 305. Therefore, the output of the sixth multiplexer 2362 is equal to the first input 305 + the second input 306. In this way, the output 308 = (the first input 305 + the second input 306) × the third input 307. The arithmetic logic unit 236 completes the function of modular addition first and then modular multiplication, as Figure 7 shown.

[0137] As Figure 6 shown, when the fifth strobe signal a is set to strobe the second input terminal (1 input terminal), the sixth strobe signal b is set to strobe the second input terminal (1 input terminal), and the seventh strobe signal c is set to strobe the first input terminal (0 input terminal), the circuit cannot complete any operation, that is, it is invalid, as Figure 7 shown. Similarly, when the fifth strobe signal a is set to strobe the second input terminal (1 input terminal), the sixth strobe signal b is set to strobe the second input terminal (1 input terminal), and the seventh strobe signal c is set to strobe the second input terminal (1 input terminal), the circuit also cannot complete any operation, that is, it is invalid, as Figure 7 shown.

[0138] Through the above simple circuit and different setting combinations of the strobe signals a, b, and c, the arithmetic logic unit 236 can realize different combinations of modular addition and modular multiplication operations, achieving the effect of completing multiple arithmetic logic operations with the same circuit and improving the efficiency of the circuit.

[0139] Flow of the deep neural network operation method according to an embodiment of the present disclosure

[0140] As Figure 9 shown, according to an embodiment of the present disclosure, a homomorphic encryption method is provided, including:

[0141] Step 410: Receive a homomorphic encryption instruction to be executed;

[0142] Step 420: Decompose the homomorphic encryption instruction to be executed into operations;

[0143] Step 430: Allocate the number-theoretic transforms included in the operations to at least one of one or more number-theoretic transform units for execution, and allocate the arithmetic operations included in the operations to at least one of one or more arithmetic logic units for execution.

[0144] Since the implementation details of the above process have been introduced in detail in the description of the foregoing apparatus embodiments, they will not be elaborated herein.

[0145] Commercial value of the embodiments of the present disclosure

[0146] Verified by experiments, the embodiments of the present disclosure propose a general, modular, and extensible homomorphic encryption accelerator architecture, which greatly reduces the deployment cost of homomorphic encryption algorithms and the redeployment cost caused by algorithm changes later, and can be reduced to 50%-80% of the original, having good market prospects.

[0147] It should be understood that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for method embodiments, since they are basically similar to the methods described in apparatus and system embodiments, the description is relatively simple, and the relevant parts can refer to the partial descriptions of other embodiments.

[0148] It should be understood that the specific embodiments of this specification have been described above. Other embodiments are within the scope of the claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0149] It should be understood that an element described herein in the singular form or shown in the drawings as only one does not represent limiting the quantity of the element to one. In addition, a module or element described or shown herein as separate can be combined into a single module or element, and a module or element described or shown herein as a single one can be split into multiple modules or elements.

[0150] It should also be understood that the terminology and expressions employed herein are for the purpose of description only and that one or more embodiments of this specification should not be limited to these terms and expressions. The use of these terms and expressions does not imply the exclusion of any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.

Claims

1. An acceleration unit, comprising: one or more number theory transform units for performing number theory transforms in a homomorphic encryption process; one or more arithmetic logic units for performing arithmetic operations in a homomorphic encryption process; a scheduler for allocating operations in a to-be-executed homomorphic encryption instruction to at least one of the one or more number theory transform units and the one or more arithmetic logic units, wherein the scheduler divides the to-be-executed homomorphic encryption instruction into at least one of the following operations: modular multiplication operation, modular addition operation, number theory transform, inverse number theory transform, modular swap, key exchange, rescaling; an instruction buffer for receiving control signals, the control signals including the to-be-executed homomorphic encryption instruction and an access memory address; an instruction fetch unit for fetching the to-be-executed homomorphic encryption instruction from the instruction buffer; an instruction decoding unit for decoding the to-be-executed homomorphic encryption instruction fetched by the instruction fetch unit and sending the decoded to-be-executed homomorphic encryption instruction to the scheduler; a memory interface for data transmission with a memory; a direct memory access unit for receiving the access memory address sent by the instruction buffer and instructing the memory interface to fetch data required by the to-be-executed homomorphic encryption instruction according to the access memory address.

2. The acceleration unit according to claim 1, wherein, for the modular multiplication operation or the modular addition operation, the scheduler allocates to at least one of the one or more arithmetic logic units.

3. The acceleration unit according to claim 1, wherein, for the number theory transform, the scheduler allocates to at least one of the one or more number theory transform units.

4. The acceleration unit according to claim 1, wherein, for the inverse number theory transform, the scheduler decomposes the inverse number theory transform into a combination of a number theory transform and a modular multiplication operation, allocates the decomposed number theory transform to at least one of the one or more number theory transform units, and allocates the decomposed modular multiplication operation to at least one of the one or more arithmetic logic units.

5. The acceleration unit according to claim 4, wherein, for the modular swap, the scheduler decomposes the modular swap into a combination of modular addition and modular multiplication and allocates to at least one of the one or more arithmetic logic units.

6. The acceleration unit according to claim 5, wherein, for the key exchange, the scheduler decomposes the key exchange into a number theory transform, an inverse number theory transform, a modular multiplication, a modular swap, and allocates according to the decomposition result to at least one of the one or more number theory transform units or the one or more arithmetic logic units.

7. The acceleration unit according to claim 5, wherein, for the rescaling, the scheduler decomposes the rescaling into a number theory transform, an inverse number theory transform, a modular swap, and allocates according to the decomposition result to at least one of the one or more number theory transform units or the one or more arithmetic logic units.

8. The acceleration unit according to claim 1, wherein, at least one of the one or more number theory transform units includes: a first polynomial coefficient storage subunit; The second polynomial coefficient storage subunit; The rotation factor storage subunit, which is used to store the rotation factors in the number-theoretic transform; The butterfly processing subunit, which is used to perform the first butterfly processing and the second butterfly processing. Among them, the first butterfly processing includes: reading a first pair of polynomial coefficients from the first polynomial coefficient storage subunit, obtaining the rotation factor corresponding to the first pair of polynomial coefficients from the rotation factor storage subunit, obtaining a second pair of polynomial coefficients based on the first pair of polynomial coefficients and the rotation factor, and writing the second pair of polynomial coefficients into the second polynomial coefficient storage subunit; the second butterfly processing includes: reading a third pair of polynomial coefficients from the second polynomial coefficient storage subunit, obtaining the rotation factor corresponding to the third pair of polynomial coefficients from the rotation factor storage subunit, obtaining a fourth pair of polynomial coefficients based on the third pair of polynomial coefficients and the rotation factor, and writing the fourth pair of polynomial coefficients into the first polynomial coefficient storage subunit.

9. The acceleration unit according to claim 8, wherein, at least one of the one or more number-theoretic transform units further includes: a control unit, which is used to control the operations of the first polynomial coefficient storage subunit, the second polynomial coefficient storage subunit, the rotation factor storage subunit, and the butterfly processing subunit.

10. The acceleration unit according to claim 9, wherein, both the first polynomial coefficient storage subunit and the second polynomial coefficient storage subunit include a plurality of memory banks, each memory bank has a corresponding index, the first pair of polynomial coefficients and the third pair of polynomial coefficients come from the memory banks with the same index, and the second pair of polynomial coefficients and the fourth pair of polynomial coefficients come from the memory banks with the same index.

11. The acceleration unit according to claim 10, wherein, there are M memory banks in the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit, and there are M / 2 butterfly processing subunits. And for a single butterfly processing subunit, the indexes of the pair of memory banks from which the first pair of polynomial coefficients come are the same as the indexes of the pair of memory banks from which the third pair of polynomial coefficients come, and the indexes of the two memory banks in this pair of memory banks differ by M / 2; the indexes of the pair of memory banks from which the second pair of polynomial coefficients come are the same as the indexes of the pair of memory banks from which the fourth pair of polynomial coefficients come, and the indexes of the two memory banks in this pair of memory banks are adjacent.

12. The acceleration unit according to claim 11, wherein, After the first butterfly processing or the second butterfly processing of log 2 for M times, the control unit transposes the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit, and then the butterfly processing subunit performs the first butterfly processing or the second butterfly processing of log 2 for M, wherein the transposition includes: taking out the polynomial coefficients queued with the same serial number in each memory bank of the first polynomial coefficient storage subunit or the second polynomial coefficient storage subunit before transposition, and placing them in a memory bank after transposition in the order of memory bank indexes.

13. The acceleration unit according to claim 9, wherein, the butterfly processing subunit includes a first multiplexer, a second multiplexer, a third multiplexer, a fourth multiplexer, a first adder, a second adder, a first subtractor, a second subtractor, and a first multiplier.

14. The acceleration unit according to claim 13, wherein, The first coefficient in the first polynomial coefficient pair or the third polynomial coefficient pair is input to the first input terminal of the first multiplexer, and after being added to the second coefficient in the first polynomial coefficient pair or the third polynomial coefficient pair by a first adder, it is input to the second input terminal of the first multiplexer. One of the first input terminal and the second input terminal is selected by the first selection signal of the first multiplexer and output to the output terminal; The second coefficient is input to the second input terminal of the third multiplexer, and after being subtracted from the first coefficient by a first subtractor, it is input to the first input terminal of the third multiplexer. One of the first input terminal and the second input terminal is selected by the third selection signal of the third multiplexer and output to the output terminal, and then multiplied by a corresponding rotation factor by a first multiplier to obtain a product signal.

15. The acceleration unit according to claim 14, wherein, The signal output from the output terminal of the first multiplexer is input to the first input terminal of the second multiplexer, and after being added to the product signal by a second adder, it is input to the second input terminal of the second multiplexer. One of the first input terminal and the second input terminal is selected by the second selection signal of the second multiplexer and output to the output terminal, serving as one coefficient in the second polynomial coefficient pair or the fourth polynomial coefficient pair; the product signal is input to the second input terminal of the fourth multiplexer, and after being subtracted from the signal output from the first multiplexer by a second subtractor, it is input to the first input terminal of the fourth multiplexer. One of the first input terminal and the second input terminal is selected by the fourth selection signal of the fourth multiplexer and output to the output terminal, serving as the other coefficient in the second polynomial coefficient pair or the fourth polynomial coefficient pair.

16. The acceleration unit according to claim 1, wherein, The arithmetic logic unit includes a modulo adder, a modulo multiplier, a fifth multiplexer, a sixth multiplexer, and a seventh multiplexer. Among them, a first input signal is input to the first input terminals of the fifth multiplexer and the sixth multiplexer, the output of the modulo adder is input to the second input terminal of the sixth multiplexer, the output of the modulo multiplier is input to the second input terminal of the fifth multiplexer. One of the first input terminal and the second input terminal of the fifth multiplexer is selected by the fifth selection signal of the fifth multiplexer and output to the output terminal, and one of the first input terminal and the second input terminal of the sixth multiplexer is selected by the sixth selection signal of the sixth multiplexer and output to the output terminal. The output terminal of the fifth multiplexer is connected to the first input terminal of the modulo adder, the second input signal is input to the second input terminal of the modulo adder, the output terminal of the sixth multiplexer is connected to the first input terminal of the modulo multiplier, the third input signal is input to the second input terminal of the modulo multiplier, the output terminal of the modulo adder is connected to the first input terminal of the seventh multiplexer, the output terminal of the modulo multiplier is connected to the second input terminal of the seventh multiplexer. One of the first input terminal and the second input terminal of the seventh multiplexer is selected by the seventh selection signal of the seventh multiplexer and output to the output terminal.

17. The acceleration unit according to claim 16, wherein, by setting the fifth strobe signal to strobe the first input terminal and setting the seventh strobe signal to strobe the first input terminal, the arithmetic logic unit is used for modular addition operation.

18. The acceleration unit according to claim 16, wherein, by setting the sixth strobe signal to strobe the first input terminal and setting the seventh strobe signal to strobe the second input terminal, the arithmetic logic unit is used for modular multiplication operation.

19. The acceleration unit according to claim 16, wherein, by setting the fifth strobe signal to strobe the second input terminal, setting the sixth strobe signal to strobe the first input terminal, and setting the seventh strobe signal to strobe the first input terminal, the arithmetic logic unit is used for modular multiplication followed by modular addition operation.

20. The acceleration unit according to claim 16, wherein, by setting the fifth strobe signal to strobe the first input terminal, setting the sixth strobe signal to strobe the second input terminal, and setting the seventh strobe signal to strobe the second input terminal, the arithmetic logic unit is used for modular addition followed by modular multiplication operation.

21. A computing device, comprising: a memory for storing homomorphic encryption instructions to be executed; the acceleration unit according to any one of claims 1-20; a processing unit for loading the homomorphic encryption instructions to be executed and distributing the homomorphic encryption instructions to be executed to the acceleration unit for execution.

22. A system on chip comprising the acceleration unit according to any one of claims 1-20.

23. A data center comprising the computing device according to claim 21.

24. A homomorphic encryption method applied to the computing device according to claim 21, comprising: receiving homomorphic encryption instructions to be executed; decomposing the homomorphic encryption instructions to be executed into operations; distributing the number theoretic transforms included in the operations to at least one of one or more number theoretic transform units for execution, and distributing the arithmetic operations included in the operations to at least one of one or more arithmetic logic units for execution.

Citation Information

Patent Citations

  • A homomorphic processing unit (HPU) for accelerating secure computations under homomorphic encryption

    CN110892393A

  • Zero-knowledge proof hardware accelerator and method thereof

    CN111373694A