System on chip and computer-implemented method

By adopting a buffer storage system and deterministic combination method in hardware accelerator, the inefficiency of memory utilization and coefficient storage during NTT calculation in PQC algorithm is solved, and efficient memory utilization and flexible algorithm support are achieved.

CN120067039APending Publication Date: 2025-05-30INFINEON TECHNOLOGIES AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411737416.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing hardware accelerators have problems with inefficiency in implementing number-theoretical transformation (NTT) in post-quantum cryptography (PQC) algorithms, especially lack of flexibility in memory utilization and coefficient storage, making it difficult to adapt to future algorithm changes.

Method used

Using a buffer storage system and a deterministic combination, effective memory utilization and flexible coefficient storage scheme are realized by processing the calculated coefficients performed at each stage of NTT. This arrangement allows the output coefficient to be written to the memory so that it is stored on the same address line as input for the next processing stage.

Benefits of technology

Improves memory utilization efficiency, reduces memory requirements, while maintaining high performance levels, and provides flexibility to support various coefficient sizes and multiple PQC algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067039A_ABST
    Figure CN120067039A_ABST
Patent Text Reader

Abstract

A system on chip and a computer-implemented method are disclosed. The described techniques improve performance and efficiency of hardware (HW) accelerators that may be used as part of post quantum cryptography (PQC) applications. Such hardware accelerators include those configured to perform so-called "butterfly operations" that process coefficients of a polynomial through which number-theoretical transformation (NTT) operations are performed. These techniques include an efficient memory storage scheme that uses a reordered buffer to ensure that a single address location can be computed when reading inputs to the next stage of the hardware accelerator. The described architecture facilitates a flexible and scalable solution to meet the high performance requirements of PQC processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the use of hardware accelerator processing architectures, and more particularly, to the use of hardware accelerator processing architectures that support memory storage solutions to facilitate efficient and high-performance solutions for unordered number theoretic transforms (NTTs) for post-quantum cryptography (PQC) applications. Background Art

[0002] In anticipation of the availability of significant processing power through the advent of quantum computing, post-quantum cryptography (PQC) algorithms have been developed. Such PQC algorithms include, for example, CRYSTALS-KYBER for key encapsulation mechanisms (KEMs), and CRYSTALS-DILITHIUM, FALCON, and SPHINCS+ for digital signature algorithms (DSAs). The security of CRYSTALS-KYBER, CRYSTALS-DILITHIUM, and FALCON (as well as other conventional PQC algorithms, such as NTRU and SABER) is based on mathematical problems related to (modular) lattices, and provides a good compromise between performance, key size, and security.

[0003] Hardware accelerators can be used as part of such PQC algorithms, which typically implement so-called number theoretic transforms (NTTs) and inverse NTTs. An NTT is a mathematical generalization of the discrete Fourier transform over a ring. Thus, an NTT is one of the most computationally intensive operations in PQC algorithm execution and requires a large amount of memory as well as complex memory access patterns. However, conventional hardware accelerators for PQC algorithms suffer from a lack of scalability and a lack of efficient coefficient storage and utilization. Summary of the Invention

[0004] Since the popularity of post-quantum cryptography (PQC) is expected to continue to increase in view of the progress of quantum computing, the acceleration of PQC is gaining momentum. However, PQC is not as stable as classical asymmetric cryptography. In particular, with respect to PQC, different ongoing standardizations are still evolving, and thus current implementations of PQC hardware accelerators require greater flexibility to accommodate expected future changes, recognizing that future products may need to support many different algorithms.

[0005] Accordingly, a hardware (HW) accelerator needs to efficiently implement multiple algorithms to meet the demand for flexibility in this currently evolving branch of cryptography. And, as described above, the use of NTT and inverse NTT is the most computationally intensive operation in lattice-based cryptography and is typically implemented as part of a PQC algorithm. To address such computationally intensive processing tasks, the input and output data for NTT calculations are divided into several stages, and within each stage, processing operations are performed on smaller groups of data inputs called coefficients. Additionally, the output coefficients from some stages are used to provide inputs to the next subsequent stage, where processing occurs across multiple stages in this way, further complicating the ability to make such calculations more efficient. Embodiments described herein relate to a processing architecture that supports a hardware accelerator configured to perform NTT calculations and addresses these issues while facilitating efficient coefficient storage and memory utilization. To that end, and as further discussed below, the embodiments described herein utilize the use of a buffer storage system and employ knowledge of the deterministic combination of coefficients for the calculations performed at each stage of the NTT. The use of buffers enables the output coefficients to be written to memory such that each group of coefficients used as inputs for the processing operations of the next processing stage is stored on the same address line.

[0006] This arrangement provides an efficient read scheme from the memory and achieves a reduction in memory compared to conventional systems while maintaining the same performance level. Additionally, the efficient memory storage scheme enables increased flexibility to support coefficient sizes of various widths to provide flexibility in supporting multiple types of PQC algorithms. Further, the processing architecture discussed herein can be advantageously scaled by adding additional hardware accelerators and accompanying logic to utilize the efficient use of the memory structure to improve performance by facilitating concurrent processing operations for each hardware accelerator. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The drawings are incorporated herein and form a part of the specification and illustrate aspects of the present disclosure and, together with the specification, further serve to explain the principles of these aspects and enable one of ordinary skill in the relevant art to implement and use these aspects.

[0008] FIG. 1 shows a block diagram of a conventional butterfly unit for performing butterfly operations according to a number theoretic transform (NTT);

[0009] Figures 2A to 2H shows a conventional ordered sequence of the first two stages of a butterfly operation for calculating an NTT;

[0010] Figure 3A and Figure 3B shows an example architecture for performing an unordered butterfly operation for calculating an NTT according to one or more embodiments of the present disclosure;

[0011] Figures 4A to 4G shows an unordered sequence for a first stage of a butterfly operation for computing an NTT according to one or more embodiments of the present disclosure;

[0012] Figures 5A to 5F shows an unordered sequence for a second stage of a butterfly operation for computing an NTT according to one or more embodiments of the present disclosure; and

[0013] Figure 6 shows an example processing flow according to an embodiment of the present disclosure.

[0014] Example aspects of the present disclosure will be described with reference to the accompanying drawings. The drawings in which elements first appear are generally denoted by the leftmost digit in the corresponding reference numerals. Detailed Description

[0015] In the following description, numerous specific details are set forth in order to provide a thorough understanding of aspects of the present disclosure. However, it will be apparent to those skilled in the art that these aspects, including structures, systems, and methods, may be practiced without these specific details. The description and representation herein are the common means used by those experienced or skilled in the art to most effectively convey the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the present disclosure.

[0016] I. Technical Overview: Number Theoretic Transform (NTT)

[0017] Similarly, the NTT is used in many post-quantum asymmetric cryptographic algorithms. In the context of these schemes, the NTT is a transformation of polynomials of degree < n with coefficients in (the two parameters may be different depending on the scheme under consideration). Thus, the input and output of the NTT each include n integers < q and are typically represented by a vector of length n, where each entry is at least w bits, where w ≥ ceil(log 2 (q)). The NTT can be decomposed into so-called "butterfly operation" stages. Each of these operations takes two inputs (two integers < q) and maps them to two outputs (two integers < q) using one multiplication, one addition, and one subtraction. One or more butterfly units, i.e., dedicated HW blocks that implement butterfly operations, are typically present in an NTT hardware accelerator.

[0018] With this in mind, a main challenge in implementing a PQC algorithm is to provide the input to the butterfly units and accept their output such that there is as little stalling as possible, i.e., to minimize the number of cycles in which the butterfly units do not perform arithmetic operations due to missing input / invalid input or backpressure from the output. Such stalling significantly affects the performance of the accelerator because butterfly units are typically designed for a throughput of one operation per cycle, and thus even a single cycle of stalling per butterfly operation results in a two-fold reduction in performance.

[0019] As discussed further below in detail, the embodiments described herein can be implemented according to any suitable hardware accelerator architecture, which can be implemented for any suitable algorithm that utilizes butterfly operations. Thus, although the embodiments described herein may be particularly advantageous for use with a hardware accelerator implemented as part of a PQC algorithm, this is by way of example and not limitation, and the embodiments described further herein can be implemented according to any suitable type of hardware accelerator architecture and / or algorithm implementation.

[0020] The embodiments described further below relate to a scalable processing architecture that supports one or more hardware accelerators, each hardware accelerator being implemented to perform an NTT by "anticipating" the execution of the next stage of the NTT calculation using paired coefficients. To this end, it should be understood that the NTT is a generalization of the FFT, and thus the embodiments discussed herein are also applicable to the use of the FFT and equally applicable to FFT hardware implementations.

[0021] The butterfly operations themselves are performed as part of a number of sequential multi-stage calculations to perform the NTT. The NTT operations are performed according to various parameters including the number of coefficients, the size of the coefficients (in bits), and the number of stages, which mainly depends on the number of coefficients. This will be described in more detail with reference to FIG. 1, which shows a conventional butterfly unit 102 including three stages and 8 coefficients. Specifically, FIG. 1 shows a radix-8 NTT divided into three stages each having 4 butterfly operations. For each stage, a pair of coefficients is used and processed through the butterfly unit 102, which can be implemented as an arithmetic unit, for example.

[0022] Thus, the NTT implementation shown in FIG. 1 represents an 8-dimensional NTT, and Figures 2A to 2G a memory architecture for performing the sequence of butterfly operations for such an 8-dimensional NTT implementation is shown in more detail in. As Figure 2AAs shown, the input to the butterfly unit 102 includes coefficients s(0) through s(7), each coefficient having a width (e.g., bit length) w. Thus, for stage 0, the output of the butterfly unit 102 for the NTT calculation is a'(0) through a'(7), and a'(0) through a'(7) also represent the input to stage 1. Additionally, for stage 1, the output of the butterfly unit 102 for the NTT calculation is a''(0) through a''(7), and a''(0) through a''(7) also represent the input to stage 2. The output of stage 2 is shown as through through representing the set of coefficients calculated for the NTT. Depending on the number of stages, different (but deterministic) combinations of coefficients are used to calculate the next stage (or final result) of the NTT. Thus, for the NTT calculation, each stage of the processing operation uses (i.e., operates on) a different pair of such coefficients.

[0023] Thus, as Figure 2A shown, for a conventional HW accelerator design, memory block A stores the input data, which is also the set of coefficients of the polynomial on which the NTT operation is to be performed. As discussed above, the magnitude of the coefficients is determined by q, which is a known parameter of the NTT operation. Then, the width of each coefficient, represented in bits per coefficient, is defined as w. The hardware supports integers up to a certain q_max (the maximum possible q), which is used to define w. This relationship can be expressed as w ≥ ceil(log 2 (q_max)). As an example, consider the coefficients s(0) through s(7) shown in FIG. 1, which again form the set of input data. The data is typically received in order, i.e., s(0), s(1), …, s(7) through the bus interface and sequentially complete the write to memory A. Then, the coefficients are stored in different RAM blocks of each memory in a predetermined manner such that any pair required as input to the butterfly operation at a given stage is distributed across two RAM blocks (e.g., RAM_A0, RAM_A1, RAM_B0, RAM_B1, etc.), enabling the concurrent reading of two elements.

[0024] Each RAM block in a separate RAM block can represent a separate physical single-port RAM that is independently addressed and accessed by the butterfly unit 102. In other words, due to the hardware configuration and the width of each coefficient stored in memory A, a single entry from two RAM blocks can be accessed at once as part of a single read operation for each RAM port, and no more than one entry can be accessed from the same RAM block. Due to this single-port design, in order to access multiple entries from the same RAM block, multiple read operations will need to be performed sequentially, which increases the latency.

[0025] II. Conventional Butterfly Operation

[0026] Figures 2A to 2H shows a conventional ordered sequence for the first two stages of the butterfly operations for computing the NTT. Specifically, Figures 2A to 2H shows an example of a conventional way of reading and writing data during the respective stages of the NTT processing operation. As Figure 2A shown, the input data of the example NTT from FIG. 1 is loaded into four different single-port RAM blocks 0, 1, 2, and 3 of memory A. The rows in each RAM block (i.e., RAM A0, A1, A2, and A3) define the number of memory lines. Referring to Figure 2A , due to the way the butterfly operations are performed, the first computation (i.e., for the first stage 0) requires the coefficient pairs s(0) and s(4). These two coefficients are written to two different locations across two separate RAM blocks since s(0) is received before s(4), and this would prevent the concurrent reading of s(0) and s(4) if the data was loaded into each address line of the same RAM block. Thus, the coefficients are stored in four different RAM blocks such that each input coefficient pair [s(0), s(4)]; [s(1), s(5)]; [s(2), s(6)]; and [s(3), s(7)] can be concurrently read and accessed via two different addresses from two different RAM blocks.

[0027] For each stage, the two outputs of the butterfly unit 102 are then stored in a similar manner to two of the other RAM blocks (i.e., RAM_B0, RAM_B1, RAM_B2, and RAM_B3) identified as memory B to enable the concurrent access of the coefficients required for the butterfly operations of the next stage. Continuing with this example, the operation results of s(0) and s(4) from the first stage are stored in RAM_B0 as a'(0) and in RAM_B2 as a'(4). This requires the computation of two different addresses in two different RAMs for the store operation. This process is repeated for each stage of the NTT operation, where each of the remaining computations in the remaining computations of the first stage is Figures 2A to 2D shown. For example, the next coefficients s(1), s(5) are read from memory A, where the results of the butterfly operations of stage 1 (i.e., a'(1), a'(5)) are stored in memory B. As Figure 2D shown, this continues until all the computations of the first stage have been completed.

[0028] Figures 2E to 2G shows the butterfly operations performed for the second stage. Thus, Figure 2E shows the coefficients a'(0), a'(2) read from memory B and used as inputs to generate the outputs a''(0), a''(2), as Figure 2FAs shown, the outputs a”(0) and a”(2) are written to memory A. As Figure 2G shown, this operation is repeated for the next stage until all the calculations for stage 2 have been performed, in which case, as Figure 2H shown, the result of the calculation (i.e., the processed data output) is stored in memory A. For the sake of brevity, the other stages of the butterfly operations are not shown in the figures. However, the process of reading the input coefficients from memory, performing the butterfly operations, and writing the outputs to memory is repeated until all the calculations for all the stages have been completed. That is, at the end of each stage, the output of each stage is accessed from one memory and stored in the other memory (either at a new address location or overwriting the previous value). After all the stages are completed, the last coefficients stored in the memory represent the final output, i.e., the coefficients associated with the actual calculation of the NTT.

[0029] Thus, it is noted that the same “pattern” of reading the input coefficients and writing the output coefficients is used for each stage, and the sequence of operations proceeds “in order” using the lowest numbered coefficients in each input pair for each stage of the device. For example, the input coefficients including s(0), a’(0), and a”(0), etc., are always used to initiate the calculations for each stage, followed by the input coefficients including s(1), a’(1), and a”(1), and so on. However, the conventional ordered sequence of performing the butterfly operations and the pattern of accessing and storing the calculations during each stage are disadvantageous.

[0030] For example, it is noted that for each processing operation, the output data (i.e., the processed data output) needs to be immediately stored in one of memories A or B. As a result, for the next stage, the butterfly unit 102 cannot access the next input pair without performing multiple address calculations. For example, for stage 0, if the inputs are s(0) and s(4), then the outputs a’(0) and a’(4) need to be immediately stored in memory B. However, for a conventional butterfly unit 102, the memory needs to be partitioned into 4 separately addressable single-port RAMs. As another example, to access s(0) and s(4), two RAM macros (RAM_A0 and RAM_A1) need to be read from. Thus, a conventional hardware accelerator requires at least 4 RAM macros to perform the concurrent reading of two coefficients required for each stage of one NTT operation. In addition, each load and store operation for each butterfly operation also requires calculating at least 2 addresses and accessing 4 different RAMs: 2 for loading and 2 for storing.

[0031] III. Example Architecture for Performing Unordered Butterfly Operations

[0032] Figure 3A and Figure 3BIllustrates an example architecture for performing unordered butterfly operations to compute the NTT according to one or more embodiments of the present disclosure. Architecture 300 can be implemented as any suitable type of circuit component, software component, or a combination of these components. For example, architecture 300 can include an integrated circuit, a system-on-chip, one or more interconnected SoCs or integrated circuits, etc. Architecture 300 can be implemented as part of any suitable system that utilizes butterfly operations as discussed herein. For example, architecture 300 can form part of any suitable type of system that uses post-quantum cryptography (PQC) for any suitable number and / or type of applications. This can include the use of PQC key generation, authentication, encryption, and / or decryption, etc. Additionally, although these techniques are discussed herein with respect to the use of butterfly operations and PQC algorithms, this is by way of example and not limitation, and the embodiments described herein can be implemented according to any suitable application that utilizes multi-stage computations in combination with predefined processing operations and memory reads and writes.

[0033] Architecture 300 includes a bus interface 308 that is configured to couple architecture 300 to any suitable number and / or type of data buses to support data communication according to any suitable number or type of communication protocols. The processing block 310 can receive instructions and / or data from other interconnected components via the bus interface 308 and can send data to other interconnected components via the bus interface 308. For example, the processing block 310 can receive a request for a PQC key or other PQC algorithms computed using butterfly operations as discussed herein, where the contents of the memory as discussed herein are transmitted in response to such a request.

[0034] Thus, the bus interface 308 can include any suitable implementation of components for this purpose, such as wires, buses, and / or corresponding terminals, ports, pins, etc. The bus interface 308 can be implemented as any suitable hardware component that enables communication between architecture 300 and other interconnected components within the applicable system. Thus, the bus interface 308 can include hardware components, software components, or a combination of these components that are generally associated with components configured to perform data communication. For example, the bus interface 308 can additionally or alternatively include any suitable number of ports, drivers, transmit and / or receive buffers, switches, etc.

[0035] Processing block 310 may include processing circuitry and / or any suitable number and / or type of dedicated hardware components, such as a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on a chip (SoC), dedicated logic, and / or other circuitry, etc. Processing block 310 may be implemented as one or more processors and / or cores, which may execute computer-readable instructions stored in program memory 312 to perform any of the various functions discussed further herein. Program memory 312 may include any suitable type of non-transitory computer-readable medium, such as volatile memory, non-volatile memory, or a combination of these memories. To the extent that architecture 300 implements a software-based solution to perform the various functions discussed herein, this may be achieved, for example, by processor block 310 (e.g., the processing circuitry associated therewith) executing instructions stored in program memory 312.

[0036] Processing block 310 is configured to monitor the operation of the various components of architecture 300 and generate control signals to facilitate the flow / transfer of data and the reading and writing of data to and from the various memories within architecture 300 as discussed further herein. Although processing block 310 is shown as a single component in Figure 3A it should be understood that processing block 310 may include any suitable number of sub-components, separate processors, logic, etc. to facilitate the operation of architecture 300 as discussed further herein, which operations may include the generation of various control signals. As discussed further herein, such control signals may include address / chip select control signals, input select control signals, result select control signals, butterfly result source select control signals, etc. For example, as discussed further herein, processing block 310 may include a memory processing and configuration block configured to generate address / chip select control signals with respect to the various memories.

[0037] Processing block 310 may additionally include a compute input / output (I / O) processing block configured to generate additional control signals that help control the data flow within architecture 300. As discussed further herein, such control signals may include, for example, input select, result select, and butterfly result select control signals, which are coupled to various multiplexers and used to control the data flow to the various components of architecture 300. Of course, the generation of any of the control signals discussed herein may alternatively be generated by processing block 310 implemented as a single processing component.

[0038] Architecture 300 includes any suitable number of butterfly units, where butterfly unit 302.1 is shown in Figure 3A which will be referred to below with reference to Figure 3BThe use of additional butterfly units is discussed in further detail. The butterfly unit 302.1 can be implemented as any suitable type of hardware component configured to perform a predetermined type of processing operation according to any suitable number of stages. For example, the butterfly unit 302.1 can be implemented as a hardware accelerator that can implement a predetermined combinational logic or other suitable arrangement of hardware components to compute processed data output from an input data set (e.g., a pair).

[0039] In the case of being implemented as a hardware accelerator, the butterfly unit 302.1 can be configured as a number-theoretic transform (NTT) hardware accelerator that is configured to perform butterfly operations as discussed above to compute the NTT. To this end, the butterfly unit 302.1 can perform a predetermined number of processing operations, such as butterfly operations, for each stage within a multi-stage computation set (e.g., the four stages as described above for the conventional butterfly unit 102). The butterfly unit 302.1 can be implemented in a similar or identical manner to the conventional butterfly unit 102 described herein and can operate in a similar manner. However, as discussed in further detail below, compared to the conventional butterfly unit 102, the butterfly unit 302.1 performs processing operations in a different order (i.e., disordered) within each stage.

[0040] Furthermore, although the butterfly unit 302.1 is described herein in terms of the same number of stages, inputs, and number of butterfly operations per stage as the conventional butterfly unit 102, these are parameters provided for illustrative purposes and by way of example rather than limitation. For example, the butterfly unit 302.1 can perform any suitable number of processing operations (e.g., butterfly operations) per stage, which can be performed on any suitable number of data inputs to generate corresponding processed data outputs for any suitable number of sequential stages. In this way, the butterfly unit 302.1 performs processing operations on a predetermined group of data inputs within the entire data input set for each sequential stage.

[0041] The data input set can include, for example, each of the coefficients used as inputs for a particular stage as discussed above for the butterfly unit 102. For example, for the initial stage 0 of the butterfly unit 302.1, as described above for the conventional butterfly unit 102, each predetermined group of data inputs includes a pair of coefficients of a polynomial on which the NTT is computed via a butterfly operation. The processed data output for each stage of the butterfly operation can likewise be referred to as coefficients herein, but it should be understood that these coefficients can be intermediate values subject to further processing operations until the final stage of the processing operations, where the final coefficients are thus output at the last stage.

[0042] In addition, each predetermined data input group received at each stage (e.g., each pair of coefficients) can be a function of the number of inputs and the number of butterfly operations performed at each stage. For example, assuming that the butterfly unit 302.1 is implemented in a manner similar to the conventional butterfly unit 102, the entire data input set for each stage will represent a total of 8 coefficients. Further assuming that 4 butterfly operations are performed within each stage, each predetermined data input group will include two coefficients. For example, for stage 0, each pair of coefficients from s(0) to s(7), for stage 1, each pair of coefficients from a'(0) to a'(7), and for stage 2, each pair of coefficients from a''(0) to a''(7).

[0043] Thus, as described above with respect to the conventional butterfly unit 102, the butterfly unit 302.1 can also output a corresponding predetermined processed data output group (e.g., a pair of output coefficients for each pair of input coefficients) that forms a processed data output set (e.g., the entire output coefficient set for each stage) for each stage in the sequential stages. Thus, each predetermined processed data output group for each stage can include, for example, each pair of coefficients from a'(0) to a'(7) for stage 0, each pair of coefficients from a''(0) to a''(7) for stage 1, and for the last stage 2 to each pair of coefficients. In other words, the predetermined data input group (e.g., coefficient pair) for one or more stages (e.g., for stages 1 and 2) is formed from the processed data output set output by the previous stage (e.g., stages 0 and 1).

[0044] Continuing to refer to Figure 3A , the architecture 300 also includes any suitable number of memory blocks, where two memory blocks are shown by way of example and not limitation in Figure 3A . The memory blocks 306A, 306B may alternatively be referred to simply as memories herein. Each of the memory blocks 306A, 306B can include any suitable number and / or type of physical and addressable memory. This can include, for example, non-volatile memory or volatile memory, such as any suitable type of random access memory (RAM) that can include static RAM, dynamic RAM, etc. Each of the memories A0, A1, B0, and B1 shown in Figure 3A can be implemented as, for example, a single-port RAM (e.g., single-port SRAM), where, as discussed above with respect to Figures 2A to 2H , one of the memory blocks 306A, 306B initially stores the input data for the first stage 0 of the butterfly unit 302.1, and the computed output for each stage is written to the other memory block 306A, 306B.

[0045] Each of SRAMs A0, A1, B0, and B1 includes any suitable number of unique addresses, which may alternatively be referred to herein as rows. The coefficients processed by the butterfly unit 302.1 at various stages of the processing operation may have any suitable width (e.g., bit length) such that SRAMs A0, A1, B0, and B1 can store multiple coefficients based on the total number of rows and the coefficient width. Other details regarding the size of the coefficients and how the architecture 300 is differently implemented for varying coefficient widths are discussed in further detail below. Additionally, and as discussed in further detail below, since each individual butterfly input pair requires one address calculation, this enables the use of a smaller memory structure. Thus, compared to the conventional solution discussed above with respect to Figures 2A to 2H each memory 306A, 306B only requires half the SRAMs.

[0046] IV. Example Sequence for Performing Unordered Butterfly Operations

[0047] Now turning to Figures 4A to 4G , note that these figures illustrate an unordered sequence of the first stage of the butterfly operations for computing the NTT according to one or more embodiments of the present disclosure. The initial steps in this sequence are as Figure 4A shown. Note that the input data may be received in sequence (e.g., via the bus interface 308 or other data source), i.e., each of the coefficients s(0) to s(7) is received in sequence by the processing block 310. However, for the first half (i.e., s(0) to s(3)), the coefficients s(0) to s(7) are loaded into SRAMs A0 and A1 in sequence, and then again for the second half (i.e., s(4) to s(7)) into SRAMs A0 and A1. Thus, and as Figure 4A shown, the coefficients s(0) to s(3) are stored in the rows (i.e., addresses) of SRAM A0, while the coefficients s(4) to s(7) are stored in the rows (i.e., addresses) of SRAM A1. As Figure 3A shown, the processing block 310 can control the manner in which the coefficients s(0) to s(7) are stored in the memory 306A by the timing generation of the address / chip select signals. The timing generation of the address / chip select control signals in this manner can be based, for example, on the knowledge of the butterfly unit 302.1 and the initial inputs of stage 0. That is, the initial coefficients s(0) to s(7) are stored in the memory 306A such that each pair of coefficients to be provided as input data to the butterfly unit 302.1 at stage 0 and used to compute the corresponding output coefficients a'(0) to a'(7) are stored in the same (but non-sequential, although deterministic) corresponding address lines. For example, and as Figure 4AAs shown, each pair of coefficients [s(0), s(4)], [s(1), s(5)], [s(2), s(6)], and [s(3), s(7)] is stored on the same address line of the memory 306A. This arrangement provides an efficient read scheme for performing stage 0 butterfly operations from the memory 306A.

[0048] For example, the input coefficients for the first calculation in stage 0 are [s(0), s(4)]. These two coefficients are written to two different SRAMs A0, A1 within the same memory block 306A and are thus single-address accessible. To perform a read operation, the processing block 310 can access the contents of [s(0), s(4)] in the memory block 306A using a single address calculation. Thus, Figure 4B shows the result of the processing circuit 310 controlling the transfer of the coefficient pair [s(0), s(4)] from the memory block 306A to the butterfly unit 302.1 by using an address / chip select control signal and an input select control signal, which control the mux 350 such that the contents of the currently selected address of the memory block 306A, rather than the contents of the memory block 306B, are provided as an input to the butterfly unit 302.1.

[0049] The result of the first butterfly operation performed on the input coefficient pair [s(0), s(4)] is Figure 4B represented as the output coefficient pair [a'(0), a'(4)]. As Figure 4C shown, the output coefficient pair [a'(0), a'(4)] is stored in a reorder buffer, which may alternatively be referred to herein simply as a buffer, and as Figure 3A shown, can be identified by one of the buffers 304.0, 304.1. The particular buffer 304.0, 304.1 used is controlled by a control signal output by the processing block 310. As an example, buffer 304.0 can be used for each stage 0 operation in stage 0 operations, where the calculation result 0 data is stored in the memory 306B by the processing block 310, and the processing block 310 controls the operations of the mux304.0, mux 352.0, and mux 354 through appropriate control signals. Then, buffer 304.1 can be used to store the results of stage 1 operations, etc., where the results of each operation of each stage are stored in one of the buffers 304.0, 304.1 in an alternating manner. According to this operation mode, the buffer 304.0, 304.1 in which the coefficients are stored at a particular time depends on the current stage of the processing operation currently being performed by the butterfly unit 302.1.

[0050] However, as another example, it may be convenient to alternatively select buffer 304.0 or 304.1 within each operation phase for storing the calculation results of each coefficient pair of a given operation phase. In other words, according to this operation mode, the results can be stored (within each phase of the processing operation) in each of buffers 304.0 and 304.1. As an illustrative example, this can include storing the results (i.e., a pair of coefficients) from one or more processing operations within phase 0 in buffer 304.0, storing the next results from one or more processing operations within phase 0 in buffer 304.1, and repeating this process in an alternating manner within each phase of the processing operation.

[0051] Thus, and as Figure 4C shown, buffers 304.0 and 304.1 may initially store the output pairs of coefficients a'(0) and a'(4) as the case may be, until the results of the next processing operation (i.e., the input pairs s(2), s(6)) are available. Then, the buffers may retain the output coefficient pairs a'(0) and a'(4), while providing the second input coefficient pair [s(2), s(6)] at the input of the butterfly unit 302.1. Similarly, the input coefficient pair [s(2), s(6)] is loaded into the same memory line of memory 306A and is thus accessible at the same address to be transferred in this manner through mux 350. The result of performing the second butterfly operation on the input coefficient pair [s(2), s(6)] is represented as the output coefficient pair [a'(2), a'(6)] in Figure 4C shown.

[0052] Then, as Figure 4D shown, the coefficient pair transferred to memory 306B includes the output coefficient a'(2) paired with the previously calculated coefficient a'(0), such that the output coefficients [a'(0), a'(2)] are stored in the same address line for the next phase of the butterfly operation (i.e., phase 1). In other words, now returning Figure 3A , each of the output coefficients [a'(0), a'(2)] is routed through muxes 352.0, 352.1, and optionally through mux 354, and is routed to memory 306B through appropriate control signals provided by the processing block 310. To this end, the output of the butterfly unit 302.1 (i.e., a'(2)) is written directly to memory 306B together with the output of the previously calculated butterfly unit 302.1 (i.e., a'(0)), which is temporarily stored in the reorder buffer to facilitate this result.

[0053] Continuing with this example, as Figure 4DAs shown, the coefficients a'(0) and a'(2) are stored at the same address line in the memory 306B, while the output coefficients a'(6) and a'(4) are stored together in the reorder buffer. Then, as Figure 4E shown, a second write is performed to transfer the contents of the reorder buffer to the memory block 306B such that the output coefficients [a'(4), a'(6)] are stored to the same (but non-sequential) address line. In this way, note that the butterfly unit 302.1 computes each predetermined data input group (i.e., coefficient pair) as part of a processing order that is based on each predetermined data input group that will be used in the next stage. For example, contrary to the conventional case, the butterfly operations performed by the butterfly unit 302.1 for each stage are performed as part of a non-sequential processing order. For example, contrary to processing the input coefficients [s(1), s(5)] after the input coefficients [s(0), s(4)], the butterfly unit 302.1 performs an unordered (i.e., non-sequential) processing sequence. This unordered processing sequence is controlled by the processing block 310 and uses the predetermined knowledge of the inputs required for the next stage of the butterfly operations. In this way, each stage can access the input coefficient pairs required for each processing operation from the same address line in the memory 306B, thus facilitating an efficient paired storage as part of the initial data loading.

[0054] This unordered sequence of processing operations continues in the same way until all inputs have been processed and all output coefficients have been stored in the memory 306B. For example, and as Figure 4F and Figure 4G shown, the sequence of operations for stage 0 would be: [s(0), s(4)], [s(2), s(6)], [s(1), s(5)], [s(3), s(7)], while a conventional butterfly unit typically performs these same processing operations in an ordered sequence [s(0), s(4)], [s(1), s(5)], [s(2), s(6)], [s(3), s(7)].

[0055] In this way, the reorder buffer is configured to store the outputs of the butterfly unit 302.1 before they are stored in the memory 306B, thus enabling the reordering of the coefficient pairs that can be accessed via the same line for the processing operations of the next stage. Thus, the processing block 310 controls the transfer of the processed data outputs provided by the butterfly unit 302.1 for each stage by generating control signals at the appropriate times. This includes transferring the processed data outputs stored in the buffer together with subsequent processed data outputs to the memory 306B at the appropriate time such that each predetermined data input group (e.g., each coefficient pair) for the next stage is read from the same address line in the memory 306B and loaded into the butterfly unit 302.1.

[0056] This unordered sequence of processing operations for the butterfly unit 302.1 can continue for each stage until the final output coefficients are computed. Thus, now referring to Figures 5A to 5F , which shows an unordered sequence for the second stage of the butterfly operations for computing the NTT according to one or more embodiments of the present disclosure. Figures 5A to 5F The processing operations shown in are related to stage 1 of the butterfly operations performed by the butterfly unit 302.1, and these processing operations are performed after stage 0 as shown in Figures 4A to 4F .

[0057] Thus, as discussed above, the memory 306B as shown in Figure 5A includes the outputs from the stage 0 processing operations. The contents of the memory 306B are now read and used as the inputs for stage 1 of the butterfly operations, where the computed results, i.e., the processed output data (e.g., coefficient pairs), are now stored in the memory 306A. Since stage 1 represents the penultimate stage, the sequence of operations for stage 1 is now ordered, where the sequence of processing operations includes the coefficient pairs [a'(0), a'(2)], [a'(1), a'(3)], [a'(4), a'(6)], [a'(5), a'(7)].

[0058] Thus, as shown in Figure 5B , the buffer holds the output coefficients a''(0), a''(2) from the first processing operation, and the next input coefficients a'(1), a'(3) are provided as the inputs for the butterfly unit 302.1 for the next processing operation, which produces the output coefficients a''(1), a''(3). As shown in Figure 5C , the coefficient pairs transmitted to the memory 306B include the output coefficient a''(1), and a''(1) is paired with the previously computed coefficient a''(0) such that the output coefficient [a''(0), a''(1)] is stored in the same address line for the next stage of the butterfly operations (i.e., stage 2).

[0059] Then, as shown in Figure 5D , a second write is performed to transfer the contents of the reorder buffer to the memory block 306A such that the output coefficients [a''(2), a''(3)] are stored in the same (but non-sequential) address line. Additionally, during this operation, Figure 5D also shows that the next output coefficients a''(4), a''(6) are stored in the buffer, and the next coefficient input pair a'(5), a'(7) are provided as the inputs for the butterfly unit 302.1 for the next processing operation, which produces the output coefficients a''(5) and a''(7). As shown in Figure 5EAs shown, the coefficient pairs transmitted to memory 306B include the output coefficient a”(5), and a”(5) is paired with the previously calculated coefficient a”(4) such that the output coefficients [a”(4), a”(5)] are stored in the same address line for the next stage of butterfly operations (i.e., stage 2). Finally, as Figure 5F shown, a second write is performed to transfer the contents of the reorder buffer to memory block 306A such that the output coefficients [a’(6), a’(7)] are stored to the same (but non-sequential) address lines. Thus, after the processing operations of stage 1 are completed, each line of memory 306A includes input coefficient pairs to be used as inputs for the final stage 2.

[0060] Thus, compared to the conventional use of butterfly units, for the same performance, the number of memory accesses is four times for each coefficient pair. However, processing block 310 only needs to calculate one address per pair instead of two addresses. Also, the processing order within each stage is based on the coefficient pairs for the next stage. In other words, the pairing of the coefficients required for the next stage is determined, and processing block 310 can be programmed or otherwise configured to utilize this information to control data transfer within architecture 300 to ensure that the processing operations are performed and the results are stored in a pattern that facilitates efficient storage and use of memories 306A, 306B as discussed above.

[0061] Similarly, in the case where each single butterfly input pair only requires one address, the calculations are performed more efficiently. This enables the use of a smaller memory structure because the test hardware around the RAM is reduced. Thus, compared to the conventional solution discussed above Figures 2A to 2H each of memories 306A, 306B only requires half the SRAM. Thus, the embodiments discussed above can achieve the same performance as the conventional solution while only requiring the use of additional buffer memories. This advantageously provides an overall reduction in area by eliminating the large amount of space required for SRAM support hardware (SSH).

[0062] V. Support for Performing NTT Calculations with Additional Coefficients

[0063] The examples discussed above enable the efficient use of memories 306A, 306B to perform NTT calculations according to coefficient n with width w, where the coefficient n can represent the size of the coefficient according to the bit length of the coefficient and / or the address size, e.g., Figure 3A as shown for memories 306A, 306B. However, the efficient use of the memory structure described above provides additional flexibility that can be utilized to advantageously extend support for configurations in which NTT can be calculated using a larger number of coefficients with a smaller width (e.g., w / 2), although the performance will be reduced.

[0064] For example, the performance degradation may be the result of reading from and writing to the same SRAM. That is, in the above example, the number of coefficients n is 8, and each of the data inputs and outputs (i.e., the initial input coefficients, intermediate coefficients, and final coefficients) is stored in two different memory blocks 306A, 306B at different stages to improve performance. However, during each processing stage, one of the memory blocks 306A, 306B is used to store the input (i.e., the coefficient pair) that is read and provided to the butterfly unit 302.1, while the other memory block 306A, 306B stores the processed data output (i.e., the output coefficient pair). Then, the memory block used for reading and writing changes between each processing stage.

[0065] However, if the number of coefficients increases (e.g., n = 16), each memory block can concurrently store the data that is read and used as input, where the computation result is stored back into the same memory from which the input data was read. In other words, for an algorithm with a smaller number of coefficients n (e.g., 8), for each stage, one of the memory blocks 306A, 306B can be used for data reading, while the other memory block 306A, 306B can be used for writing the computed data of the previous coefficient pair, enabling high-performance computation.

[0066] Note that there may be variations in efficiency regarding the coefficient width. To provide an illustrative example, each pair of coefficients can be represented as 22 bits, and the algorithm can use a total of 8 coefficients. That is, each individual coefficient can have a size of 11 bits in length. This would enable four coefficient pairs to be stored in memory block 306A, such that the results of four computation steps can be stored within memory block 306B for each computation cycle. Thus, one butterfly operation can be performed per loop.

[0067] However, to provide another illustrative example, 8 coefficient pairs can be used, with each coefficient again having a length of 11 bits as described above. However, four coefficient pairs can be stored in memory block 306A, and the remaining four coefficient pairs can be stored in memory block 306B. In this way, each butterfly operation is spread over two cycles, which may double the computation time.

[0068] However, by reading the input of the butterfly unit 302.1 from one of the memory blocks 306A, 306B in one cycle and then writing the computation result into the same memory block 306A, 306B in the next cycle, the same memory architecture as discussed above can be achieved for algorithms with a larger number of coefficients (e.g., 16, 256, etc.).

[0069] In other words, for a larger number of coefficients, the alternating read and write operations on the same memory block within the same processing stage of the butterfly unit 302.1 can be performed by the processing block 310. Thus, even without adding more memory, the architecture 300 provides the flexibility to support NTTs with a larger number of coefficients with reduced performance, while maintaining higher performance for NTT calculations with a smaller number of coefficients.

[0070] Thus, the butterfly unit 302.1 can be optimized for an input with n coefficients, i.e., it supports an input with one butterfly operation per clock cycle for that number of coefficients (e.g., for a pair of coefficients in the case of n = 8). The butterfly unit 302.1 can also support 2n (e.g., 16) coefficients by reducing the performance to one butterfly operation every two clock cycles (e.g., for a pair of coefficients in the case of n = 16). Thus, compared to the calculation of the number of coefficients for which the accelerator is optimized, the performance of the latter is reduced by a factor of two. However, in either case, the butterfly unit 302.1 achieves a significant acceleration compared to a software implementation. This can be achieved through the memory configuration and the architecture 300 described above.

[0071] VI. Extending the Architecture for Additional Butterfly Units to Improve Performance

[0072] In an additional embodiment, the memory configuration discussed above can be modified to support additional scalability. For example, each of the memory blocks 306A, 306B includes two SRAMs A0, A1 and B0, B1, respectively. This represents dividing the total memory width 2w into two halves, where, as Figure 3A shown, each of the SRAMs A0, A1, B0, B1 is configured to store coefficients of width w.

[0073] However, the memory blocks 306A, 306B can be reformatted without reducing the total width 2w of the memory blocks by subdividing the total width of the memory blocks (i.e., 2w) into four parts instead of two parts as discussed above. Thus, as discussed above, each row of SRAMs A0, A1, B0, and B1 can alternatively store a total of four coefficients of length w / 2 instead of two coefficients of length w per row. Thus, a 2-fold performance improvement can be achieved for an algorithm that implements smaller-sized coefficients of width w / 2.

[0074] For illustrative examples, for Dilithium, the coefficient width required in the memory can be expressed as w_D ≥ 23 (w_D ≤ w), and for Kyber, the coefficient width required in the memory can be expressed as w_K = 12 (w_K < w / 2). Thus, both memories 306A and 306B can be configured to store two w_Ds (i.e., 48 bits) or four Kyber coefficients of width (w_K) in a single memory line (i.e., row).

[0075] To this end, reference is now made to Figure 3B , which shows additional butterfly units 302.2, additional buffers 304.2, 304.3, and additional muxes 352.2, 352.3. As the components shown in Figure 3B can be configured the same or similarly to the similar components previously described with respect to Figure 3A . For example, as shown and described above with reference to Figure 3A , the butterfly units 302.2, buffers 304.2, 304.3, and muxes 352.2, 352.3 can be configured as butterfly unit 302, buffers 304.0, 304.1, and muxes 352.1, 352.2, respectively, and operate in the same manner as them.

[0076] As the additional components shown in Figure 3B can be referred to herein as additional architecture 380 and can form a scaled extension of architecture 300 as shown in Figure 3A . Thus, when present, the additional architecture 380 can be coupled to other components as shown in Figure 3A and Figure 3B . Specifically, the inputs of each of the butterfly units 302.1, 302.2 are each coupled to mux 350, and the input selection control signal generated by the processing block 310 can facilitate the content of either of the memory blocks 306A, 306B being provided as an input to either of the butterfly units 302.1, 302.2. In addition, the outputs of each of the butterfly units 302.1, 302.2 are coupled to mux 354, and the butterfly result source selection control signal generated by the processing block 310 can facilitate the result of either of the butterfly units 302.1, 302.2 being written to either of the memory blocks 306A, 306B. Although only one additional architecture is shown and discussed herein, this is by way of example and not limitation. The architecture 300 can be scaled in a similar manner by extending the concept of the additional architecture 380 to include any suitable number of additional architectures, each additional architecture including components similar or identical to those shown in Figure 3B .

[0077] Continuing with the above Kyber PQC algorithm and asFigures 3A to 3B The example shown in uses two butterfly units, and such a configuration enables serving two butterfly units 302.1, 302.2 in a single read operation. Thus, for each processing stage, two pairs of computations can be computed in parallel (i.e., concurrently), with each coefficient having a width of w / 2. Thus, as Figure 3A and Figure 3B shown in , the inputs of the butterfly units 302.1, 302.2 can be read from one of the memories 306A, 306B, and subsequently the output computations for these two pairs of coefficients can be stored according to the paired computations described above for the single coefficient pair operation. Thus, these two pairs of coefficients can be computed in parallel and independently by each respective butterfly unit 302.1, 302.2. This enables using the same memory size of the memory blocks 306A, 306B while significantly improving the performance for coefficients of smaller width.

[0078] VII. Example Processing Flow

[0079] Figure 6 FIG. shows an example processing flow according to an embodiment of the present disclosure. The processing flow 600 may include a method performed by and / or otherwise associated with any suitable number and / or type of components (e.g., one or more processors (processing circuits), hardware components, executed instructions (e.g., software components), or a combination of these components). These components may be associated with one or more components of the architectures 300, 380 discussed herein. For example, the blocks shown in Figure 6 can be performed by any one of the butterfly units 302.1, 302.2, the processing block 310, etc. The processing flow 600 may include alternative or additional blocks not shown in Figure 6 for the sake of brevity, and may be executed in an order different from that shown. Additionally, some blocks may be optional.

[0080] The processing flow 600 begins by receiving (block 602) a predetermined set of data inputs from a memory. Again, as discussed above, the predetermined set of data inputs may include any suitable number (e.g., a pair) of coefficients to be processed according to a butterfly operation. For example, the data inputs may be received (block 602) from any suitable memory and provided as inputs to the hardware accelerator. As discussed above, the data inputs may be accessed via the same address lines of the memory.

[0081] The processing flow 600 further includes performing (block 604) a processing operation on a predetermined data input group. As described above, this can include, for example, performing a butterfly operation on the data input. The result of the processing operation can generate a corresponding predetermined processed data output group. For example, if the predetermined data input group includes coefficient pairs a(0), a(4), then the predetermined processed data output group will include coefficients a'(0) and a'(4).

[0082] The processing flow 600 further includes controlling (block 606) the transfer of the predetermined processed data output group to a buffer and / or memory. As discussed above, this can include, for example, the processing block 310 generating (block 606) appropriate control signals to facilitate the transfer of data output by the butterfly units 302.1 and / or 302.2 to one or more of the buffers or memories 306A, 306B.

[0083] The processing flow 600 further includes determining (block 608) whether all data input groups have been processed. Using the above-mentioned butterfly operation for an 8-coefficient NTT as an example, this can be determined once the last group among the 4 coefficient pairs has been processed, thereby terminating the current processing stage. If not (block 608, N), then the next predetermined data input group is received (block 602), and the processing flow is repeated until all data input groups have been processed.

[0084] Once the current processing stage has been terminated (block 608, yes), the processing flow proceeds to a further determination (block 610) as to whether the current processing stage is the last stage. If so (block 610, Y), then the processing is complete (block 612), and the stored processed data output represents the final coefficients for the NTT calculation.

[0085] Otherwise, the processing flow 600 repeats by receiving (block 602) the next predetermined data input group for the next processing stage. Again, as described above, the transfer of the processed data is controlled (block 606) within each processing stage such that the predetermined data input groups for each stage can be accessed via a single line of the memory, and thus only a single address needs to be calculated.

[0086] Example

[0087] The techniques of the present disclosure can also be described in the following examples.

[0088] Example 1. A system-on-chip (SoC) includes: a hardware accelerator configured to, for each stage in a set of sequential stages, (i) perform a processing operation on each data input in a predetermined group of data inputs in a data input set and (ii) output a corresponding predetermined group of processed data outputs in a processed data output set; a buffer configured to store one or more of the processed data outputs in the predetermined group of processed data outputs before the one or more processed data outputs are stored in a first memory or a second memory, wherein, for one or more stages in the set of sequential stages, the data input set is formed from a predetermined group of processed data output sets output by a previous stage; and a processing circuit configured to control one or more of the processed data outputs in the predetermined group of processed data outputs stored in the buffer to be transferred to the first memory or the second memory such that, for one or more stages in the set of sequential stages, each predetermined group of data inputs is read from the same address line in the first memory or the second memory.

[0089] Example 2. The SoC according to Example 1, wherein the hardware accelerator includes a number theoretic transform (NTT) hardware accelerator.

[0090] Example 3. The SoC according to any combination of Examples 1 to 2, wherein the processing operation includes a butterfly operation performed to compute a number theoretic transform (NTT).

[0091] Example 4. The SoC according to any combination of Examples 1 to 3, wherein the hardware accelerator is configured to, for a first stage in one or more stages, compute each predetermined group of data inputs as part of a processing order based on a predetermined group of data inputs for a subsequent second stage.

[0092] Example 5. The SoC according to any combination of Examples 1 to 4, wherein the processing order is non-sequential.

[0093] Example 6. The SoC according to any combination of Examples 1 to 5, wherein, for an initial stage in the set of sequential stages, each data input in the predetermined group of data inputs includes a pair of coefficients of a polynomial through which an NTT is to be computed by the butterfly operation.

[0094] Example 7. The SoC according to any combination of Examples 1 to 6, wherein, for each stage in the set of sequential stages, a predetermined group of data inputs is read from one of the first memory or the second memory while writing the predetermined group of processed data outputs to the other of the first memory or the second memory.

[0095] Example 8. An SoC according to any combination of Examples 1 to 7, wherein, for each stage in the set of sequential stages, a predetermined set of data inputs is read from the first memory or the second memory, and subsequently the processed data outputs of the predetermined set are written to the same memory of the first memory or the second memory.

[0096] Example 9. An SoC according to any combination of Examples 1 to 8, wherein the first memory and the second memory include one or more single-port static random access memories (SRAMs).

[0097] Example 10. An SoC according to any combination of Examples 1 to 9, further comprising: another hardware accelerator, and wherein, for each stage in the set of sequential stages, the hardware accelerator and the another hardware accelerator are each configured to concurrently perform processing operations on corresponding predetermined sets of data inputs in the data input set.

[0098] Example 11. An SoC according to any combination of Examples 1 to 10, wherein the first memory and the second memory have a plurality of memory lines, and the width of each memory line in the plurality of memory lines is equal to twice the width of each data input in the data input set and equal to twice the width of each data output in the processed data output set.

[0099] Example 12. An SoC according to any combination of Examples 1 to 11, wherein the set of sequential stages includes a first stage and a second stage, and upon completion of the first stage, the first memory or the second memory stores the processed data outputs of each corresponding predetermined set from the first stage, and wherein the processed data outputs of each corresponding predetermined set from the first stage are stored at the same corresponding address line in the first memory or the second memory.

[0100] Example 13. A computer-implemented method, comprising: for each stage in the set of sequential stages, performing a processing operation on each data input in a predetermined set of data inputs in the data input set and outputting the processed data outputs of the corresponding predetermined set in the processed data output set; storing one or more of the processed data outputs in the predetermined set of processed data outputs in a buffer before being stored in the first memory or the second memory, wherein, for one or more stages in the set of sequential stages, the predetermined set of data inputs is formed by the processed data outputs of the predetermined set output by the previous stage; and controlling one or more of the processed data outputs in the predetermined set of processed data outputs stored in the buffer to be transferred to the first memory or the second memory such that, for one or more stages, each corresponding predetermined set of data inputs is read from the same address line in the first memory or the second memory.

[0101] Example 14. The computer-implemented method according to Example 13, wherein the hardware accelerator includes a number theoretic transform (NTT) hardware accelerator.

[0102] Example 15. The computer-implemented method according to any combination of Examples 13 to 14, wherein the processing operation includes a butterfly operation that is performed to compute a number theoretic transform (NTT).

[0103] Example 16. The computer-implemented method according to any combination of Examples 13 to 15, further comprising: for a first stage in one or more stages, computing data inputs of each predetermined group as part of a processing order based on data inputs of a predetermined group for a subsequent second stage.

[0104] Example 17. The computer-implemented method according to any combination of Examples 13 to 16, wherein the processing order is non-sequential.

[0105] Example 18. The computer-implemented method according to any combination of Examples 13 to 17, wherein for an initial stage in a set of sequential stages, each data input in the predetermined group of data inputs includes a pair of coefficients of a polynomial through which an NTT is to be computed by means of a butterfly operation.

[0106] Example 19. The computer-implemented method according to any combination of Examples 13 to 18, wherein for each stage in a set of sequential stages, a predetermined group of data inputs is read from one of a first memory or a second memory while a predetermined group of processed data outputs is written to the other of the first memory or the second memory.

[0107] Example 20. The computer-implemented method according to any combination of Examples 13 to 19, wherein for each stage in a set of sequential stages, a predetermined group of data inputs is read from a first memory or a second memory and subsequently a predetermined group of processed data outputs is written to the same memory of the first memory or the second memory.

[0108] Example 21. The computer-implemented method according to any combination of Examples 13 to 20, wherein the first memory and the second memory include one or more single-port static random access memories (SRAMs).

[0109] Example 22. The computer-implemented method according to any combination of Examples 13 to 21, further comprising: for each stage in a set of sequential stages, concurrently performing processing operations on corresponding predetermined groups of data inputs in a data input set by a hardware accelerator and another hardware accelerator.

[0110] Conclusion

[0111] Although specific embodiments have been shown and described herein, it should be understood that any arrangement designed to achieve the same purpose may be substituted for the specific embodiments shown. The present disclosure is intended to cover any and all adaptations or variations of various embodiments. After reviewing the above description, combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art.

[0112] It should also be noted that the specific terms used in this specification and the claims may be interpreted in a very broad sense. For example, the term "circuit" as used herein should be interpreted in the sense of including not only hardware, but also software, firmware, or any combination thereof. The term "data" may be interpreted to include any form of represented data. In addition to any form of digital information, the term "information" may also include other forms of represented information. In an embodiment, the term "entity" or "unit" may include any device, equipment circuit, hardware, software, firmware, chip, or other semiconductor, as well as a logical unit or physical implementation of a protocol layer, etc. Furthermore, the term "coupled" or "connected" may be interpreted in a broad sense, covering not only direct coupling but also indirect coupling.

[0113] It should also be noted that the methods disclosed in the specification or claims may be implemented by a device having modules for performing each corresponding step of these methods.

[0114] Although specific embodiments have been shown and described herein, those of ordinary skill in the art will understand that various alternative implementations and / or equivalent implementations may be substituted for the specific embodiments shown and described without departing from the scope of the present disclosure. The present disclosure is intended to cover any adaptations or variations of the specific embodiments discussed herein.

Claims

1. A system on chip SoC, comprising: a hardware accelerator configured to, for each stage in the set of sequential stages, perform a processing operation on each data input in a predetermined set of data inputs in the set of data inputs and output a corresponding predetermined set of processed data outputs in the set of processed data outputs; a buffer configured to store one or more of the predetermined set of processed data outputs before the one or more processed data outputs are stored in the first memory or the second memory, wherein, for one or more stages in said set of sequential stages, said set of data inputs is formed by said predetermined set of processed data outputs output by a preceding stage; as well as processing circuitry configured to control one or more of the predetermined groups of processed data outputs stored in the buffer to be transferred to the first memory or the second memory such that for one or more stages in the set of sequential stages, each predetermined group of data inputs is read from the same address line in the first memory or the second memory.

2. The SoC according to claim 1, wherein: The hardware accelerator includes a number theory transformation hardware accelerator.

3. The SoC according to claim 1, wherein: The processing operations include butterfly operations performed to compute number-theoretic transforms.

4. The SoC according to claim 1, wherein: The hardware accelerator is configured to compute, for a first stage of the one or more stages, each predetermined set of data inputs as part of a processing sequence based on the predetermined set of data inputs for a subsequent second stage.

5. The SoC according to claim 4, wherein: The processing order is non-sequential.

6. The SoC according to claim 3, wherein: For an initial stage in the set of sequential stages, each data input in the predetermined set of data inputs comprises a coefficient pair of a polynomial through which a number-theoretic transformation is to be computed by the butterfly operation.

7. The SoC according to claim 1, wherein: For each stage in the set of sequential stages, the predetermined set of data inputs is read from one of the first memory or the second memory, while the predetermined set of processed data outputs is written to the other of the first memory or the second memory.

8. The SoC according to claim 1, wherein: For each stage in the set of sequential stages, the predetermined set of data inputs is read from the first memory or the second memory, and the predetermined set of processed data outputs is subsequently written to the same memory in the first memory or the second memory.

9. The SoC according to claim 1, wherein: The first memory and the second memory include one or more single-port static random access memories.

10. The SoC according to claim 1, further comprising: another hardware accelerator, and Wherein, for each stage in the set of sequential stages, the hardware accelerator and the another hardware accelerator are each configured to concurrently perform processing operations on a corresponding predetermined group of data inputs in the set of data inputs.

11. The SoC according to claim 1, wherein: The first memory and the second memory have a plurality of memory lines, each memory line of the plurality of memory lines having a width equal to twice a width of each data input in the set of data inputs and equal to twice a width of each data output in the set of processed data outputs.

12. The SoC according to claim 1, wherein: The set of sequential stages includes a first stage and a second stage, and upon completion of the first stage, the first memory or the second memory stores each respective predetermined set of processed data output from the first stage, and Wherein each respective predetermined set of processed data output from the first stage is stored at the same respective address line in the first memory or the second memory.

13. A computer-implemented method comprising: for each stage in the set of sequential stages, performing a processing operation on each of the predetermined set of data inputs in the set of data inputs and outputting a corresponding predetermined set of processed data outputs in the set of processed data outputs; storing one or more of the predetermined set of processed data outputs in a buffer before being stored in the first memory or the second memory, wherein, for one or more stages in said set of sequential stages, said predetermined set of data inputs is formed by said predetermined set of processed data outputs output by a preceding stage; as well as Controlling one or more of the predetermined groups of processed data outputs stored in the buffer to be transferred to the first memory or the second memory so that for one or more stages, each corresponding predetermined group of data inputs is read from the same address line in the first memory or the second memory.

14. The computer-implemented method of claim 13, wherein: The hardware accelerator includes a number theory transformation hardware accelerator.

15. The computer-implemented method of claim 13, wherein: The processing operations include butterfly operations performed to compute number-theoretic transforms.

16. The computer-implemented method of claim 13, further comprising: For a first stage of the one or more stages, each predetermined set of data inputs is calculated as part of a processing sequence based on the predetermined set of data inputs for a subsequent second stage.

17. The computer-implemented method of claim 16, wherein: The processing order is non-sequential.

18. The computer-implemented method of claim 15, wherein: For an initial stage in the set of sequential stages, each data input in the predetermined set of data inputs comprises a coefficient pair of a polynomial through which a number-theoretic transformation is to be computed by the butterfly operation.

19. The computer-implemented method of claim 13, wherein: For each stage in the set of sequential stages, the predetermined set of data inputs is read from one of the first memory or the second memory, while the predetermined set of processed data outputs is written to the other of the first memory or the second memory.

20. The computer-implemented method of claim 13, wherein: For each stage in the set of sequential stages, the predetermined set of data inputs is read from the first memory or the second memory, and the predetermined set of processed data outputs is subsequently written to the same memory in the first memory or the second memory.

21. The computer-implemented method of claim 13, wherein: The first memory and the second memory include one or more single-port static random access memories.

22. The computer-implemented method of claim 13, further comprising: For each stage in the set of sequential stages, processing operations are concurrently performed by the hardware accelerator and another hardware accelerator on a corresponding predetermined group of data inputs in the set of data inputs.