ML-DSA-oriented low-resource number theory transform acceleration system and method

CN121193420BActive Publication Date: 2026-08-28中电信量子信息科技集团有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511390872.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-08-28
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

[0005]本申请的目的在于,针对上述现有技术中的不足,提供一种面向ML-DSA的低资源数论变换加速系统及方法,以解决现有技术中迭代次数多,处理效率低的问题

Benefits of technology

[0020]The beneficial effects of this application are as follows: The low-resource number theory transformation acceleration system for ML-DSA includes a control addressing module, a first storage module, a second storage module, and a computation module. The first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit. Both the first and second storage units are used to store polynomial coefficients, and each coefficient address in both the first and second storage units stores two coefficients. The third storage unit stores twiddle factors, and the fourth storage unit stores inverse twiddle factors. The control addressing module, based on the current computation mode, the current computation round, and the current computation count under the current computation round, determines multiple target coefficient addresses in the first or second storage unit, and determines multiple target twiddle factor addresses. The first and second computation units in the computation module read one coefficient from each target coefficient address and the twiddle factor from each target twiddle factor address, respectively, perform operations on the read coefficients, and store the operation results in the second or first storage unit at each target coefficient address. In this embodiment, the arithmetic module performs butterfly operations on multiple coefficients per clock cycle under the current in-memory architecture, achieving high throughput. Furthermore, the control addressing module determines the target coefficient address and target twiddle factor address based on the current operation round and the current operation count, thus simplifying the addressing method and doubling the read/write efficiency compared to existing technologies. Additionally, this embodiment utilizes a first and second storage unit for coefficient access, coupled with the pipelined operation design of the arithmetic module, resulting in low read/write latency, reduced storage unit overhead, and a simple and efficient overall process. The system of this embodiment enables efficient signature and verification to meet real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121193420B_ABST
    Figure CN121193420B_ABST
Patent Text Reader

Abstract

The application provides an ML-DSA-oriented low-resource number theory transformation acceleration system and method. The system comprises a control addressing module, a first storage module, a second storage module and an operation module. The first storage unit and the second storage unit in the storage module are used for storing the coefficients of a polynomial, and the third storage unit and the fourth storage unit are used for storing rotation factors and inverse rotation factors. The control addressing module determines a plurality of target coefficient addresses and a plurality of target rotation factor addresses based on the current operation mode, the current operation round and the current operation count. The first operation unit and the second operation unit in the operation module read the coefficients and the rotation factors respectively, operate the read coefficients, and store the operation results to the target coefficient addresses. The addressing method is simple, the operation unit accesses multiple coefficients at a time, and the reading and writing efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of post-quantum cryptography, and more specifically, to a low-resource number theory transformation acceleration system and method for ML-DSA. Background Technology

[0002] With the rapid development of quantum computing technology, traditional public-key cryptography faces a significant threat of being broken by quantum algorithms, driving the research and standardization of post-quantum cryptography. Among these, cryptographic schemes based on lattice-based hard problems have become the mainstream direction of post-quantum cryptography due to their provable security and good performance balance. The Module-Lattice Digital Signature Algorithm (ML-DSA), as a lattice-based digital signature scheme, is suitable for various scenarios requiring identity authentication and data integrity assurance. However, the signature generation and verification process of ML-DSA relies on a large number of polynomial ring multiplication operations, and its computational complexity directly limits the deployment of the algorithm in resource-constrained scenarios such as IoT terminals and edge computing devices. Therefore, a method to reduce the computational complexity of polynomial multiplication is urgently needed.

[0003] In existing technologies, number-theoretic transformation (NTT) is commonly used to reduce the computational complexity of polynomial multiplication. NTT transforms the polynomial from the coefficient domain to the frequency domain and performs point-wise multiplication (PWM), thereby reducing the original complexity from... Down to This method has become the core technology for hardware acceleration of lattice cryptography algorithms.

[0004] However, in existing technologies, the NTT transformation process involves a large number of iterations, with each loop containing a large number of repetitive butterfly operations, resulting in a lengthy overall process. This leads to high time consumption for signing and verification, making it difficult to meet real-time requirements. Summary of the Invention

[0005] The purpose of this application is to provide a low-resource number theory transformation acceleration system and method for ML-DSA, addressing the shortcomings of the prior art, so as to solve the problems of high iteration count and low processing efficiency in the prior art.

[0006] To achieve the above objectives, the technical solution adopted in this application is as follows: In a first aspect, this application provides a low-resource number theory transformation acceleration system for ML-DSA, the low-resource number theory transformation acceleration system for ML-DSA includes: a control addressing module, a first storage module, a second storage module and a computation module, the first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit. The first and second storage units are both used to store the coefficients of the polynomial, and each coefficient address in the first and second storage units stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor. The control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current operation mode, the current operation round and the current operation count under the current operation round, and to determine multiple target rotation factor addresses. The first and second arithmetic units in the arithmetic module are used to read a coefficient from each of the target coefficient addresses and a rotation factor from each of the target rotation factor addresses, respectively, perform calculations on the read coefficients, and store the calculation results in the second storage unit or each of the target coefficient addresses in the first storage unit.

[0007] Optionally, the step of determining multiple target coefficient addresses in the first or second storage unit based on the current operation mode, the current operation round, and the current operation count under the current operation round, and determining multiple target twiddle factor addresses, includes: The address of the first target coefficient is determined based on the current operation round and the current operation count; The second target coefficient address is determined based on the first target coefficient address and the current calculation round; Based on the current operation mode, and according to the current operation round and the current operation count, the address of the first target rotation factor is determined; Based on the current operation mode and according to the address of the first target rotation factor, the address of the second target rotation factor is determined.

[0008] Optionally, determining the address of the first target coefficient based on the current operation round and the current operation count includes: If the current operation round is not a first preset high round, and the current operation round is not a first preset low round, then based on the formula...

[0009] The address of the first target coefficient is calculated, where, The first target coefficient address is defined as cnt, the current operation count is defined as cnt, and the current operation round is defined as stage. If the current operation round is the first preset low round, then the current operation count is used as the address of the first target coefficient; If the current operation round is the first preset high round, then the value of shifting the current operation count one bit to the left is used as the address of the first target coefficient.

[0010] Optionally, determining the second target coefficient address based on the first target coefficient address and the current computation round includes: If the current calculation round is the second preset low round, then based on the formula

[0011] The address of the second target coefficient is calculated, where, For the address of the second target coefficient, The address of the first target coefficient is `stage`, and `stage` is the current computation round. If the current operation round is the second preset high round, then the sum of the first target coefficient address and 1 is used as the second target coefficient address.

[0012] Optionally, determining the address of the first target rotation factor based on the current operation mode and according to the current operation round and the current operation count includes: If the current operation mode is a positive operation mode, and the current operation round is not a third preset high round, and the current operation round is not a third preset low round, then based on the formula...

[0013] The address of the first target rotation factor is calculated, where, The first target rotation factor address is given, cnt is the current operation count, and stage is the current operation round. If the current operation mode is positive operation mode, and the current operation round is the third preset low round, then 1 is used as the address of the first target rotation factor; If the current operation mode is positive operation mode and the current operation round is the third preset high round, then 128 and the value of the current operation count shifted one position to the left are used as the address of the first target rotation factor.

[0014] Optionally, determining the second target rotation factor address based on the first target rotation factor address includes: If the current operation mode is positive operation mode, and the current operation round is the fourth preset low round, then the first target rotation factor address is used as the second target rotation factor address; If the current operation mode is positive operation mode and the current operation round is the fourth preset high round, then the sum of 128 and the first target rotation factor address is used as the second target rotation factor address.

[0015] Optionally, the step of reading one coefficient from each of the target coefficient addresses and one coefficient from each of the target rotation factor addresses, performing calculations on the read coefficients, and storing the calculation results in the second storage unit or each of the target coefficient addresses in the first storage unit includes: The first arithmetic unit reads a first coefficient from the first target coefficient address of the first storage unit, reads a second coefficient from the second target coefficient address of the first storage unit, and reads a first rotation factor from the first target rotation factor address; The first calculation unit performs calculations based on the first coefficient, the second coefficient, and the first rotation factor to obtain a first calculation result and a second calculation result. The first calculation result is stored in the first target coefficient address in the second storage unit, and the second calculation result is stored in the second target coefficient address in the second storage unit. The second arithmetic unit reads the third coefficient from the first target coefficient address of the first storage unit, reads the fourth coefficient from the second target coefficient address of the first storage unit, and reads the second rotation factor from the second target rotation factor address; The second operation unit performs operations based on the third coefficient, the fourth coefficient, and the second rotation factor to obtain a third operation result and a fourth operation result. The third operation result is stored in the first target coefficient address in the second storage unit, and the fourth operation result is stored in the second target coefficient address in the second storage unit.

[0016] Optionally, the first calculation unit performs calculations based on the first coefficient, the second coefficient, and the first rotation factor to obtain a first calculation result and a second calculation result, including: The first arithmetic unit calculates the multiplication result of the second coefficient and the first rotation factor, and compresses the multiplication result through modulo reduction to obtain a compressed result; Calculate the modulus sum of the first coefficient and the compression result, and use the modulus sum as the first calculation result; Calculate the reduction of the first coefficient and the compression result, and use the reduction result as the second calculation result.

[0017] Optionally, the calculation of the multiplication result of the second coefficient and the first rotation factor includes: Based on the Karasuba divide-and-conquer algorithm, and using DSP primitives in a field-programmable gate array (FPGA), the multiplication result of the second coefficient and the first twitch factor is calculated.

[0018] Thirdly, this application provides a low-resource number-theoretic transformation acceleration method for ML-DSA, the method being applied to the low-resource number-theoretic transformation acceleration system for ML-DSA as described in the first aspect, the method comprising: The first and second storage units are both used to store the coefficients of the polynomial, and each coefficient address in the first and second storage units stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor. The control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current operation mode, the current operation round and the current operation count under the current operation round, and to determine multiple target rotation factor addresses. The first and second arithmetic units in the arithmetic module are used to read a coefficient from each of the target coefficient addresses and a rotation factor from each of the target rotation factor addresses, respectively, perform calculations on the read coefficients, and store the calculation results in the second storage unit or each of the target coefficient addresses in the first storage unit.

[0019] Thirdly, this application provides an electronic device on which a low-resource number-theoretic transformation acceleration system for ML-DSA as described in the first aspect is deployed.

[0020] The beneficial effects of this application are as follows: The low-resource number theory transformation acceleration system for ML-DSA includes a control addressing module, a first storage module, a second storage module, and a computation module. The first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit. Both the first and second storage units are used to store polynomial coefficients, and each coefficient address in both the first and second storage units stores two coefficients. The third storage unit stores twiddle factors, and the fourth storage unit stores inverse twiddle factors. The control addressing module, based on the current computation mode, the current computation round, and the current computation count under the current computation round, determines multiple target coefficient addresses in the first or second storage unit, and determines multiple target twiddle factor addresses. The first and second computation units in the computation module read one coefficient from each target coefficient address and the twiddle factor from each target twiddle factor address, respectively, perform operations on the read coefficients, and store the operation results in the second or first storage unit at each target coefficient address. In this embodiment, the arithmetic module performs butterfly operations on multiple coefficients per clock cycle under the current in-memory architecture, achieving high throughput. Furthermore, the control addressing module determines the target coefficient address and target twiddle factor address based on the current operation round and the current operation count, thus simplifying the addressing method and doubling the read / write efficiency compared to existing technologies. Additionally, this embodiment utilizes a first and second storage unit for coefficient access, coupled with the pipelined operation design of the arithmetic module, resulting in low read / write latency, reduced storage unit overhead, and a simple and efficient overall process. The system of this embodiment enables efficient signature and verification to meet real-time requirements. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the structure of a low-resource number-theoretic transform acceleration system for ML-DSA provided in an embodiment of this application; Figure 2 This is a storage diagram of a computing unit provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a first arithmetic unit provided in an embodiment of this application; Figure 4This is a flowchart illustrating a low-resource number-theoretic transformation acceleration method for ML-DSA provided in an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0024] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0025] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0026] In existing technologies, NTT (Network Transformer) is commonly used to reduce the computational complexity of polynomial multiplication. NTT reduces the computational complexity by transforming the polynomial from the coefficient domain to the frequency domain and then performing PWM (Pulse Width Modulation). Down to This method has become the core technology for hardware acceleration of lattice-based cryptography algorithms. Field-Programmable Gate Arrays (FPGAs), due to their combination of reconfigurability and parallel computing capabilities, can flexibly adapt to the parallel operation characteristics of NTT (Network Transformation), and are considered an ideal platform for achieving efficient hardware acceleration of ML-DSA. However, in existing technologies, the NTT transformation process involves numerous iterations, with each loop containing a large number of repetitive butterfly operations, resulting in a lengthy overall process and high time consumption for signature and verification, making it difficult to meet real-time requirements. Furthermore, in FPGA-based NTT processes, the indexing of coefficient addresses is complex, resulting in high address read / write overhead. The storage of a large number of intermediate states of coefficients during the operation makes it impossible to achieve a lightweight design. Additionally, ML-DSA has a large modulus parameter, for example... As a result, modular multiplication is inefficient, and this process may consume a lot of FPGA on-chip lookup table (LUT) resources, which may also make it impossible to implement lightweight designs.

[0027] Based on this, this application proposes a low-resource number theory transformation acceleration system for ML-DSA. The control addressing module in this system determines multiple target coefficient addresses and multiple target twitch factor addresses in a first or second storage unit based on the current operation mode, the current operation round, and the current operation technique. The first and second operation units respectively read a coefficient from one target coefficient address and a coefficient from each target twitch factor address, perform operations on the read coefficients, and store the operation results in the second or first storage unit at each target coefficient address. In this application, the control addressing module efficiently searches for multiple target coefficient addresses and multiple target twitch factor addresses based on the current operation round and the current operation count. Due to the high address search efficiency, the NTT loop process is accelerated, and the signing and verification time is relatively small, which is insufficient to meet real-time requirements.

[0028] Next, we will introduce the ML-DSA of this application. As a standardized post-quantum signature scheme, ML-DSA's core operations rely on efficient computation on polynomial rings. Specifically, ML-DSA uses polynomial rings as... ,in, It is an integer residue class ring modulo q. In this embodiment, we take q=8380417 and a bit length of 23 as an example for illustration. It is a cycloid polynomial, which defines the reduction rules for polynomials in a ring, that is, n is the degree of the polynomial, which determines the number of coefficients of the polynomial in the ring. In this embodiment, n=256 is used for illustration.

[0029] Polynomial ring Multiplication is the core operation for ML-DSA signature generation and verification, but the complexity of direct computation is O(n log n). This makes it difficult to meet the efficiency requirements of resource-constrained scenarios. NTT is a key technology to reduce this complexity: NTT maps polynomials from the coefficient domain to the frequency domain, transforming polynomial multiplication into PWM, thus reducing the complexity to... The implementation of NTT relies on the primitive root of unity ζ. In this embodiment, ζ is 1753 as an example. The 2n-th primitive root of unity on the x-th surface, that is: And for any 0 <k<2n, .

[0030] To support efficient NTT operations, it is necessary to pre-calculate the powers of ζ. These values ​​form an array of zetas (rotation factors). This indicates that 8 bits of k are reversed, which is a crucial operation for NTT address generation, ensuring the correct alignment of coefficients in the frequency domain. The rotation factor is pre-calculated and stored in memory before algorithm execution, avoiding the resource overhead of real-time computation.

[0031] Next, refer to Figure 1 This paper introduces a low-resource number-theory transform acceleration system for ML-DSA. This system can be applied to FPGAs. Figure 1 This is a schematic diagram of a low-resource number-theoretic transformation acceleration system for ML-DSA provided in an embodiment of this application.

[0032] Optionally, the low-resource number theory transformation acceleration system for ML-DSA includes: a control addressing module, a first storage module, a second storage module, and a computation module. The first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit.

[0033] Optionally, the control addressing module may include a control unit and an addressing unit. The control unit may be a finite state machine (FSM) controller. The control unit controls the finite state machine of the NTT transformation process, the addressing unit determines the coefficient address and sends the coefficient address to the control unit, and the control unit receives the address sent by the addressing unit and sends the address to the first arithmetic unit and the second arithmetic unit accordingly.

[0034] Optionally, the control addressing module is connected to the first end of the first storage module, the first end of the second storage module, and the first end of the arithmetic module, respectively, and the second end of the arithmetic module is connected to the second end of the first storage module and the second end of the second storage module.

[0035] Optionally, both the first and second storage units are used to store the coefficients of the polynomial, and each coefficient address in the first and second storage units stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor.

[0036] Optionally, the coefficients of the polynomial are pre-stored in the first and second storage units. The contents stored in the first and second storage units can be the same before the first round. The first and second storage units have the same dual-port structure to transmit two data items simultaneously. During initialization, the width of the first and second storage units is instantiated to... If q is 8380417, then the width is 46, and the depth is half the degree of the polynomial. For example, if the degree of the polynomial is 256, then the depth of the first and second storage units is 128, and the corresponding address width is... For example, 128 units require 7 bits of binary numbering, which covers 0-127. Optionally, in the first and second storage units, one coefficient address can store two coefficients, namely 2k and 2k+1, where k can be a coefficient between 0 and 127.

[0037] Optionally, the first and second storage units form a ping-pong structure to achieve the theoretically lowest latency for coefficient read and write operations.

[0038] Optionally, the third and fourth storage units can pre-store the rotation factor and the inverse rotation factor, respectively, before starting the operation. Specifically, the width of the third and fourth storage units is instantiated during initialization. If q is 8380417, then the width is 23, the depth is the power of a polynomial, for example, 256, and the corresponding address width is... For example, 256 units require 8-bit binary numbers, which perfectly cover 0-255. Furthermore, both the third and fourth storage units are dual-port structures.

[0039] Optionally, the control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current operation mode, the current operation round and the current operation count under the current operation round, and to determine multiple target rotation factor addresses.

[0040] Optionally, the current operation mode can include a forward operation mode and an inverse operation mode. The forward operation mode performs number-theoretical transformations when a polynomial needs to be transformed from the coefficient domain to the frequency domain. This is implemented on the FPGA using a Cooley-Tukey (CT) architecture. Specifically, a large NTT is decomposed into multiple smaller NTTs, and the results of the sub-transformations are gradually merged through time decimation. The inverse operation mode performs inverse number-theoretical transformations when a pointwise product result in the frequency domain needs to be transformed back to the coefficient domain. This is implemented on the FPGA using a Gentleman-Sande (GS) architecture, synthesizing the transformation result through frequency decimation. It is dual to the CT structure in terms of the order of operations and the use of the twiddle factor.

[0041] Optionally, the total number of operation rounds can be calculated based on the polynomial degree. For example, if the polynomial degree is 256, then the total number of operation rounds is... Layers. As an optional implementation, the current round of computation can be determined by the control unit in the control addressing module. As the number of rounds increases, the total number of coefficients remains unchanged, but the grouping rule changes, that is, from a grouping rule of multiple groups with small intervals to a rule of fewer groups with large intervals.

[0042] Optionally, the control addressing module further includes a counting unit. The counting unit is used to count the current operation count in the current operation round and send the current operation count to the addressing module.

[0043] As an optional implementation, the control addressing module can determine multiple target coefficient addresses in the first or second storage unit based on the current operation round and the current operation count under the current operation round, and determine multiple target rotation factor addresses based on the current operation mode, the current operation round and the current operation count under the current operation round.

[0044] Optionally, after determining multiple target coefficient addresses and multiple target rotation factor addresses, the control addressing module can send these addresses to the corresponding first and second arithmetic units, respectively, to schedule the state of the arithmetic modules. This allows the first and second arithmetic units to read the corresponding coefficients and rotation factors from the storage units based on the target coefficient addresses and rotation factor addresses, thereby completing the calculation. During the calculation, the counter increments by 1 in each clock cycle. The control module determines the cycle number based on the current count of the counter and controls the state transition. Simultaneously, the addressing unit determines the target coefficient address and target rotation factor address based on the current cycle number and the current count.

[0045] Optionally, the first and second arithmetic units in the arithmetic module are used to read a coefficient from each target coefficient address and a rotation factor from each target rotation factor address, respectively, perform calculations on the read coefficients, and store the calculation results in the second storage unit or the target coefficient addresses in the first storage unit.

[0046] It is worth noting that the storage units used for reading and storing are different within the same computation cycle. For example, in the current cycle, the computation module reads the coefficients from the first storage unit and stores the computation result in the second storage unit.

[0047] As an optional implementation, the first and second arithmetic units can read one coefficient from each target coefficient address from either the first or second storage unit, and according to the current arithmetic mode, read the rotation factor from either the third or fourth storage unit according to the target rotation factor address via a selector. Specifically, if the current arithmetic mode is forward arithmetic mode, the rotation factor is read from each of the third storage units; if the current arithmetic mode is inverse arithmetic mode, the rotation factor is read from each of the fourth storage units. If there are two target coefficient addresses, the first and second arithmetic units read two coefficients from either the first or second storage unit, and read one rotation factor from the target rotation factor address, thus achieving a butterfly transformation with four input coefficients. Optionally, the first storage module further includes a coefficient processing unit connected to the control addressing module, which, based on control instructions from the control addressing module, enables the first and second arithmetic units to read and store coefficients.

[0048] Specifically, in the current round of computation, the first computation unit in the computation module reads a coefficient from each target coefficient address from either the first or second storage unit, and reads a rotation factor from each target rotation factor address from either the third or fourth storage unit. The first computation unit performs computation based on the read coefficients, obtaining two results: result A and result B. The second computation unit in the computation module reads a coefficient from each target coefficient address from either the first or second storage unit, and reads a rotation factor from each target rotation factor address from either the third or fourth storage unit. The second computation unit performs computation based on the read coefficients, obtaining two results: result C and result D. Result A, result B, result C, and result D are all used as coefficients and stored in either the second or first storage unit according to the target coefficient addresses.

[0049] The above process represents an operation under a count in the current operation cycle. While the first and second operation units read coefficients, they simultaneously store the result of the previous operation count, and while storing coefficients for the current operation count, they simultaneously read coefficients for the next operation count. In other words, the first and second operation units continuously read data from one storage unit while simultaneously storing data in another. This pipelined operation method allows for butterfly operations on multiple coefficients to be completed in a single clock cycle.

[0050] After completing the current round of computation, that is, after completing one round of computation on all coefficients in memory, the memory units used for reading and storing are flipped, and the next round of computation is executed. For example, in the first round of computation, the computation module reads coefficients from the first memory unit and stores the computation result in the second memory unit. In the second round of computation, the computation module reads coefficients from the second memory unit and stores the computation result in the first memory unit. In the third round of computation, the computation module reads coefficients from the first memory unit and stores the computation result in the second memory unit. In this way, low-latency coefficient reading and writing can be achieved, and the overhead of memory units is reduced.

[0051] For example, Figure 2 This is a storage diagram of a computing unit provided in an embodiment of this application. For example... Figure 2 As shown, two coefficients are stored in one address of the first storage unit RAM0 and the second storage unit RAM1. The storage units used for reading and storing are reversed in 2k (k=1,2,3,4) and 2k+1 (k=0,1,2,3) cycle iterations. Specifically, in the 2k+1 cycle iteration, the first arithmetic unit BFU0 and the second arithmetic unit BFU1 read the coefficients from the first storage unit RAM0 and simultaneously write the coefficients to the second storage unit RAM1. In the 2k cycle iteration, the first arithmetic unit BFU0 and the second arithmetic unit BFU1 read the coefficients from the second storage unit RAM1 and simultaneously write the coefficients to the first storage unit RAM0. (See figure) These are the address pairs for the dual ports of the first storage unit RAM0. These are the address pairs for the dual ports of the second storage unit RAM1.

[0052] In this embodiment, the low-resource number theory transformation acceleration system for ML-DSA includes a control addressing module, a first storage module, a second storage module, and a computation module. The first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit. Both the first and second storage units store polynomial coefficients, and each coefficient address in both units stores two coefficients. The third storage unit stores twiddle factors, and the fourth storage unit stores inverse twiddle factors. The control addressing module, based on the current computation mode, the current computation round, and the current computation count in the current round, determines multiple target coefficient addresses in the first or second storage unit, and also determines multiple target twiddle factor addresses. The first and second computation units in the computation module read one coefficient from each target coefficient address and the twiddle factor from each target twiddle factor address, perform operations on the read coefficients, and store the results in the second or first storage unit. In this embodiment, the computation module implements butterfly operations on multiple coefficients per clock cycle under the current memory-computing architecture, achieving high throughput. Furthermore, the addressing module determines the target coefficient address and target rotation factor address based on the current operation round and the current operation count. Therefore, the addressing method is simple, and the read / write efficiency is doubled compared to existing technologies. Additionally, this embodiment uses a first storage unit and a second storage unit for coefficient access. Combined with the pipelined operation design of the operation module, read / write latency is low, storage unit overhead is reduced, and the overall process is simple and efficient. The system of this embodiment can achieve efficient signature and verification to meet real-time requirements.

[0053] Next, we will introduce the method by which the control addressing module determines the addresses of multiple target coefficients in the first or second memory unit and the addresses of multiple target rotation factors based on the current operation mode, the current operation round, and the current operation count under the current operation round.

[0054] As an optional implementation, the truncated high-order bits of the current operation count correspond to the address increment between butterfly operation blocks in the current operation round, and the low-order bits of the current operation count correspond to the address increment within a single operation unit block. The high-order bits are shifted according to the corresponding block spacing, and the shifted result is added to the low-order bits to obtain the target twiddle factor address. Each butterfly operation block can perform one counted butterfly operation.

[0055] The following describes one implementation method for the above approach: Optionally, the address of the first target coefficient can be determined based on the current operation round and the current operation count.

[0056] Optionally, different calculation rounds can correspond to different determination methods. Specifically, the calculation rounds can be divided into high rounds, medium rounds, and low rounds, and the methods for determining the address of the first target coefficient are different for high rounds, medium rounds, and low rounds.

[0057] Optionally, the second target coefficient address is determined based on the first target coefficient address and the current operation round.

[0058] Optionally, the larger the address of the first target coefficient, the larger the address of the second target coefficient.

[0059] Optionally, the address of the first target rotation factor is determined based on the current operation mode and according to the current operation round and the current operation count.

[0060] The current operation module can be in either forward or inverse operation mode. The forward and inverse operation modes determine the address of the first target rotation factor in opposite ways. Specifically, the forward operation mode uses a CT structure, while the inverse operation mode uses a GS structure. The inverse operation mode only needs to complete the inverse process of the forward operation mode, making the addressing process simple.

[0061] Optionally, the second target rotation factor address is determined based on the current operation mode and according to the first target rotation factor address.

[0062] Optionally, the larger the address of the first target rotation factor, the larger the address of the second rotation factor.

[0063] In this embodiment, a first target coefficient address is determined based on the current operation round and the current operation count. A second target coefficient address is then determined based on the first target coefficient address and the current operation round. Finally, a first target rotation factor address is determined based on the current operation mode, the current operation round, and the current operation count. The second target rotation factor address is then determined based on the current operation mode and the first target rotation factor address. This embodiment efficiently determines the target coefficient address and the target rotation factor address using this simple method.

[0064] Next, we will explain the specific method for determining the address of the first target coefficient.

[0065] Optionally, if the current calculation round is not the first preset higher round and the current calculation round is not the first preset lower round, then based on the formula...

[0066] The address of the first target coefficient is calculated, where, `cnt` is the address of the first target coefficient, `stage` is the current operation count, and `stage` is the current operation round.

[0067] Optionally, if the current operation round is the first preset low round, then the current operation count is used as the address of the first target coefficient.

[0068] Optionally, if the current operation round is the first preset high round, then the value of shifting the current operation count one bit to the left is used as the address of the first target coefficient.

[0069] For example, when the total number of operation rounds is 8, the first preset high round can be the 7th round and the 8th round, and the first preset low round can be the 1st round.

[0070] Optionally, the address of the first target coefficient can be calculated using the following formula (1):

[0071] Specifically, when the current operation round is the 2nd to 6th round, the high-order part of the binary representation of cnt is shifted to the left by stage+1 bits, and the low-order part is added to obtain the address of the first target coefficient. When the current operation round is the 1st round, cnt is used as the address of the first target coefficient. When the current operation round is the 7th or 8th round, cnt is shifted to the left by 1 bit to obtain the address of the first target coefficient.

[0072] For example, if the polynomial degree is 256 and the current operation round is 2, then 7-stage=5, that is, the high-order bits are taken as cnt[5:5], 6-stage=4, that is, the low-order bits are taken as cnt[4:0], stage+1=3, that is, the high-order bits are shifted left by three bits. If the current binary value of cnt is 00101101, which is 45 in decimal, then by using the above formula (1), the high-order bits are 0, so the high-order bits are shifted left by 3 bits to get 0, and the low-order bits cnt[4:0] are 10101, which is 21 in decimal, thus obtaining the address of the first target coefficient. In other words, when the current operation count is 45, the corresponding address of the first target coefficient is 21.

[0073] In this embodiment, by dividing the address of the first target coefficient according to the current round of operation, the operation of the address index only includes truncation, shifting and a small amount of addition, which greatly reduces the complexity and enables the hardware to efficiently obtain the coefficients required for calculation.

[0074] Next, we will explain the specific method for determining the address of the second target coefficient.

[0075] Optionally, if the current calculation round is the second preset lower round, then based on the formula

[0076] The address of the second target coefficient is calculated, where, For the address of the second target coefficient, `stage` is the address of the first target coefficient, and `stage` is the current computation round.

[0077] Optionally, if the current operation round is the second preset high round, the sum of the first target coefficient address and 1 is used as the second target coefficient address.

[0078] For example, when the total number of operation rounds is 8, the second preset high round can be the 8th round, and the second preset low round can be the 1st to 7th rounds.

[0079] Alternatively, the address of the second target coefficient can be calculated using the following formula (2):

[0080] For example, if the current operation round is 2 and the address of the first target coefficient is 21, then the address of the second target coefficient is... =53.

[0081] In this embodiment, by dividing the address of the second target coefficient according to the current round of operation, the operation of the address index only involves a small number of additions, which greatly reduces the complexity and allows the hardware to efficiently obtain the coefficients required for calculation.

[0082] Next, we will explain the specific method for determining the address of the first target rotation factor.

[0083] Optionally, if the current operation mode is positive operation mode, and the current operation round is not the third preset high round, and the current operation round is not the third preset low round, then based on the formula...

[0084] The address of the first target rotation factor is calculated, where, `cnt` is the address of the first target rotation factor, `stage` is the current operation count, and `stage` is the current operation round.

[0085] Optionally, if the current operation mode is positive operation mode and the current operation round is the third preset low round, then 1 is used as the address of the first target rotation factor.

[0086] Optionally, if the current operation mode is positive operation mode and the current operation round is the third preset high round, then 128 and the value shifted one bit to the left of the current operation count are used as the address of the first target rotation factor.

[0087] For example, when the total number of operation rounds is 8, the third preset high round can be the 8th round, and the third preset low round can be the 1st round.

[0088] Optionally, when the current operation mode is positive operation mode, the address of the first target rotation factor can be calculated using the following formula (3):

[0089] Specifically, when the current calculation round is round 2-7, the calculation is performed. And compare the calculation result with the binary representation of cnt. The two values ​​are added together to obtain the address of the first target rotation factor. When the current operation round is 1, 1 is used as the address of the first target rotation factor. When the current operation round is 8, the binary value of cnt is shifted left by one bit and 128 is added to obtain the address of the first target rotation factor.

[0090] As an optional implementation, if the current operation mode is the inverse operation mode, the method for determining the forward operation mode is reversed to obtain the address of the first target rotation factor. For example, if the current operation mode is the inverse operation mode and the current operation round is 1, then the address of the first target rotation factor is... .

[0091] In this embodiment, by dividing the address of the first target rotation factor according to the current operation round, the operation of the address index only includes a small number of truncation, shift and addition operations, which greatly reduces the complexity and enables the hardware to efficiently obtain the rotation factor required for calculation.

[0092] Next, we will explain the specific method for determining the address of the second target rotation factor.

[0093] Optionally, if the current operation mode is positive operation mode and the current operation round is the fourth preset low round, then the address of the first target rotation factor is used as the address of the second target rotation factor.

[0094] Optionally, if the current operation mode is positive operation mode and the current operation round is the fourth preset high round, then the sum of 128 and the first target rotation factor address is used as the second target rotation factor address.

[0095] For example, when the total number of operation rounds is 8, the fourth preset high round can be the 8th round, and the fourth preset low round can be the 1st to 7th rounds.

[0096] Optionally, when the current operation mode is positive operation mode, the address of the first target rotation factor can be calculated using the following formula (3): For example, when the total number of operation rounds is 8, the third preset high round can be the 8th round, and the third preset low round can be the 1st round.

[0097] Optionally, when the current operation mode is positive operation mode, the address of the second target rotation factor can be calculated using the following formula (4):

[0098] in, This is the address of the second target rotation factor.

[0099] As an optional implementation, if the current operation mode is the inverse operation mode, the method for determining the forward operation mode is reversed to obtain the address of the second target rotation factor. For example, if the current operation mode is the inverse operation mode and the current operation round is 1, then the address of the second target rotation factor is... .

[0100] In this embodiment, by dividing the address of the second target rotation factor according to the current operation round, the operation of the address index only involves a small number of additions, which greatly reduces the complexity and enables the hardware to efficiently obtain the rotation factor required for calculation.

[0101] In existing technologies, NTT designs typically require rearranging the addresses of the coefficient array after completing one loop. However, the addressing method in this application maintains the unchanged geometric structure of the coefficients, avoiding the rearrangement operation and thus further improving addressing and read / write efficiency.

[0102] Next, the operation process of the operation module will be described in detail.

[0103] Specifically, the first and second arithmetic units respectively read a coefficient from each target coefficient address and a coefficient from each target rotation factor address, perform calculations on the read coefficients, and store the calculation results in the second storage unit or the target coefficient addresses in the first storage unit. The specific process is as follows: Optionally, the first arithmetic unit reads the first coefficient from the first target coefficient address of the first storage unit, reads the second coefficient from the second target coefficient address of the first storage unit, and reads the first rotation factor from the first target rotation factor address.

[0104] Specifically, if the current operation mode is the forward operation mode, the first rotation factor is read from the first target rotation factor address in the third storage unit; if the current operation mode is the reverse operation mode, the first rotation factor is read from the first target rotation factor address in the fourth storage unit.

[0105] Specifically, after determining the address of the first target coefficient, the address of the second target coefficient, and the address of the first target rotation factor, the control addressing module controls the first arithmetic unit to read the coefficients and the rotation factor.

[0106] Optionally, the first arithmetic unit performs calculations based on the first coefficient, the second coefficient, and the first rotation factor to obtain a first calculation result and a second calculation result, and stores the first calculation result in the first target coefficient address in the second storage unit, and stores the second calculation result in the second target coefficient address in the second storage unit.

[0107] Optionally, each address in the first and second storage units stores two coefficients. Therefore, after determining the first and second target coefficient addresses, the control addressing module obtains four coefficients, and then controls which coefficient the first and second arithmetic units to use for calculation.

[0108] Optionally, the first arithmetic unit performs NTT operations. Specifically, if the current arithmetic mode is positive arithmetic mode, then the first arithmetic unit implements... The operation. In positive operation mode, A is the first operation result, B is the second operation result, a is the first coefficient, b is the second coefficient, and z is the first rotation factor.

[0109] If the current operation mode is the inverse operation mode, then the first operation unit implements... The operation is as follows. In the inverse operation mode, 'a' is the result of the first operation, 'b' is the result of the second operation, 'A' is the first coefficient, and 'B' is the second coefficient. is the first rotation factor.

[0110] After obtaining the first and second operation results, the control addressing module stores the first and second operation results into the second storage unit based on the first target coefficient address and the second target coefficient address, respectively.

[0111] Optionally, the second arithmetic unit reads the third coefficient from the first target coefficient address of the first storage unit, reads the fourth coefficient from the second target coefficient address of the first storage unit, and reads the second rotation factor from the second target rotation factor address.

[0112] Specifically, after determining the address of the first target coefficient, the address of the second target coefficient, and the address of the first target rotation factor, the control addressing module controls the second arithmetic unit to read the coefficients and the rotation factor.

[0113] Optionally, the second operation unit performs operations based on the third coefficient, the fourth coefficient, and the second rotation factor to obtain the third operation result and the fourth operation result, and stores the third operation result in the first target coefficient address in the second storage unit, and stores the fourth operation result in the second target coefficient address in the second storage unit.

[0114] The operation process of the second operation unit is the same as that of the first operation unit, and will not be described again here.

[0115] In this embodiment, the first arithmetic unit reads the first coefficient, the second coefficient, and the first rotation factor, and obtains the first and second arithmetic results. The first and second arithmetic results are then stored in the second storage unit. The second arithmetic unit reads the third coefficient, the fourth coefficient, and the second rotation factor, and obtains the third and fourth arithmetic results. The third and fourth arithmetic results are then stored in the second storage unit. The above method achieves a 4-coefficient butterfly operation with high throughput and efficient NTT acceleration.

[0116] Next, taking the first arithmetic unit as an example, the structure of the first arithmetic unit will be introduced. Optionally, the structure of the second arithmetic unit can be the same as that of the first arithmetic unit.

[0117] Optionally, the first arithmetic unit calculates the multiplication result of the second coefficient and the first rotation factor, and compresses the multiplication result through modular reduction to obtain a compressed result.

[0118] Optionally, during the operation of the first arithmetic unit, the data is aligned under each clock cycle, thereby forming a single-stage pipeline, so that the first arithmetic unit under any single clock cycle can receive new coefficients and rotation factors to perform butterfly operations.

[0119] Optionally, the first arithmetic unit can calculate the multiplication result of the second coefficient and the first twitch factor using a multiplier. The multiplier supports 24-bit * 24-bit calculations. The multiplication result can be compressed using a compressor to obtain a compressed result.

[0120] Optionally, the first and second coefficients can be 23 bits, and the multiplication result of the second coefficient and the first twitch factor can be 46 bits. The compressor can compress the 46-bit multiplication result into 23 bits.

[0121] Optionally, during the processing of the multiplier and compressor, the first coefficient is pipelined and timed to ensure clock synchronization.

[0122] Optionally, the first coefficient and the compression result are calculated and the result of the modulus addition is used as the first operation result.

[0123] Optionally, the first arithmetic unit can perform modular addition calculations using a modular adder.

[0124] Optionally, the first coefficient and the modulus reduction of the compression result are calculated, and the modulus reduction result is used as the second operation result.

[0125] Optionally, the first arithmetic unit can perform modular addition calculations using a modular subtractor.

[0126] In this embodiment, the first arithmetic unit calculates the multiplication result of the second coefficient and the first twitch factor, and compresses the multiplication result through modulo reduction to obtain a compressed result. It then calculates the modulo addition of the first coefficient and the compressed result, using the modulo addition result as the first arithmetic result. Finally, it calculates the modulo subtraction of the first coefficient and the compressed result, using the modulo subtraction result as the second arithmetic result. This embodiment achieves low-power and high-efficiency NTT acceleration through deep pipeline design.

[0127] Optionally, in the above process, the process of calculating the multiplication result of the second coefficient and the first rotation factor is as follows: based on the Karasuba divide-and-conquer algorithm, and using the digital signal processor (DSP) primitives in the FPGA, the multiplication result of the second coefficient and the first rotation factor is calculated.

[0128] The Karasuba divide-and-conquer algorithm is a fast algorithm for large integer multiplication. When handling large number multiplications, it uses addition and subtraction to avoid partial multiplications, thereby reducing the number of multiplications and the space resources required. In this embodiment, the large-area multiplication is split into three smaller-area multiplications, thus reducing the space resources by one-quarter.

[0129] Alternatively, Digital Signal Processor (DSP) primitives refer to the low-level instructions or operation units in a DSP used to implement basic signal processing operations. Using DSP primitives for computation can reduce the consumption of LUTs (Low-Level Units).

[0130] Optionally, the dedicated hardware resource units in the DSP of mainstream FPGAs support maximum direct multiplication. Therefore, the Karasuba divide-and-conquer algorithm is used to divide and conquer the second coefficient and the first twitch factor, and the multiplication operation is performed using DSP primitives.

[0131] In this embodiment, the multiplication result of the second coefficient and the first twitch factor is calculated using the Kara-Suba divide-and-conquer algorithm and the digital signal processor (DSP) primitives in the FPGA. The Kara-Suba divide-and-conquer algorithm can reduce the space resources consumed and decrease the consumption of LUTs.

[0132] Next, refer to Figure 3 The structure of the first arithmetic unit is described below. Figure 3 This is a schematic diagram of the structure of a first arithmetic unit provided in an embodiment of this application.

[0133] exist Figure 3In the first arithmetic unit, the first arithmetic unit reads the first coefficient *a*, the second coefficient *b*, and the first twitch factor *z*. All three coefficients are 23-bit integers. The multiplier *mul* performs optimized multiplication on the second coefficient *b* and the first twitch factor *z*, outputting a 46-bit multiplication result *u*. This result *u* is then input to the compressor *Barrett* for compression, yielding a 23-bit compressed result *t*. During the optimized multiplication and compression processes, the first coefficient is stored in registers for pipelined timing. Once the compressed result *t* is obtained and the clock is aligned, the first coefficient *a* and the compressed result *t* are input to the modulo adder *mod_add* and the modulo subtractor *mod_sub*, respectively, for computation, resulting in a 23-bit first operation result *A* and a 23-bit second operation result *B*.

[0134] In the multiplier mul*, a Carasuba divide-and-conquer approach is first used. Specifically, the second coefficient b is split into a high-order part bhigh and a low-order part blow, and the first twitch factor z is split into a high-order part zhigh and a low-order part zlow. Then, (bhigh × zhigh), (blow × zlow), and (bhigh + blow) × (zhigh + zlow) are calculated using the DSP primitive x*, respectively. The results are then combined through shifting and addition to obtain the multiplication result u, where <<24 indicates a left shift of 24 bits, and <<12 indicates a left shift of 12 bits. The multiplier replaces two large multiplications with three small multiplications and one addition, thereby reducing the hardware load.

[0135] In the Barrett compressor, the multiplication result u is progressively adjusted from 46 bits to 23 bits through truncation, shifting (<<), and multiple multiplications (mul*). After clock alignment, the modulo subtractor mod_sub finally corrects it to the modulo q range, outputting the compressed result u. Intermediate registers are pipelined and time-stamped, allowing the circuit to input and output simultaneously, thus improving throughput. Furthermore, the multiplier mul* can be reused during compression, improving both time and area efficiency.

[0136] In the mod_add function, the algorithm (first coefficient a + compression result u) mod q is run. If the sum is greater than or equal to q, then q is subtracted. In the mod_sub function, the algorithm (first coefficient a - compression result u) mod q is run. If the difference is less than q, then q is added.

[0137] In the inverse operation mode, a division by 2 operation needs to be performed after each round. Alternatively, the data in the result can be multiplied by 2 all at once after all rounds of operation. The operation is as follows. It is evident that the computational resources used by the inverse operation mode and PWM do not exceed the resources of the first arithmetic unit; therefore, the first arithmetic unit can be reused by adjusting the calling flow.

[0138] Based on the aforementioned dual-operation unit and its internal structure design, the low-resource number theory transformation acceleration system for ML-DSA can perform a butterfly operation with 4 coefficients per clock cycle. A single-level loop requires only 64 clock cycles to complete, and the complete NTT process consumes only 576cc. This demonstrates a significant improvement in NTT computational efficiency. Furthermore, the simplified internal structure design of the operation unit, the concise addressing method, and the in-memory computation structure design result in low system area resource consumption: approximately 800 LUTs, approximately 300 flip-flops, and 4 DSPs, thus achieving a lightweight NTT design.

[0139] Reference Figure 4 As shown in the embodiments of this application, a low-resource number-theoretic transformation acceleration method for ML-DSA is also provided. Figure 4 This is a flowchart illustrating a low-resource number-theoretical transform acceleration method for ML-DSA provided in an embodiment of this application. The method is applied to the aforementioned low-resource number-theoretical transform acceleration system for ML-DSA and includes: S401, the first storage unit and the second storage unit are both used to store the coefficients of the polynomial, and each coefficient address of the first storage unit and the second storage unit stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor.

[0140] S402, the control addressing module is connected to the first end of the first storage module, the first end of the second storage module and the first end of the arithmetic module respectively. The control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current arithmetic mode, the current arithmetic round and the current arithmetic count under the current arithmetic round, and to determine multiple target rotation factor addresses.

[0141] S403, the first and second arithmetic units in the arithmetic module are used to read a coefficient from each target coefficient address and a coefficient from each target rotation factor address, respectively, perform calculations on the read coefficients, and store the calculation results in the second storage unit or the target coefficient address in the first storage unit.

[0142] This application also provides an electronic device on which the above-described low-resource number-theoretic transformation acceleration system for ML-DSA is deployed.

[0143] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A low-resource number-theoretic transform acceleration system for ML-DSA, characterized in that, The low-resource number theory transformation acceleration system for ML-DSA includes: a control addressing module, a first storage module, a second storage module, and a computation module. The first storage module includes a first storage unit and a second storage unit, the second storage module includes a third storage unit and a fourth storage unit, and the computation module includes a first computation unit and a second computation unit. The first and second storage units are both used to store the coefficients of the polynomial, and each coefficient address in the first and second storage units stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor. The control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current operation mode, the current operation round and the current operation count under the current operation round, and to determine multiple target rotation factor addresses. The first and second operation units in the operation module are used to read a coefficient from each of the target coefficient addresses and a rotation factor from each of the target rotation factor addresses, respectively, perform operations on the read coefficients, and store the operation results in the second storage unit or each of the target coefficient addresses in the first storage unit; Based on the current operation mode, and according to the current operation round and the current operation count under the current operation round, the determination of multiple target coefficient addresses in the first storage unit or the second storage unit, and the determination of multiple target rotation factor addresses, includes: The address of the first target coefficient is determined based on the current operation round and the current operation count; The second target coefficient address is determined based on the first target coefficient address and the current calculation round; Based on the current operation mode, and according to the current operation round and the current operation count, the address of the first target rotation factor is determined; Based on the current operation mode and according to the address of the first target rotation factor, determine the address of the second target rotation factor; Determining the address of the first target coefficient based on the current operation round and the current operation count includes: If the current operation round is not a first preset high round, and the current operation round is not a first preset low round, then based on the formula... The address of the first target coefficient is calculated, where, The first target coefficient address is defined as cnt, the current operation count is defined as cnt, and the current operation round is defined as stage. If the current operation round is the first preset low round, then the current operation count is used as the address of the first target coefficient; If the current operation round is the first preset high round, then the value of shifting the current operation count one bit to the left is used as the address of the first target coefficient.

2. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 1, characterized in that, The step of determining the second target coefficient address based on the first target coefficient address and the current calculation round includes: If the current calculation round is the second preset low round, then based on the formula The address of the second target coefficient is calculated, where, For the address of the second target coefficient, The address of the first target coefficient is `stage`, and `stage` is the current computation round. If the current operation round is the second preset high round, then the sum of the first target coefficient address and 1 is used as the second target coefficient address.

3. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 1, characterized in that, The step of determining the address of the first target rotation factor based on the current operation mode and according to the current operation round and the current operation count includes: If the current operation mode is a positive operation mode, and the current operation round is not a third preset high round, and the current operation round is not a third preset low round, then based on the formula... The address of the first target rotation factor is calculated, where, The first target rotation factor address is given, cnt is the current operation count, and stage is the current operation round. If the current operation mode is positive operation mode, and the current operation round is the third preset low round, then 1 is used as the address of the first target rotation factor; If the current operation mode is positive operation mode and the current operation round is the third preset high round, then 128 and the value of the current operation count shifted one position to the left are used as the address of the first target rotation factor.

4. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 1, characterized in that, The step of determining the second target rotation factor address based on the first target rotation factor address includes: If the current operation mode is positive operation mode, and the current operation round is the fourth preset low round, then the first target rotation factor address is used as the second target rotation factor address; If the current operation mode is positive operation mode and the current operation round is the fourth preset high round, then the sum of 128 and the first target rotation factor address is used as the second target rotation factor address.

5. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 1, characterized in that, The step of reading one coefficient from each of the target coefficient addresses and one coefficient from each of the target rotation factor addresses, performing calculations on the read coefficients, and storing the calculation results in the second storage unit or each of the target coefficient addresses in the first storage unit includes: The first arithmetic unit reads a first coefficient from the first target coefficient address of the first storage unit, reads a second coefficient from the second target coefficient address of the first storage unit, and reads a first rotation factor from the first target rotation factor address; The first calculation unit performs calculations based on the first coefficient, the second coefficient, and the first rotation factor to obtain a first calculation result and a second calculation result. The first calculation result is stored in the first target coefficient address in the second storage unit, and the second calculation result is stored in the second target coefficient address in the second storage unit. The second arithmetic unit reads the third coefficient from the first target coefficient address of the first storage unit, reads the fourth coefficient from the second target coefficient address of the first storage unit, and reads the second rotation factor from the second target rotation factor address; The second operation unit performs operations based on the third coefficient, the fourth coefficient, and the second rotation factor to obtain a third operation result and a fourth operation result. The third operation result is stored in the first target coefficient address in the second storage unit, and the fourth operation result is stored in the second target coefficient address in the second storage unit.

6. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 5, characterized in that, The first arithmetic unit performs calculations based on the first coefficient, the second coefficient, and the first rotation factor to obtain a first calculation result and a second calculation result, including: The first arithmetic unit calculates the multiplication result of the second coefficient and the first rotation factor, and compresses the multiplication result through modulo reduction to obtain a compressed result; Calculate the modulus sum of the first coefficient and the compression result, and use the modulus sum as the first calculation result; Calculate the reduction of the first coefficient and the compression result, and use the reduction result as the second calculation result.

7. The low-resource number-theoretic transform acceleration system for ML-DSA according to claim 6, characterized in that, The calculation of the multiplication result of the second coefficient and the first rotation factor includes: Based on the Karasuba divide-and-conquer algorithm, and using DSP primitives in a field-programmable gate array (FPGA), the multiplication result of the second coefficient and the first twitch factor is calculated.

8. A low-resource number-theoretic transformation acceleration method for ML-DSA, characterized in that, The method is applied to a low-resource number-theoretic transform acceleration system for ML-DSA as described in any one of claims 1-7, and the method includes: The first and second storage units are both used to store the coefficients of the polynomial, and each coefficient address in the first and second storage units stores two coefficients. The third storage unit is used to store the twitch factor, and the fourth storage unit is used to store the inverse twitch factor. The control addressing module is used to determine multiple target coefficient addresses in the first storage unit or the second storage unit based on the current operation mode, the current operation round and the current operation count under the current operation round, and to determine multiple target rotation factor addresses. The first and second operation units in the operation module are used to read a coefficient from each of the target coefficient addresses and a rotation factor from each of the target rotation factor addresses, respectively, perform operations on the read coefficients, and store the operation results in the second storage unit or each of the target coefficient addresses in the first storage unit; Based on the current operation mode, and according to the current operation round and the current operation count under the current operation round, the determination of multiple target coefficient addresses in the first storage unit or the second storage unit, and the determination of multiple target rotation factor addresses, includes: The address of the first target coefficient is determined based on the current operation round and the current operation count; The second target coefficient address is determined based on the first target coefficient address and the current calculation round; Based on the current operation mode, and according to the current operation round and the current operation count, the address of the first target rotation factor is determined; Based on the current operation mode and according to the address of the first target rotation factor, the address of the second target rotation factor is determined; Determining the address of the first target coefficient based on the current operation round and the current operation count includes: If the current operation round is not a first preset high round, and the current operation round is not a first preset low round, then based on the formula... The address of the first target coefficient is calculated, where, The first target coefficient address is defined as cnt, the current operation count is defined as cnt, and the current operation round is defined as stage. If the current operation round is the first preset low round, then the current operation count is used as the address of the first target coefficient; If the current operation round is the first preset high round, then the value of shifting the current operation count one bit to the left is used as the address of the first target coefficient.

9. An electronic device, characterized in that, The electronic device is equipped with a low-resource number-theoretic transformation acceleration system for ML-DSA as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Hardware implementation method and device for number-theory transformation and security chip

    CN118820655A