Data encryption and decryption task execution method and device based on post quantum cryptography algorithm, equipment and medium
By using a target register for pipelined operation in the post-quantum cryptography algorithm, the problem of low efficiency in polynomial multiplication is solved, hardware resources are optimized and computing performance is improved, making it suitable for high-speed cryptographic card applications in corporate cryptographic cards and cloud servers.
Patent Information
- Application Number
- CN202511442392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-11
AI Technical Summary
In post-quantum cryptography algorithms, polynomial multiplication operations consume a lot of time and resources, and butterfly operations are complex and computationally inefficient, affecting hardware resource utilization and communication bandwidth.
By using the target register to participate in the pipelined operation of the post-quantum cryptography algorithm, the polynomial coefficients are stored in the target register, and the pipelined operation of the butterfly arithmetic unit is executed in parallel, reducing memory resource consumption and improving computational efficiency.
It reduces hardware resource consumption and improves the computing performance and efficiency of cloud servers, making it suitable for corporate cryptographic card projects and high-speed cryptographic card scenarios on cloud servers.
Smart Images

Figure CN120934758A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, and medium for performing data encryption and decryption tasks based on post-quantum cryptography algorithms. Background Technology
[0002] In post-quantum cryptography's Kyber lattice cryptography schemes, polynomial multiplication is a critical step, consuming the majority of time and resources. Common schemes utilize number theoretic transforms (NTTs) based on butterfly operations for fast polynomial multiplication. However, NTT computation involves multiple butterfly operations, including modular addition, subtraction, multiplication, and reduction, significantly impacting overall hardware efficiency and resource consumption. Secondly, the NTT control logic contains multiple nested loops, resulting in a complex computational structure. Finally, the storage and scheduling of polynomial coefficients are complex and require high communication bandwidth; this affects the effectiveness of post-quantum cryptography algorithms in specific scenarios, such as corporate cryptographic card projects or high-speed cryptographic cards on cloud servers.
[0003] It is evident that balancing computational efficiency and resource consumption is a problem that needs to be solved in the post-quantum cryptography data processing process. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a data encryption / decryption task execution method, apparatus, and medium based on a post-quantum cryptography algorithm. The target register participates in the pipeline operation of the post-quantum cryptography algorithm, which can reduce the use of memory resources. The butterfly operation, performing parallel computation, can improve computational efficiency and performance. The specific solution is as follows:
[0005] Firstly, this application provides a data encryption / decryption task execution method based on a post-quantum cryptography algorithm, applied to a cloud server, including:
[0006] The coefficients to be processed, read from memory, are stored in a target register; the target register includes a first register and a second register based on data bit markers; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm.
[0007] Using a preset number of butterfly operation units, pipelined operations are performed in parallel on the current coefficients of preset data bits read from the first register and the second register respectively, to obtain the corresponding first processing result and second processing result; the pipelined operations are the number theory transformation, inverse number theory transformation and circular convolution algorithm-related processing operations of the post-quantum cryptography algorithm;
[0008] The first processing result and the second processing result are stored in the memory, and relevant data encryption and decryption tasks are performed based on all processing results.
[0009] Optionally, storing the coefficients to be processed read from the memory into the target register includes:
[0010] The polynomial coefficients to be processed, obtained by sampling through a binomial distribution sampler, are stored in memory;
[0011] The current coefficients to be processed are read from the memory and stored sequentially in the target register.
[0012] Optionally, if the current pipelined operation is a processing operation related to the number theory transformation, then the step of performing a pipelined operation in parallel on the current coefficients of a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes:
[0013] The current coefficient of the preset data bit length read from the first register and the second register respectively;
[0014] By using a preset number of butterfly operation units, the current coefficient is multiplied in parallel to obtain the corresponding first data;
[0015] Perform modulo reduction on the first data to obtain the corresponding second data;
[0016] Modular addition and modular subtraction operations are performed on the second data respectively to obtain the corresponding first and second processing results.
[0017] Optionally, if the current pipelined operation is a processing operation related to the inverse number theory transformation, then the step of performing a pipelined operation in parallel on the current coefficients of a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes:
[0018] The current coefficient of the preset data bit length read from the first register and the second register respectively;
[0019] Using a preset number of butterfly operation units, the current coefficient is subjected to modular addition and modular subtraction operations in parallel to obtain the corresponding third and fourth data.
[0020] Multiply the third and fourth data respectively to obtain the corresponding fifth and sixth data;
[0021] Modular reduction operations are performed on the fifth and sixth data respectively to obtain the corresponding first and second processing results.
[0022] Optionally, if the current pipelined operation is a processing operation related to the circular convolution algorithm, then the step of performing a pipelined operation in parallel on the current coefficients of a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes:
[0023] The current coefficient of the preset data bit length read from the first register and the second register respectively;
[0024] By using a preset number of butterfly operation units, the current coefficient is multiplied in parallel to obtain the corresponding seventh data.
[0025] Perform modulo reduction on the seventh data to obtain the corresponding eighth data;
[0026] Modular addition is performed on the data in the eighth data that corresponds to the current register, and the ninth data obtained from the corresponding operation, along with the eighth data, are determined as the first processing result and the second processing result.
[0027] Optionally, the execution of pipeline operations includes:
[0028] During the pipelined operation, the target register, modular addition unit, modular subtraction unit, multiplier, and modular reduction unit are reused.
[0029] Optionally, storing the first processing result and the second processing result in the memory, and performing relevant data encryption / decryption tasks based on all the processing results, includes:
[0030] The first processing result and the second processing result are stored in the corresponding positions in the target register. After obtaining all the processing results corresponding to the coefficient to be processed, all the processing results in the target register are written into the memory, and related data encryption and decryption tasks are performed based on all the processing results.
[0031] Secondly, this application provides a data encryption / decryption task execution device based on a post-quantum cryptography algorithm, applied to a cloud server, comprising:
[0032] A data storage module is used to store the coefficients to be processed read from the memory into a target register; the target register includes a first register and a second register based on data bit markers; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm;
[0033] The data processing module is used to perform pipelined operations on the current coefficients of a preset number of data bits read from the first register and the second register respectively, through a preset number of butterfly operation units, to obtain the corresponding first processing result and second processing result; the pipelined operations are number theory transformations, inverse number theory transformations, and circular convolution algorithm-related processing operations of the post-quantum cryptography algorithm;
[0034] The result storage module is used to store the first processing result and the second processing result into the memory, and to perform relevant data encryption and decryption tasks based on all processing results.
[0035] Thirdly, this application provides an electronic device, comprising:
[0036] Memory, used to store computer programs;
[0037] A processor for executing the computer program to implement the data encryption / decryption task execution method based on the post-quantum cryptography algorithm as described above.
[0038] Fourthly, this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the data encryption / decryption task execution method based on the post-quantum cryptography algorithm described above.
[0039] Therefore, in this application, the cloud server can store the coefficients to be processed read from the memory into a target register. The target register includes a first register and a second register based on data bit width marking. The coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm. Then, through a preset number of butterfly operation units, pipelined operations are performed in parallel on the current coefficients with a preset number of data bits read from the first register and the second register respectively to obtain corresponding first and second processing results. The pipelined operations are processing operations related to number theory transformation, inverse number theory transformation, and circular convolution algorithm of the post-quantum cryptography algorithm. Afterward, the first and second processing results are stored in the memory, and related data encryption and decryption tasks are performed based on all processing results. In this way, this scheme reduces the use of memory resources by participating in the pipelined operations related to the post-quantum cryptography algorithm through the target register. The binomial distributed collector can sample simultaneously without affecting the use of a preset number of butterfly operations to perform data operations in parallel, thereby improving computational efficiency and the computational performance of the cloud server. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0041] Figure 1 This is a flowchart of a data encryption / decryption task execution method based on a post-quantum cryptography algorithm disclosed in this application;
[0042] Figure 2 This application discloses a flowchart of a specific data encryption / decryption task execution method based on a post-quantum cryptography algorithm.
[0043] Figure 3 This is a schematic diagram of an NTT operation structure disclosed in this application;
[0044] Figure 4 This is a schematic diagram of register access for a butterfly arithmetic unit disclosed in this application;
[0045] Figure 5 This application discloses an NTT pipeline operation flowchart;
[0046] Figure 6 This application discloses an INTT pipeline operation flowchart;
[0047] Figure 7 This application discloses a CWM pipeline operation flowchart;
[0048] Figure 8 This is a design block diagram of a resource reuse method disclosed in this application;
[0049] Figure 9 This is a schematic diagram of a data encryption / decryption task execution device based on a post-quantum cryptography algorithm disclosed in this application;
[0050] Figure 10 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Understandably, for Kyber's algorithm, the polynomial has 256 coefficients. Since it's split into two 128th-degree polynomials, a total of seven rounds of butterfly operations are required, making memory (Random Access Memory, RAM) access more complex. Current designs mostly employ multi-RAM interleaving, meaning that eight butterfly operations require 32 blocks of 16x16 (16-bit depth, 16-bit width) RAM0, RAM1, RAM2, RAM3…RAM31 to store intermediate results, and the entire process involves seven rounds of butterfly operations. Although all eight butterfly arithmetic units can operate in parallel, the use of 32 16x16 RAM blocks (RAM0, RAM1, RAM2, RAM3...RAM31) for storing intermediate results results in a large number of RAM blocks, a large total size, and very small block size, leading to RAM fragmentation and failing to leverage the advantages of large storage capacity and small area. Furthermore, the entire NTT operation requires seven rounds of butterfly operations, meaning that these 32 RAM blocks are occupied continuously during these seven rounds, preventing parallel processing of other operations and limiting performance. In contrast, this solution can participate in the pipelined operations related to post-quantum cryptography algorithms through the target register, reducing memory resource usage. The parallel binomial distributed collector can sample simultaneously without affecting the use of a preset number of butterfly operations for parallel data processing, thus improving computational efficiency and enhancing the computing performance of the cloud server.
[0053] See Figure 1 As shown, this invention discloses a data encryption / decryption task execution method based on a post-quantum cryptography algorithm, applied to a cloud server, including:
[0054] Step S11: The coefficients to be processed read from the memory are stored in the target register; the target register includes a first register and a second register based on data bit mark; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm.
[0055] In this embodiment, the polynomial coefficients related to the post-quantum cryptography algorithm are originally stored in memory. When performing the post-quantum cryptography algorithm operation, the coefficients to be processed in memory can be stored in the target register, that is, 256 polynomial coefficients are read from memory into a 256x16 (depth 256, bit width 16) register. In order to optimize the efficiency of subsequent butterfly operation, the polynomial coefficients in the target memory can be marked according to the number of data bits, and marked as the first register reg0 and the second register reg1 according to the number of data bits, which facilitates subsequent operation.
[0056] In one specific embodiment, storing the coefficients to be processed read from the memory into the target register may include: storing the polynomial coefficients to be processed obtained by a binomial distributed sampler into the memory; reading the current coefficients to be processed from the memory and storing them into the target register sequentially. Specifically, in the post-quantum cryptography algorithm, matrix generation is performed using a binomial distributed sampler, the polynomial coefficients are sampling-related coefficients, and all related polynomial coefficients are stored in the memory; then, during algorithm execution, the coefficients can be read sequentially from the memory into the target register, through which the relevant operations of the quantum cryptography algorithm are completed. It is understood that the algorithm operation process, combined with the target register, allows the memory to be freed up for storing sampling-related polynomial coefficients, reducing memory usage.
[0057] Step S12: Using a preset number of butterfly operation units, pipeline operations are performed in parallel on the current coefficients of the preset number of data bits read from the first register and the second register respectively, to obtain the corresponding first processing result and second processing result; the pipeline operation is the number theory transformation, inverse number theory transformation and circular convolution algorithm related processing operations of the post-quantum cryptography algorithm.
[0058] In this embodiment, the polynomial coefficients to be processed can be read into the target register through the above steps. The coefficients stored in the target register are also marked as two copies based on the number of data bits; that is, the first register and the second register, with the corresponding number of data bits being 0-127 and 128-255. Then, the current coefficients with a certain number of data bits can be read from addresses 0 and 128 simultaneously, such as reading 8 half-words of data at a time, and the read current coefficients are sent to a preset number (8) of butterfly operation units for related pipelined operations, that is, 16 data are sent to 8 butterfly operation units at a time, and the corresponding first processing result and second processing result can be obtained. It should be noted that the butterfly operation unit is the core of the number theory transformation, inverse number theory transformation (inverse NNT, INNT) and circular convolution algorithm (coefficient-wise multiplication, CWM) calculation in the post-quantum cryptography algorithm, and the operation is performed through a pipelined structure.
[0059] In a specific embodiment, if the current pipelined operation is a number theory transformation-related processing operation, then the step of performing a pipelined operation in parallel on the current coefficients with a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units may include: reading the current coefficients with a preset number of data bits from the first register and the second register respectively; performing multiplication operations on the current coefficients in parallel through a preset number of butterfly operation units to obtain the corresponding first data; performing modulo reduction and subtraction operations on the first data to obtain the corresponding second data; and performing modulo addition and modulo subtraction operations on the second data respectively to obtain the corresponding first processing result and second processing result. Specifically, the number theory transformation process in the post-quantum cryptography algorithm is implemented through a pipelined structure. The first clock cycle fetches eight 16-bit data points from reg0 and reg1 and inputs them into the multiplier, performing parallel multiplication of the current coefficients. The second clock cycle buffers the multiplied data, i.e., the first data. The third clock cycle inputs the data into the modulo-subtraction function, performing modulo-subtraction on the first data. The fourth clock cycle buffers the modulo-subtracted data, i.e., the second data. The fifth clock cycle performs both modulo-addition and modulo-subtraction simultaneously. The sixth clock cycle buffers the data from the fifth clock cycle, i.e., the first and second processing results. Finally, the seventh clock cycle writes the data back to reg0 and reg1 simultaneously. Furthermore, through this cyclic reading and pipelined operation, the number theory transformation processing operations of the post-quantum cryptography algorithm can be implemented.
[0060] In another specific embodiment, if the current pipelined operation is a processing operation related to the inverse number theory transformation, then the step of performing pipelined operations in parallel on the current coefficients with a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly arithmetic units may include: reading the current coefficients with a preset number of data bits from the first register and the second register respectively; performing modular addition and modular subtraction operations on the current coefficients in parallel through a preset number of butterfly arithmetic units to obtain corresponding third and fourth data; performing multiplication operations on the third and fourth data respectively to obtain corresponding fifth and sixth data; and performing modular reduction and subtraction operations on the fifth and sixth data respectively to obtain corresponding first and second processing results. Specifically, the inverse number theory transformation process in the post-quantum cryptography algorithm is implemented through a pipelined structure. The first clock cycle inputs the data at the corresponding addresses of reg0 and reg1 into the modular addition and subtraction modules, performing modular addition and subtraction operations on the current coefficients. The second clock cycle buffers the data after modular addition and subtraction, i.e., the third and fourth data. The third clock cycle buffers the modular addition data from the second clock cycle, while simultaneously inputting the modular subtraction data from the second clock cycle into the multiplier for multiplication. The fourth clock cycle buffers the data from the third clock cycle, i.e., the fifth and sixth data. The fifth clock cycle performs modular reduction simultaneously. The sixth clock cycle buffers the data from the fifth clock cycle, i.e., the first and second processing results. Then, the seventh clock cycle writes the results back to reg0 and reg1 simultaneously. This loop implements the inverse number theory transformation processing operations of the post-quantum cryptography algorithm.
[0061] In another specific embodiment, if the current pipelined operation is a processing operation related to the circular convolution algorithm, then the step of performing a pipelined operation in parallel on the current coefficients with a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units may include: reading the current coefficients with a preset number of data bits from the first register and the second register respectively; performing a multiplication operation on the current coefficients in parallel through a preset number of butterfly operation units to obtain the corresponding seventh data; performing a modulo reduction operation on the seventh data to obtain the corresponding eighth data; performing a modulo addition operation on the data in the eighth data corresponding to the current register, and determining the ninth data obtained from the corresponding operation and the eighth data as the first processing result and the second processing result. Specifically, the circular convolution algorithm in the post-quantum cryptography algorithm is implemented through a pipelined structure. In the first clock cycle, the data at the corresponding addresses of reg0 and reg1 are input into the multiplier. In the second clock cycle, the data from the first clock cycle is buffered, i.e., the seventh data. In the third clock cycle, modulo reduction is performed to obtain the eighth data. In the fourth clock cycle, the data from the third clock cycle is buffered, and a modulo addition operation is performed on the data in the eighth data corresponding to reg1, yielding the ninth data as the processing result. The data in the eighth data corresponding to reg0 is the corresponding processing result. In the fifth clock cycle, the data from the fourth clock cycle is buffered, i.e., the first and second processing results corresponding to reg0 and reg1 are buffered respectively. Then, in the sixth clock cycle, the data is simultaneously written back to reg0 and reg1. This cycle repeats to implement the related processing operations of the circular convolution algorithm in the post-quantum cryptography algorithm.
[0062] In one specific embodiment, to reduce hardware resource consumption and costs, the target register, modular addition unit, modular subtraction unit, multiplier, and modular reduction unit can be reused in the pipeline of number theory transformation, inverse number theory transformation, and circular convolution algorithms in post-quantum cryptography. Furthermore, the sampling process can also be executed in parallel.
[0063] Step S13: Store the first processing result and the second processing result in the memory, and perform relevant data encryption and decryption tasks based on all processing results.
[0064] In this embodiment, the above steps can be combined with the target register, and the coefficients in the target register can be operated on by the parallel butterfly operation unit to obtain the corresponding operation results, namely the first processing result and the second processing result. It should be noted that the first processing result and the second processing result can be stored in the corresponding position in the target register, that is, the operation result is stored back to the corresponding address of the register after the operation. Then, after obtaining all the processing results corresponding to the coefficients to be processed, all the processing results in the target register are written into the memory, and then the relevant data encryption and decryption tasks can be performed based on all the processing results.
[0065] Therefore, this scheme reduces memory resource usage by using the target register in the pipelined operations related to post-quantum cryptography algorithms. Furthermore, the binomial distributed collectors used for parallel operations can sample simultaneously without affecting the use of a preset number of butterfly operations to perform data operations in parallel. By implementing the butterfly operation units in parallel, and combining the reuse of NTT, INTT, CWM operations and registers, hardware resource consumption can be further reduced, and computational efficiency can be improved, thereby enhancing the computing performance and practicality of cloud servers.
[0066] like Figure 2 As shown, this embodiment discloses a data encryption / decryption task execution method based on a post-quantum cryptography algorithm, applied to a cloud server, including:
[0067] In this embodiment, the control module is used to control the invocation of NTT, INTT, and CWM operators. The post-quantum cryptography algorithm is a computation module with the core of the post-quantum cryptography algorithm. The bus interface can be an AHB bus, which can be directly connected to the AHB bus. The computation module includes a memory, registers (including reg0 and reg1), and a butterfly operation unit.
[0068] Understandably, post-quantum cryptography Kyber is a key encapsulation mechanism based on lattice-hard problems, used to securely negotiate a symmetric key, that is, to generate a shared symmetric key. This symmetric key can be used to encrypt actual data using symmetric encryption algorithms such as AES. It can be applied to specific scenarios such as company cryptographic card projects and high-speed cryptographic cards for cloud servers. The actual data can be the identity information of company employees, dynamic passwords, sensitive resources, etc.
[0069] Furthermore, the NTT transform (number-theoretic transform) of the Kyber algorithm (the Kyber algorithm's operational unit is 12 bits) is designed for parallel operation by eight butterfly arithmetic units. Specifically, the NTT operation needs to ensure its own parallelism while simultaneously ensuring that the sampler operation performed by the binomial distributed sampler can also be performed in parallel. Here, a 256x16 register is used to complete the entire NTT operation. The first step reads the 256 polynomial coefficients of the NTT from the original RAM storage sequentially into the 256x16 register. The second step sends 16 data points from the register to the eight butterfly arithmetic units every clock cycle, and the result is stored back in the register. This process is repeated 16 times, or 16 clock cycles, for 7 rounds, totaling 112 clock cycles. Finally, the result is written back to the original RAM. Figure 3 As shown, during NTT operations, the original data is read from RAM into the Register. After pipelined operations are performed through the register and the butterfly arithmetic unit, the corresponding operation results are written back to RAM. The 256x16 register is divided into reg0 and reg1, each 128x16 in size. In the first round, parameters are sequentially stored in reg0 and reg1. Then, starting from addresses 0 and 128, data in reg0 and reg1 is read simultaneously, with eight halfwords read at a time. These halfwords are then sent to eight butterfly arithmetic units, and the results are stored back to the same address in reg0 and reg1. In the next round, data is again read from reg0 and reg1 simultaneously and sent to the eight butterfly arithmetic units, with the results again stored back in reg0 and reg1. This cycle repeats, using alternating iterative operations on a single register to achieve ordered parameter access in each round. The register access mode of the butterfly arithmetic unit is as follows: Figure 4 As shown; the butterfly arithmetic unit is the core of NTT, INTT, and CWM computations, and is implemented using a pipelined structure. The pipelined operation of NTT is as follows: Figure 5 As shown, the first clock cycle: inputs eight 16-bit data bits from reg0 and reg1 into the multiplier; the second clock cycle: buffers the data after multiplication; the third clock cycle: inputs the data into the modulo-reduction function; the fourth clock cycle: buffers the data after modulo-reduction; the fifth clock cycle: performs modulo-addition and modulo-subtraction simultaneously; the sixth clock cycle: buffers the data from the fifth clock cycle; the seventh clock cycle: writes the data back to reg0 and reg1 simultaneously.
[0070] Accordingly, the INTT transform (inverse number theory transform) of the Kyber algorithm is designed as eight butterfly operations. The INTT operations need to ensure parallelism within themselves, and also guarantee that other operations can be performed in parallel simultaneously. This is achieved through pipelined operations using registers and parallel butterfly arithmetic units; the pipelined operations are as follows... Figure 6As shown, the process is as follows: 1st clock: Input the data at the corresponding addresses of reg0 and reg1 into the modulo addition and subtraction module while performing modulo addition and subtraction; 2nd clock: Buffer the data after modulo addition and subtraction; 3rd clock: Buffer the modulo addition data from the 2nd clock, while inputting the modulo subtraction data from the 2nd clock into the multiplier; 4th clock: Buffer the data from the 3rd clock; 5th clock: Perform modulo reduction and subtraction simultaneously; 6th clock: Buffer the data from the 5th clock; 7th clock: Write back to reg0 and reg1 simultaneously.
[0071] Furthermore, the CWM transform of the Kyber algorithm is designed as eight butterfly operations. The CWM operation needs to ensure its own parallelism, and also guarantee that other operations can be performed in parallel simultaneously. This is achieved through pipelined operations using registers and parallel butterfly operation units; the pipelined operation is as follows... Figure 7 As shown, the first clock cycle: inputs the data at the corresponding addresses of reg0 and reg1 into the multiplier; the second clock cycle: buffers the data from the first clock cycle; the third clock cycle: performs modulo reduction; the fourth clock cycle: buffers the data from the third clock cycle and performs modulo addition on one path; the fifth clock cycle: buffers the data from the fourth clock cycle; the sixth clock cycle: simultaneously writes the data back to reg0 and reg1.
[0072] In a specific embodiment, the NTT, INTT, and CWM operations of the Kyber algorithm can be multiplexed using registers, allowing these three operations to reuse a 256x16 register. The modulo addition, modulo subtraction, multiplier, and modulo reduction / subtraction operation units are also multiplexed, further reducing hardware resource consumption and hardware costs. The register reuse design is as follows: Figure 8 As shown.
[0073] Therefore, this solution reduces memory resource usage by using the target register in the pipelined operations related to post-quantum cryptography algorithms. Furthermore, the parallel binomial distributed collector can sample simultaneously without affecting the use of a preset number of butterfly operations for parallel data processing. By implementing the butterfly operation units in parallel, combined with the reuse of NTT, INTT, and CWM operations and registers, hardware resource consumption can be further reduced. In specific scenarios such as company cryptographic card projects and high-speed cryptographic cards for cloud servers, this solution can improve computational efficiency and enhance the computing performance and practicality of cloud servers.
[0074] like Figure 9 As shown in the figure, this application discloses a data encryption / decryption task execution device based on a post-quantum cryptography algorithm, applied to a cloud server, including:
[0075] Data storage module 11 is used to store the coefficients to be processed read from the memory into a target register; the target register includes a first register and a second register based on data bit markers; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm;
[0076] Data processing module 12 is used to perform pipelined operations on the current coefficients of a preset number of data bits read from the first register and the second register respectively, through a preset number of butterfly operation units, to obtain the corresponding first processing result and second processing result; the pipelined operations are number theory transformations, inverse number theory transformations and circular convolution algorithm-related processing operations of the post-quantum cryptography algorithm;
[0077] The result storage module 13 is used to store the first processing result and the second processing result into the memory, and to perform relevant data encryption and decryption tasks based on all processing results.
[0078] Therefore, this scheme reduces the use of memory resources by participating in the pipeline operations related to post-quantum cryptography algorithms through the target register. Furthermore, the binomial distributed collectors that perform parallel operations can sample simultaneously without affecting the use of a preset number of butterfly operations to perform data operations in parallel, thereby improving computational efficiency and enhancing the computing performance of cloud servers.
[0079] In one specific embodiment, the data storage module 11 may include:
[0080] The first data storage unit is used to store the polynomial coefficients to be processed obtained by the binomial distribution sampler into the memory;
[0081] The second data storage unit is used to read the current coefficients to be processed from the memory and store the coefficients to be processed into the target register in sequence.
[0082] In one specific embodiment, the data processing module 12 may include:
[0083] The current coefficient reading unit is used to read the current coefficient with a preset number of data bits from the first register and the second register respectively;
[0084] The first arithmetic unit is used to perform multiplication operations on the current coefficient in parallel using a preset number of butterfly arithmetic units to obtain the corresponding first data;
[0085] The second processing unit is used to perform modulo reduction and subtraction on the first data to obtain the corresponding second data.
[0086] The third arithmetic unit is used to perform modular addition and modular subtraction operations on the second data respectively to obtain the corresponding first processing result and second processing result.
[0087] In another specific embodiment, the data processing module 12 may include:
[0088] The current coefficient reading unit is used to read the current coefficient with a preset number of data bits from the first register and the second register respectively;
[0089] The fourth arithmetic unit is used to perform modular addition and modular subtraction operations on the current coefficient in parallel using a preset number of butterfly arithmetic units to obtain the corresponding third and fourth data.
[0090] The fifth arithmetic unit is used to perform multiplication operations on the third and fourth data respectively to obtain the corresponding fifth and sixth data;
[0091] The sixth arithmetic unit is used to perform modulo reduction and subtraction operations on the fifth data and the sixth data respectively to obtain the corresponding first processing result and second processing result.
[0092] In yet another specific embodiment, the data processing module 12 may include:
[0093] The current coefficient reading unit is used to read the current coefficient with a preset number of data bits from the first register and the second register respectively;
[0094] By using a preset number of butterfly operation units, the current coefficient is multiplied in parallel to obtain the corresponding seventh data.
[0095] The seventh arithmetic unit is used to perform modulo reduction on the seventh data to obtain the corresponding eighth data;
[0096] The eighth arithmetic unit is used to perform a modulo addition operation on the data in the eighth data that corresponds to the current register, and to determine the ninth data obtained by the corresponding operation and the eighth data as the first processing result and the second processing result.
[0097] In one specific embodiment, the data processing module 12 may include:
[0098] The data processing unit is used to multiplex the target register, the modular addition unit, the modular subtraction unit, the multiplier, and the modular reduction unit during the execution of pipelined operations.
[0099] In one specific embodiment, the result storage module 13 may include:
[0100] The result storage unit is used to store the first processing result and the second processing result in the corresponding positions in the target register, and after obtaining all the processing results corresponding to the coefficient to be processed, write all the processing results in the target register into the memory, and perform relevant data encryption and decryption tasks based on all the processing results.
[0101] Furthermore, embodiments of this application also disclose an electronic device, Figure 10 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0102] Figure 10 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the data encryption / decryption task execution method based on the post-quantum cryptography algorithm disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0103] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0104] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0105] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the data encryption / decryption task based on the post-quantum cryptography algorithm executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0106] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned data encryption / decryption task execution method based on a post-quantum cryptography algorithm. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0107] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0108] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0109] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0110] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0111] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for executing data encryption and decryption tasks based on post-quantum cryptography algorithms, characterized in that, Applied to cloud servers, including: The coefficients to be processed, read from memory, are stored in a target register; the target register includes a first register and a second register based on data bit markers; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm. Using a preset number of butterfly operation units, pipelined operations are performed in parallel on the current coefficients of preset data bits read from the first register and the second register respectively, to obtain the corresponding first processing result and second processing result; the pipelined operations are the number theory transformation, inverse number theory transformation and circular convolution algorithm-related processing operations of the post-quantum cryptography algorithm; The first processing result and the second processing result are stored in the memory, and relevant data encryption and decryption tasks are performed based on all processing results.
2. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to claim 1, characterized in that, The step of storing the coefficients to be processed, read from memory, into the target register includes: The polynomial coefficients to be processed, obtained by sampling through a binomial distribution sampler, are stored in memory; The current coefficients to be processed are read from the memory and stored sequentially in the target register.
3. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to claim 1, characterized in that, If the current pipelined operation is a number theory transformation-related processing operation, then the pipelined operation performed in parallel on the current coefficients of a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes: The current coefficient of the preset data bit length read from the first register and the second register respectively; By using a preset number of butterfly operation units, the current coefficient is multiplied in parallel to obtain the corresponding first data; Perform modulo reduction on the first data to obtain the corresponding second data; Modular addition and modular subtraction operations are performed on the second data respectively to obtain the corresponding first and second processing results.
4. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to claim 1, characterized in that, If the current pipelined operation is a processing operation related to the inverse number theory transformation, then the pipelined operation performed in parallel on the current coefficients of a preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes: The current coefficient of the preset data bit length read from the first register and the second register respectively; Using a preset number of butterfly operation units, the current coefficient is subjected to modular addition and modular subtraction operations in parallel to obtain the corresponding third and fourth data. Multiply the third and fourth data respectively to obtain the corresponding fifth and sixth data; Modular reduction operations are performed on the fifth and sixth data respectively to obtain the corresponding first and second processing results.
5. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to claim 1, characterized in that, If the current pipelined operation is a processing operation related to the circular convolution algorithm, then the pipelined operation performed in parallel on the current coefficients of the preset number of data bits read from the first register and the second register respectively through a preset number of butterfly operation units includes: The current coefficient of the preset data bit length read from the first register and the second register respectively; By using a preset number of butterfly operation units, the current coefficient is multiplied in parallel to obtain the corresponding seventh data. Perform modulo reduction on the seventh data to obtain the corresponding eighth data; Modular addition is performed on the data in the eighth data that corresponds to the current register, and the ninth data obtained from the corresponding operation, along with the eighth data, are determined as the first processing result and the second processing result.
6. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to claim 1, characterized in that, The execution of pipeline operations includes: During the pipelined operation, the target register, modular addition unit, modular subtraction unit, multiplier, and modular reduction unit are reused.
7. The data encryption / decryption task execution method based on post-quantum cryptography algorithm according to any one of claims 1 to 6, characterized in that, The step of storing the first processing result and the second processing result in the memory, and performing relevant data encryption and decryption tasks based on all processing results, includes: The first processing result and the second processing result are stored in the corresponding positions in the target register. After obtaining all the processing results corresponding to the coefficient to be processed, all the processing results in the target register are written into the memory, and related data encryption and decryption tasks are performed based on all the processing results.
8. A data encryption / decryption task execution device based on a post-quantum cryptography algorithm, characterized in that, Applied to cloud servers, including: A data storage module is used to store the coefficients to be processed read from the memory into a target register; the target register includes a first register and a second register based on data bit markers; the coefficients to be processed are polynomial coefficients related to the post-quantum cryptography algorithm; The data processing module is used to perform pipelined operations on the current coefficients of a preset number of data bits read from the first register and the second register respectively, through a preset number of butterfly operation units, to obtain the corresponding first processing result and second processing result; the pipelined operations are number theory transformations, inverse number theory transformations, and circular convolution algorithm-related processing operations of the post-quantum cryptography algorithm; The result storage module is used to store the first processing result and the second processing result into the memory, and to perform relevant data encryption and decryption tasks based on all processing results.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the data encryption / decryption task execution method based on the post-quantum cryptography algorithm as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the data encryption / decryption task execution method based on any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method in polynomial multiplier, polynomial multiplier and processor
CN114371829A
Number-theory transformation processing device based on post-quantum encryption algorithm
CN119341735A
Hardware architecture for memory organization for fully homomorphic encryption
US20220385447A1