Parallel implementation method and system of an XMSS signature method resistant to quantum attacks for GPUs
By performing GPU adaptation and multi-level parallel optimization on the XMSS algorithm, the problem that the XMSS signature method cannot use GPU to run in parallel is solved, and efficient XMSS signature operation is achieved.
Patent Information
- Application Number
- CN202211601017.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-13
AI Technical Summary
The existing XMSS signature method cannot effectively utilize the GPU multi-core architecture for parallel operation, resulting in inefficient operation.
By GPU adapting the XMSS algorithm code, adjusting the code structure and compiling options, and determining a multi-level parallel solution based on the GPU architecture, including structural parallelism and Winternitz one-time signature process and the parallelism of L-tree construction, the parallel execution of the XMSS signature method is realized.
It significantly improves the operation efficiency of the XMSS signature method, and can quickly generate key pairs, sign and verify signatures on the GPU, meeting the needs of efficient parallel computing.
Smart Images

Figure CN116015635B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a parallel implementation method and system of an XMSS signature method for resisting quantum attacks for GPUs. Background Art
[0002] Network security is a comprehensive discipline, including network device security, network information security, and network software security. The secure transmission of information is a part of network information security, mainly involving knowledge of cryptography. Signature algorithms can be used to prove the authenticity of the information sent by the sender and are a ubiquitous network security transmission technology. Among them, the XMSS signature method is a hash-based signature algorithm that can resist quantum threats and was standardized in 2020.
[0003] Quantum threats refer to the fact that the Shor algorithm and the Grover algorithm can use quantum computers to crack existing cryptographic systems such as RSA, DSA, and ECDSA. If large-scale quantum computers are developed, it will become feasible to crack existing cryptographic systems. Therefore, it is extremely urgent to propose and improve algorithms that can resist quantum threats. Among them, the XMSS signature method and the LMS algorithm are the first and only signature algorithms against quantum attacks standardized by NIST. Compared with other algorithms against quantum attacks, they have advantages in both security and operating efficiency. Both of these algorithms are stateful algorithms, that is, a private key can only be used for one signature. Although there are some limitations in use, they are recommended for use in the following scenarios: 1. Needing to deploy a digital signature scheme recently; 2. Being deployed for a long time; 3. Once the algorithm is deployed, not transitioning to other algorithms.
[0004] Some standardized parameter implementations of the XMSS signature method still have the problem of too long running time. For example, when running the official code on a Xeon Gold 5218R (released in 2020, 14 nm, 2.1 GHz) processor, the time to generate a signature can be up to 6535 s at most. Therefore, in order to improve the usability of the XMSS signature method, it is urgent to improve the operating efficiency of the algorithm. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a parallel implementation method and system of an XMSS signature method for resisting quantum attacks for GPUs to solve the technical problem that XMSS cannot utilize the multi-core architecture of GPUs to run in parallel in view of the above-mentioned prior art.
[0006] The present invention adopts the following technical solutions:
[0007] A parallel implementation method of an XMSS signature method for resisting quantum attacks for GPUs includes the following steps:
[0008] S1. Adapt the XMSS algorithm code to the GPU, adjust the code structure and compilation options, and obtain the code adapted to the GPU.
[0009] S2. Determine the multi-level parallelism scheme corresponding to the XMSS parameters according to the GPU architecture, including one-level or two-level parallelism schemes.
[0010] S3. Deploy the first-level parallelism scheme based on the code adapted to the GPU obtained in step S1 and the multi-level parallelism scheme determined in step S2, and obtain the code with the first-level parallelism scheme deployed. The first-level parallelism scheme is structural parallelism, and the calculations of each node at the same level are independent and used for parallel computing.
[0011] S4. Deploy the second-level parallelism scheme based on the code with the first-level parallelism scheme deployed obtained in step S3, including the Winternitz one-time signature process and the construction of the L-tree, and complete the parallel implementation of the XMSS signature method against quantum attacks.
[0012] Specifically, step S1 is as follows:
[0013] S101. Adapt the C language code implementation of the XMSS algorithm to the GPU, provide a data transfer interface between the CPU and the GPU, adjust the parameter processing method in the code to adapt to the GPU compiler, and adjust the compilation parameters to adapt to the GPU compiler.
[0014] S102. Transplant the hash functions to the GPU, including the SHA256 and SHAKE256 hash functions; at the same time, optimize the hash functions according to the characteristics of the GPU architecture.
[0015] S103. Use a correctness verification program to verify the correctness of the transplanted code, and at the same time test the signature upper limit of XMSS.
[0016] Furthermore, in step S103, the correctness verification includes generating XMSS key pairs, XMSS signatures, and XMSS signatures.
[0017] Specifically, step S2 is as follows:
[0018] S201. Use a GPU architecture information collection tool to collect information on the number of GPU cores, including the number of multiprocessors, the number of cores on each multiprocessor, and the total number of cores.
[0019] S202. Determine the corresponding number of leaf nodes through the NIST standard parameters.
[0020] S203. Judge the parallelism scheme used for the parameters based on the number of cores obtained in step S201 and the number of XMSS tree leaf nodes obtained in step S202.
[0021] Further, in step S203, if the number of cores of the GPU does not exceed the number of leaf nodes, a first-level parallel scheme, i.e., parallelism in structure, is deployed; otherwise, a two-level parallel scheme is deployed.
[0022] Specifically, step S3 is as follows:
[0023] S301. Deploy the first-level parallel scheme for generating key pairs. The parallel part in the code is the process of constructing the XMSS tree;
[0024] S302. Deploy the first-level parallel scheme for XMSS signature. The processes of generating WOTS+ key pairs and constructing the XMSS tree in the code are executed in parallel, and at the same time, the process of constructing the XMSS tree is executed in parallel. The first-level parallel scheme is not deployed for the XMSS signature verification part.
[0025] Further, in step S301, multi-threaded independent calculation of subtrees is adopted, and the parallel construction of the XMSS tree is completed in combination with the synchronization technology of the GPU.
[0026] Specifically, step S4 is as follows:
[0027] S401. Adjust the generation of WOTS+ key pairs for constructing XMSS tree leaf nodes and the construction of the L tree in the process of generating XMSS key pairs into parallel interfaces; adjust the WOTS+ signature process in the XMSS signature process into a parallel interface, and the generation of WOTS+ key pairs for constructing XMSS tree leaf nodes and the construction of the L tree into parallel interfaces; adjust the acquisition of WOTS+ public keys and the construction of the L tree in the XMSS signature verification process into parallel interfaces;
[0028] S402. Implement the WOTS+ parallel interface. The WOTS+ parallel interface includes three interfaces: WOTS+ key pair generation, WOTS+ signature, and WOTS+ public key acquisition. The calculation hotspot is the calculation process of the W chain. Utilize the characteristic that every N bytes of the W chain calculation is independently calculated to subdivide the calculation tasks and allocate them to each thread for separate execution; the processes of WOTS+ key pair generation and WOTS+ signature both include the process of generating the WOTS+ private key, and the calculation of every N bytes therein is independent and is processed by multiple threads. The parallelism of WOTS+ private key generation and the parallelism of W chain calculation are allocated the same number of threads;
[0029] S403. Implement the parallel interface for the L tree construction process by using the method of separately processing the maximum complete binary tree and the extra leaves, or the method of evenly distributing node calculations.
[0030] Further, in step S403, the method of separately processing the maximum complete binary tree and the extra leaves is specifically as follows:
[0031] For the first 64 leaf nodes, 8 threads respectively process the subtrees of 8 leaf nodes, and then one thread processes the additional three nodes; then, using GPU synchronization technology, the subsequent calculations are performed in a manner of evenly distributing and synchronizing layer by layer until the value of the root node is calculated, at which point the calculation is completed;
[0032] The specific method for evenly distributing node calculations is as follows:
[0033] The calculations for each layer of nodes are evenly distributed, and synchronization is performed after the calculations for each layer are completed.
[0034] In a second aspect, an embodiment of the present invention provides a system for parallel implementation of a GPU-resistant quantum attack XMSS signature method, including:
[0035] An adaptation module that adapts the XMSS algorithm code to the GPU, adjusts the code structure and compilation options, and obtains the adapted GPU code;
[0036] A parallel module that determines the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme;
[0037] A deployment module that deploys the first-level parallel scheme according to the adapted GPU code obtained in step S1 and the multi-level parallel scheme determined in step S2, and obtains the code for deploying the first-level parallel scheme. The first-level parallel scheme is structural parallelism, and the calculations for each node at the same level are independent and used for parallel calculation;
[0038] An implementation module that deploys the second-level parallel scheme according to the code for deploying the first-level parallel scheme obtained in step S3, including the Winternitz one-time signature process and the construction of the L-tree, and completes the parallel implementation of the GPU-resistant quantum attack XMSS signature method.
[0039] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for parallel implementation of the GPU-resistant quantum attack XMSS signature method are implemented.
[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned method for parallel implementation of the GPU-resistant quantum attack XMSS signature method are implemented.
[0041] Compared with the prior art, the present invention has at least the following beneficial effects:
[0042] A parallel implementation method of the XMSS signature method resistant to quantum attacks for GPUs realizes the parallelization of XMSS on GPUs through four steps: adapting to the GPU architecture, determining the parallel scheme, the first-level parallelism, and the second-level parallelism. In this way, an adapted parallel scheme can be deployed according to the specific GPU model and XMSS parameters. At the same time, due to the decision of the parallelism in the first-level structure and the parallelism of the WOTS+ and L-tree constructions in the second level, efficient parallelization from small scale to ultra-large scale can be achieved, providing a parallelization theory for making full use of the GPU architecture.
[0043] Furthermore, an interface for XMSS to adapt to GPU programming is provided, which can ensure efficient data transmission between the CPU and the GPU; efficient compilation options for XMSS on the CPU and GPU are provided to ensure the generation of efficient compiled code for XMSS on the CPU and GPU; a method for verifying correctness is provided, which can ensure accurate result verification when XMSS runs on the GPU.
[0044] Furthermore, by verifying the XMSS key pair, XMSS signature generation, and XMSS signature authentication, it can be ensured that the results of XMSS running on the GPU are the same as those running on the CPU, ensuring correctness.
[0045] Furthermore, by obtaining the number of leaf nodes in the GPU information and NIST parameters and selecting the parallel scheme, it can be ensured that the parallel scheme with the highest parallel efficiency is selected.
[0046] Furthermore, since the efficiency of the first-level parallelism is higher than that of the second-level parallelism, when the number of threads is small, the first-level parallelism should be preferentially deployed, while when the number of threads is large, deploying two-level parallelism can provide a higher degree of parallelism.
[0047] Furthermore, by deploying the first-level parallel scheme for key pair generation, the XMSS tree can be constructed in parallel, significantly reducing the key pair generation time; by parallelizing the WOTS+ key pair generation and constructing the XMSS tree during XMSS signature, more parts that can be executed in parallel are explored.
[0048] Furthermore, the scheme of independent calculation of subtrees by multiple threads can minimize the synchronization overhead, and combined with the GPU synchronization technology to ensure correctness, which can maximize the parallel efficiency of constructing the XMSS tree.
[0049] Furthermore, the second-level parallelism is realized by parallelizing the WOTS+ and L-tree constructions. In WOTS+, parallelization is carried out according to the independence of each N-byte calculation. In the L-tree, two parallel schemes are provided.
[0050] Furthermore, the solution of separately processing the maximum complete binary tree and the redundant leaves can minimize the synchronization overhead because it is implemented by constructing subtrees.
[0051] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0052] In summary, the present invention provides various parallel construction schemes for XMSS tree construction, L tree construction, and WOTS+; in terms of strategy, when performing two-level parallel construction, it provides how to select the parallel level, and when performing XMSS signature, it provides a scheme for parallel execution of WOTS+ and XMSS tree, resulting in a parallel and efficient implementation of the XMSS signature method resistant to quantum attacks on a GPU.
[0053] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0054] Figure 1 It is a schematic diagram of a scheme for constructing an XMSS tree using a two-level parallel framework;
[0055] Figure 2 It is a schematic diagram of two schemes for parallel constructing an L tree, where (a) is a method of separately processing the maximum complete binary tree and the redundant leaves, and (b) is a method of evenly distributing node calculations. Detailed Embodiments
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0058] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0059] It should be further understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates an "or" relationship between the preceding and following related objects.
[0060] It should be understood that although terms such as first, second, and third may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range can also be referred to as the second preset range, and similarly, the second preset range can also be referred to as the first preset range.
[0061] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0062] Structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are merely exemplary, and in practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0063] The present invention provides a parallel implementation method for a GPU-based quantum-resistant XMSS signature method, which adopts a two-layer parallel method, including parallelism in the first layer structure, parallelism in the Winternitz one-time signature (WOTS+) related structure and L-tree construction in the second layer, and combines the decision of a multi-level parallel scheme to obtain a parallel and efficient implementation of a quantum-resistant XMSS signature algorithm on a GPU.
[0064] A parallel implementation method for a GPU-based quantum-resistant XMSS signature method of the present invention includes the following steps:
[0065] S1. Adapt the official XMSS algorithm code to the GPU, adjust the code structure and compilation options so that it can run correctly and efficiently on the GPU;
[0066] S101. Implement the C language code of the official XMSS algorithm on the GPU, provide a data transfer interface between the CPU and the GPU, adjust the parameter processing method in the code to adapt to the GPU compiler, and adjust the compilation parameters to adapt to the GPU compiler;
[0067] S102. Port the hash functions recommended by the National Institute of Standards and Technology (NIST) of the United States to the GPU, including the SHA256 and SHAKE256 hash functions. At the same time, the hash functions need to be optimized according to the characteristics of the GPU architecture;
[0068] S103. Use the official correctness verification program to verify the correctness of the ported code. At the same time, test the signature upper limit of XMSS to ensure that there are no security issues when the signature count exceeds the upper limit.
[0069] S2. Determine the multi-level parallelism scheme corresponding to the XMSS parameters according to the GPU architecture, including one-level or two-level parallel methods;
[0070] S201. Use the GPU architecture information collection tool to collect information about the number of GPU cores, including the number of multiprocessors, the number of cores on each multiprocessor, and the total number of cores;
[0071] S202. Determine the corresponding number of leaf nodes through the parameters of the NIST standard;
[0072] S203. Determine the parallelism scheme used for this parameter based on the number of GPU cores and the number of XMSS tree leaf nodes.
[0073] S3. Deploy the first-level parallelism scheme based on the code obtained in step S1 and the parallelism scheme determined in step S2. The first level refers to the parallelism in structure, such as the parallelism of leaf nodes and branch nodes on the XMSS tree. The calculations of each node at the same level are independent and can be parallelized;
[0074] S301. Deploy the first-level parallelism scheme for generating key pairs. The part of the code that can be parallelized is the process of constructing the XMSS tree. Since the calculations of each node at each layer are independent during the construction process, multi-threaded independent calculation of subtrees is used and combined with the synchronization technology of the GPU to complete the parallel construction of the XMSS tree, as Figure 1 shown;
[0075] S302. Deploy the first-level parallelism scheme for XMSS signatures. The processes of generating WOTS+ key pairs and constructing the XMSS tree in the code can be executed in parallel, and they are independent of each other. At the same time, the process of constructing the XMSS tree can also be executed in parallel. The parallelism scheme is shown in S301. Due to the dependence of algorithm execution, the first-level parallelism scheme cannot be deployed for the XMSS signature verification part.
[0076] S4. Deploy the second-level parallelism scheme to be used according to the code obtained in step S3. The second level refers to the sub-algorithms that can be parallelized in the XMSS algorithm, and these sub-algorithms will be called during the processes of generating XMSS key pairs, XMSS signatures, and XMSS signature verifications, including the processes related to Winternitz one-time signature (WOTS+) and the construction process of the L-tree.
[0077] S401. Adjust the WOTS+ key pair generation and L-tree construction for building the leaf nodes of the XMSS tree during the process of generating XMSS key pairs into parallel interfaces; adjust the WOTS+ signature process during the XMSS signature process into a parallel interface, as well as the WOTS+ key pair generation and L-tree construction for building the leaf nodes of the XMSS tree into parallel interfaces; adjust the WOTS+ public key acquisition and L-tree construction during the XMSS signature verification process into parallel interfaces;
[0078] S402. Implement the parallel interfaces related to WOTS+;
[0079] The parallelism scheme of the parallel interfaces related to WOTS+ is similar, including three interfaces: WOTS+ key pair generation, WOTS+ signature, and WOTS+ public key acquisition.
[0080] Among them, the computational hot spot is the calculation process of the W-chain. Utilizing the characteristic that every N bytes of the W-chain calculation is independently calculated, the calculation tasks are further subdivided and assigned to each thread for separate execution.
[0081] In addition, the processes of WOTS+ key pair generation and WOTS+ signature both include the process of generating the WOTS+ private key. The calculation of every N bytes in this process is independent and is processed by multiple threads. Since the parallelism degree of WOTS+ private key generation is the same as that of the W-chain calculation, the same threads can be assigned, and no synchronization is required between the two processes;
[0082] S403. Implement the parallel interface for the L-tree construction process.
[0083] Please refer to Figure 2 , since the L-tree is not a complete binary tree, the present invention provides two parallel construction methods for the L-tree, aiming to balance the load and synchronization overhead.
[0084] Method 1: The method of separately processing the maximum complete binary tree and the redundant leaves, which is specifically as follows:
[0085] The maximum complete binary tree is constructed by using a parallel construction method for the XMSS tree similar to S301. This method can reduce the synchronization overhead, but may lead to unbalanced computational loads.
[0086] Method 2: The method of evenly distributing node calculations, which is specifically as follows:
[0087] The computing load of each thread is made as balanced as possible. However, since synchronization is required after each layer of computation is completed, the synchronization overhead is relatively large.
[0088] In another embodiment of the present invention, a parallel implementation system for the GPU-based quantum-resistant XMSS signature method is provided. This system can be used to implement the parallel implementation method of the GPU-based quantum-resistant XMSS signature method. Specifically, the parallel implementation system for the GPU-based quantum-resistant XMSS signature method includes an adaptation module, a parallel module, a deployment module, and an implementation module.
[0089] Among them, the adaptation module adapts the XMSS algorithm code to the GPU, adjusts the code structure and compilation options, and obtains the code adapted to the GPU.
[0090] The parallel module determines the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme.
[0091] The deployment module deploys the first-level parallel scheme according to the code adapted to the GPU obtained in step S1 and the multi-level parallel scheme determined in step S2, and obtains the code for deploying the first-level parallel scheme. The first-level parallel scheme is structural parallelism, and the computation of each node at the same level is independent and used for parallel computation.
[0092] The implementation module deploys the second-level parallel scheme according to the code for deploying the first-level parallel scheme obtained in step S3, including the Winternitz one-time signature process and the construction of the L-tree, and completes the parallel implementation of the quantum-resistant XMSS signature method.
[0093] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the parallel implementation method of the GPU anti-quantum attack XMSS signature method, including:
[0094] Perform GPU adaptation on the XMSS algorithm code, adjust the code structure and compilation options to obtain the GPU-adapted code; determine the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme; according to the obtained GPU-adapted code and the determined multi-level parallel scheme, deploy the first-level parallel scheme to obtain the code of the deployed first-level parallel scheme. The first-level parallel scheme is structural parallelism, and the calculations of each node at the same level are independent and are used for parallel calculation; according to the obtained code of the deployed first-level parallel scheme, deploy the second-level parallel scheme, including the Winternitz one-time signature process and the construction of the L-tree, to complete the parallel implementation of the anti-quantum attack XMSS signature method.
[0095] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory.
[0096] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the parallel implementation method of the XMSS signature method for GPU against quantum attacks in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:
[0097] Perform GPU adaptation on the XMSS algorithm code, adjust the code structure and compilation options to obtain the adapted GPU code; determine the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme; according to the obtained adapted GPU code and the determined multi-level parallel scheme, deploy the first-level parallel scheme to obtain the code of the deployed first-level parallel scheme. The first-level parallel scheme is structural parallelism, and the calculations of each node at the same level are independent and are used for parallel computing; according to the obtained code of the deployed first-level parallel scheme, deploy the second-level parallel scheme, including the Winternitz one-time signature process and the construction of the L-tree, to complete the parallel implementation of the XMSS signature method against quantum attacks.
[0098] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0099] Taking the XMSSMT-40 / 4 parameters using the SHA256 hash function as an example
[0100] When running the official code, the time for generating XMSS key pairs, XMSS signatures, and XMSS signature verification on the CPU (Xeon Gold 5218R) is 1.32s, 5.02s, and 2.77ms respectively, and the time for generating XMSS key pairs, XMSS signatures, and XMSS signature verification on the GPU (GTX 3090) is 33.99s, 136.97s, and 91.08ms respectively.
[0101] After deploying the present invention, the time for generating XMSS key pairs, XMSS signatures, and XMSS signature verification on the GPU (GTX 3090) is 3.68ms, 13.72ms, and 3.26ms respectively.
[0102] In summary, a parallel implementation method and system for an XMSS signature method resistant to quantum attacks for GPUs according to the present invention realizes the efficient parallel execution of the XMSS algorithm on GPUs through a two-layer parallel method, including parallelism in the first-layer structure, parallelism in the relevant structure and L-tree construction in the second layer, and combining the decision-making of a multi-level parallel scheme.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0107] The above is only to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU, characterized in that, It includes the following steps: S1. Adapt the XMSS algorithm code to the GPU, adjust the code structure and compilation options to obtain the code adapted to the GPU; S2. Determine the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme. Specifically: S201. Use the GPU architecture information collection tool to collect the information of the number of GPU cores, including the number of multiprocessors, the number of cores on each multiprocessor, and the total number of cores; S202. Determine the corresponding number of leaf nodes through the NIST standard parameters; S203. Judge the parallel scheme used by the parameters based on the number of cores obtained in step S201 and the number of XMSS tree leaf nodes obtained in step S202. If the number of GPU cores does not exceed the number of leaf nodes, deploy the first-level parallel scheme, that is, the parallelism in structure. Otherwise, deploy the two-level parallel scheme; S3. According to the code adapted to the GPU obtained in step S1 and the multi-level parallel scheme determined in step S2, deploy the first-level parallel scheme to obtain the code of the deployed first-level parallel scheme. The first-level parallel scheme is the structural parallelism, and the calculation of each node at the same level is independent and used for parallel calculation; S4. According to the code of the deployed first-level parallel scheme obtained in step S3, deploy the second-level parallel scheme, including the Winternitz one-time signature process and the construction of the L tree, to complete the parallel implementation of the XMSS signature method against quantum attacks.
2. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 1, characterized in that, Step S1 is specifically: S101. Adapt the C language code implementation of the XMSS algorithm to the GPU, provide the data transfer interface between the CPU and the GPU, adjust the parameter processing method in the code to adapt to the GPU compiler, and adjust the compilation parameters to adapt to the GPU compiler; S102. Transplant the hash functions to the GPU, including the SHA256 and SHAKE256 hash functions; at the same time, optimize the hash functions according to the characteristics of the GPU architecture; S103. Use the correctness verification program to verify the correctness of the transplanted code, and at the same time test the signature upper limit of XMSS.
3. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 2, characterized in that, In step S103, the correctness verification includes generating the XMSS key pair, generating the XMSS signature, and verifying the XMSS signature.
4. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 1, characterized in that, Step S3 is specifically: S301. Deploy the first-level parallel scheme for generating the key pair. The parallel part in the code is the process of constructing the XMSS tree; S302. Deploy the first-level parallel scheme for XMSS signature. The process of generating the WOTS+ key pair and constructing the XMSS tree in the code is executed in parallel, and at the same time, the process of constructing the XMSS tree is executed in parallel. The first-level parallel scheme is not deployed for the XMSS signature verification part.
5. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 4, characterized in that, In step S301, use multi-threads to independently calculate the subtrees and combine the synchronization technology of the GPU to complete the parallel construction of the XMSS tree.
6. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 1, characterized in that, Step S4 is specifically: S401. Adjust the WOTS+ key pair generation and L-tree construction for building XMSS tree leaf nodes during the XMSS key pair generation process into a parallel interface; adjust the WOTS+ signature process during the XMSS signature process into a parallel interface, and also adjust the WOTS+ key pair generation and L-tree construction for building XMSS tree leaf nodes into a parallel interface; adjust the WOTS+ public key acquisition and L-tree construction during the XMSS signature verification process into a parallel interface. S402. Implement the WOTS+ parallel interface. The WOTS+ parallel interface includes three interfaces: WOTS+ key pair generation, WOTS+ signature, and WOTS+ public key acquisition. The calculation hotspot is the calculation process of the W-chain. Utilize the characteristic that every N bytes of the W-chain calculation are independently calculated to subdivide the calculation tasks and allocate them to each thread for separate execution. Both the WOTS+ key pair generation and WOTS+ signature processes include the WOTS+ private key generation process. Among them, the calculation of every N bytes is independent and is processed by multiple threads. The same number of threads are allocated for the parallelism of WOTS+ private key generation and the parallelism of W-chain calculation. S403. Implement the parallel interface for the L-tree construction process using the method of separately processing the maximum complete binary tree and the extra leaves, or the method of evenly distributing node calculations.
7. The parallel implementation method of the XMSS signature method resistant to quantum attacks for GPU according to claim 6, characterized in that, In step S403, the method of separately processing the maximum complete binary tree and the extra leaves is specifically as follows: For the first 64 leaf nodes, 8 threads respectively process the subtrees of 8 leaf nodes, and then use one thread to process the additional three nodes; then use the GPU synchronization technology, and the subsequent calculations are carried out in the way of layer-by-layer average distribution and synchronization until the value of the root node is calculated, which means the calculation is completed. The method of evenly distributing node calculations is specifically as follows: The calculations of each layer of nodes are all evenly distributed, and synchronization is performed after the calculations of each layer are completed.
8. A parallel implementation system of the XMSS signature method resistant to quantum attacks for GPU, characterized in that, It includes: An adaptation module that adapts the XMSS algorithm code to the GPU, adjusts the code structure and compilation options, and obtains the code adapted to the GPU. A parallel module that determines the multi-level parallel scheme corresponding to the XMSS parameters according to the GPU architecture, including a one-level or two-level parallel scheme. Specifically: Use the GPU architecture information collection tool to collect information about the number of GPU cores, including the number of multiprocessors, the number of cores on each multiprocessor, and the total number of cores; determine the corresponding number of leaf nodes through the NIST standard parameters; judge the parallel scheme used for the parameters based on the obtained number of cores and the number of XMSS tree leaf nodes. If the number of GPU cores does not exceed the number of leaf nodes, deploy the first-level parallel scheme, that is, the parallelism in structure. Otherwise, deploy the two-level parallel scheme. A deployment module that deploys the first-level parallel scheme according to the code adapted to the GPU obtained by the adaptation module and the multi-level parallel scheme determined by the parallel module, and obtains the code for deploying the first-level parallel scheme. The first-level parallel scheme is the parallelism in structure, and the calculations of each node at the same level are independent and are used for parallel calculation. An implementation module, based on the code of the first-level parallel deployment solution obtained by the deployment module, deploys the second-level parallel solution, including the Winternitz one-time signature process and the construction of the L-tree, to complete the parallel implementation of the quantum-resistant XMSS signature method.
Citation Information
Patent Citations
Tree link signature method for resisting quantum computing attack
CN114338030A
Parallel processing techniques for HASH-based signature algorithms
US20190319802A1