A parallel scheduling implementation method, system, medium and device based on SHA-256 algorithm

By designing a parallel scheduling implementation method on the Shenwei SW26010 processor, optimizing DMA transmission and instruction operation, the problem of inefficient operation of SHA-256 algorithm on the processor is solved, and an efficient network secure transmission environment is achieved, and the operation efficiency is improved by 2.77 times.

CN115934092BActive Publication Date: 2025-05-13XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210907513.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-05-13
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

The lack of efficient long message SHA-256 algorithm implementation on the Shenwei SW26010 processor has led to difficulties in building a safe and efficient network security transmission environment, which seriously hinders network security.

Method used

By designing customized instruction optimization methods and instruction set parallel strategies, combined with efficient parallel scheduling algorithms, DMA transmission, parallel scheduling, instruction operation and assembly code are optimized to realize parallel scheduling of SHA-256 algorithm.

Benefits of technology

Compared with the original SHA-256 algorithm implementation that optimizes memory access efficiency, the operating efficiency has been improved by 2.77 times, effectively improving the operating efficiency of the SHA-256 algorithm on the Shenwei SW26010 processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115934092B_ABST
    Figure CN115934092B_ABST
Patent Text Reader

Abstract

The present invention discloses a parallel scheduling implementation method, system, medium and device based on SHA-256 algorithm. For the Shenwei 26010 processor, the parallel scheduling implementation of SHA-256 algorithm is designed. Through all-round performance optimization, including compilation option optimization, DMA transmission optimization, parallel scheduling optimization, instruction reduction optimization, and assembly-level loop unrolling and dual-emission technology optimization, an efficient parallel scheduling implementation of SHA-256 algorithm is obtained. Compared with the original SHA-256 algorithm implementation with optimized over-memory access efficiency, the running time for hash calculation of 1MB message is reduced from 45.01s to 16.25s, and the running efficiency is improved by 2.77 times.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a parallel scheduling implementation method, system, medium and equipment based on a SHA-256 algorithm. Background Art

[0002] Network security is a comprehensive subject, including network equipment security, network information security and network software security. The secure transmission of information is a part of network information security, which mainly involves the knowledge of cryptography. Cryptographic hash algorithms are indispensable in the process of secure transmission, and are used to obtain the summary of messages securely and efficiently. Among them, Secure Hash Algorithms (SHAs) are the most widely used cryptographic hash functions.

[0003] Secure hash functions are often used in other cryptographic algorithms, such as digital signature algorithms, keyed hash message authentication codes, and random number generation. Secure hash functions ensure security by ensuring two types of uncomputability: one is that it is uncomputable to find the message corresponding to the message digest, and the other is that it is uncomputable to find two different messages that produce the same message digest. At the same time, secure hash functions are widely used in quantum computing-resistant cryptography, which can be used to resist attacks from quantum computers, but these algorithms usually require more computation and have higher performance requirements. Therefore, with the development of quantum computers and the greater demand for quantum computing resistance, more efficient secure hash functions have become more necessary.

[0004] There are some sensitive data on supercomputers that need to be protected. As one of the fastest supercomputers in the world, Sunway TaihuLight has an increasing demand for data security. Recently, symmetric encryption algorithms have been deployed on Sunway TaihuLight, such as AES and CHACHA20. However, to achieve secure network transmission, security negotiation is also required before data is securely transmitted, which requires the participation of secure hash functions. Currently, the most secure and efficient SHA algorithm is SHA-256, so it is ideal to deploy this algorithm on Sunway TaihuLight. After deployment, further secure transmission construction can be carried out on Sunway TaihuLight, thus ultimately realizing a complete secure data transmission system.

[0005] There is a lack of efficient implementation of the long-message SHA-256 algorithm on the Shenwei 26010 chip, and a lack of technical solutions to adapt to the chip architecture. This is not conducive to the construction of a safe and efficient network security transmission environment for domestic supercomputers, and seriously hinders network security. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a parallel scheduling implementation method, system, medium and device based on the SHA-256 algorithm in view of the deficiencies in the above-mentioned prior art. By designing a customized instruction optimization method and instruction set parallel strategy for the Shenwei SW26010 processor, and combining it with an efficient parallel scheduling algorithm design, the operating efficiency is improved by 2.77 times compared to the original SHA-256 algorithm implementation that optimizes the access efficiency. The problem of low operating efficiency of the SHA-256 algorithm on the Shenwei SW26010 processor can be solved, and the efficiency of other cryptographic algorithms that need to use the SHA-256 algorithm can also be effectively improved.

[0007] The present invention adopts the following technical solutions:

[0008] A parallel scheduling implementation method based on the SHA-256 algorithm comprises the following steps:

[0009] S1. Port the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively;

[0010] S2, using the optimized DMA transmission mode to process the slave core code and the master core code obtained in step S1, to obtain the master core code and the slave core code after DMA transmission optimization;

[0011] S3, using the master core code obtained in step S2 to provide a data interface to be transmitted, and performing vector parallelization on the scheduling part of the SHA-256 algorithm according to the slave core code obtained in step S2 to obtain the final parallel scheduling result, and then obtaining the slave core code after DMA transmission optimization and parallel scheduling optimization through the scalar register;

[0012] S4, according to the slave core code after DMA transmission optimization and parallel scheduling optimization obtained in step S3, the instructions of Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm are equivalently replaced to obtain the slave core code after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization;

[0013] S5. Perform loop expansion on the slave core code obtained in step S4 after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order, and limit the number of intermediate registers used to achieve parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0014] Specifically, step S1 is as follows:

[0015] S101, transplant the C code of SHA-256 algorithm in OpenSSL to the Shenwei SW26010 processor, provide a function interface, input is the message length and the pointer of the message, and output is the message digest;

[0016] S102. Use sw5gcc to compile the main core code of the SHA-256 algorithm with the compile options -O2 and -msimd.

[0017] S103. Use sw5gcc compiler to compile the slave core code with the compile option -O2-msimd-funroll-all-loops-mslave.

[0018] S104, using the sw5gcc compiler to link the main core intermediate file and the slave core intermediate file generated by the compilation, linking with the file calling the function interface, generating an executable file, and adding -sw3run swrun-5a running parameters to run the executable file generated by the sw5gcc compilation.

[0019] Specifically, step S2 is as follows:

[0020] S201, in the Shenwei SW26010 processor, a static space is allocated from the local memory of the core for DMA transmission;

[0021] S202. The data transmission amount used by the DMA transmission interface is adjusted, and the maximum number of blocks is set to M. For messages with a block number less than M, only one DMA transmission is performed. For messages with a block number greater than M, multiple DMA transmissions are performed to obtain a slave core code optimized for DMA transmission.

[0022] Specifically, step S3 is as follows:

[0023] S301, judging the parallel dispatch interface to be called according to the number of blocks of the message, for messages with a length greater than 8 blocks, using 8 dispatch interfaces to process them in sequence, when the number of remaining blocks is less than 8 blocks, calling interfaces for processing different numbers of blocks according to the number of remaining blocks; for messages with a length less than or equal to 8 blocks, directly calling the corresponding interface;

[0024] S302, load the S blocks of the message into the vector register, S≤8, convert the little-endian data into big-endian data, and obtain the final message scheduling result; unload the scheduling result from the vector register, complete the message compression calculation serially, and obtain the slave core code that has been optimized for DMA transmission and parallel scheduling.

[0025] Specifically, in step S4, the Ch operation defined in the SHA-256 standard is transformed into an equivalent mathematical operation; the Maj operation defined in the SHA-256 standard is transformed into an equivalent mathematical operation; the circular shift operation defined in the SHA-256 standard is transformed into an equivalent hardware instruction, and the instruction is called using the simd interface or by calling the assembly language interface to obtain the slave core code that has been optimized for DMA transmission, parallel scheduling, and instruction operation.

[0026] Furthermore, the data v is circularly shifted to the left by n bits to realize the equivalent transformation of the circular shift operation defined by the SHA-256 standard and the hardware instruction, and the conversion is as follows:

[0027] (((v)<<(n))|((v)>>(32-(n))))→vrotlw(v,n)

[0028] Among them, the '<<' and '>>' operations represent logical left shift and logical right shift operations, and vrotlw represents the circular shift instruction provided by Shenwei 26010.

[0029] Specifically, step S5 is as follows:

[0030] S501, loop expansion is performed on the slave core code in the assembly code, so that the parallel scheduling and message compression parts are loop expanded;

[0031] S502: After the parallel scheduling and message compression part loops of step S501 are expanded, memory access instructions are staggered to adjust the instruction sequence;

[0032] S503. Use 8 registers, 3 of which are registers for storing variables and 5 are intermediate registers for message scheduling. Store instructions in the registers to implement parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0033] In a second aspect, an embodiment of the present invention provides a parallel scheduling implementation system based on the SHA-256 algorithm, characterized in that it includes a transplantation module, a replacement module, a parallel module, an operation module and an implementation module;

[0034] The transplantation module is used to transplant the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively;

[0035] A replacement module is used to process the slave core and master core codes obtained by the transplantation module using the optimized DMA transmission mode to obtain the slave core code optimized by DMA transmission;

[0036] The parallel module is used to provide a data interface to be transmitted by using the master core code obtained by the replacement module, and to perform vector parallelization on the scheduling part of the SHA-256 algorithm according to the slave core code obtained by the replacement module to obtain the final parallel scheduling result, and then obtain the slave core code optimized by DMA transmission and parallel scheduling through the scalar register;

[0037] An operation module is used to perform equivalent replacement of the instructions of the Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm according to the slave core code optimized by DMA transmission and parallel scheduling obtained by the parallel module, so as to obtain the slave core code optimized by DMA transmission, parallel scheduling and instruction operation;

[0038] The implementation module is used to loop expand the slave core code obtained by the operation module after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order, and limit the number of intermediate registers used to achieve parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0039] In a third aspect, a computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned parallel scheduling implementation method based on the SHA-256 algorithm when executing the computer program.

[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned parallel scheduling implementation method based on the SHA-256 algorithm.

[0041] Compared with the prior art, the present invention has at least the following beneficial effects:

[0042] The invention discloses a parallel scheduling implementation method based on SHA-256 algorithm, which obtains an efficient implementation of parallel scheduling of SHA-256 algorithm through all-round performance optimization, including compilation option optimization, DMA transmission optimization, parallel scheduling optimization, instruction reduction optimization, and assembly-level loop unrolling and dual-emission technology optimization.

[0043] Furthermore, a compilation option that is most beneficial to the efficient operation of the SHA-256 algorithm is provided. The sw5gcc compiler of Shenwei is used, combined with the optimization level option and the loop unrolling control option to achieve efficient assembly code generation in the compilation. In order to run the executable file compiled by sw5gcc, the running command needs to be modified and the -sw3run swrun-5a running parameter needs to be added.

[0044] Furthermore, a memory access solution is provided that is beneficial to improving the memory access performance of the SHA-256 algorithm, thereby greatly improving the execution efficiency of the SHA-256 algorithm.

[0045] Furthermore, taking advantage of the fact that the scheduling part of the SHA-256 algorithm can be parallelized, a parallel scheduling implementation scheme was designed to significantly reduce the time occupied by the scheduling part and provide a parallel optimization method for messages of different lengths, so that the program can increase the degree of parallelism as much as possible, thereby efficiently utilizing parallel scheduling technology.

[0046] Furthermore, a special instruction optimization solution for the SHA-256 algorithm on Shenwei 26010 is provided, including Ch operation, Maj operation, big-endian conversion operation and circular shift operation, so as to further improve the execution efficiency of the SHA-256 algorithm on Shenwei 260101.

[0047] Furthermore, by using the circular shift instructions provided by the sw5gcc compiler to replace the circular shift instructions implemented by software, the number of cycles used by the circular shift operation can be reduced from four to one, thereby improving the operating efficiency.

[0048] Furthermore, an optimization solution for the assembly-level SHA-256 algorithm on the Shenwei 26010 is provided, including loop unrolling, improving dual-transmission efficiency, and reducing the number of registers, to further improve the algorithm execution efficiency.

[0049] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0050] In summary, the present invention provides a parallel scheduling method for the SHA-256 algorithm that fully utilizes the computing performance of the domestic Shenwei-26010 by optimizing DMA transmission, parallel scheduling, instruction operation and assembly code.

[0051] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a schematic diagram of parallel scheduling of the present invention;

[0053] Figure 2 is a schematic diagram of a computer device provided by an embodiment of the present invention;

[0054] Figure 3 is a flow chart of the method of the present invention;

[0055] Figure 4 Schematic diagram of the system of the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.

[0058] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0059] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship.

[0060] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0061] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0062] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0063] The present invention provides a parallel scheduling implementation method based on the SHA-256 algorithm. For the domestic Shenwei 26010 processor, the parallel scheduling implementation of the SHA-256 algorithm is designed. Through all-round performance optimization, including compilation option optimization, DMA transmission optimization, parallel scheduling optimization, instruction reduction optimization, and assembly-level loop unrolling and dual-emission technology optimization, an efficient parallel scheduling implementation of the SHA-256 algorithm is obtained. Compared with the original SHA-256 algorithm implementation that optimizes the over-access efficiency, the operating efficiency is improved by 2.77 times.

[0064] See also Figure 3 The present invention provides a parallel scheduling implementation method based on the SHA-256 algorithm, comprising the following steps:

[0065] S1. Port the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively;

[0066] S101, transplant the C code of SHA-256 algorithm in OpenSSL to the Shenwei SW26010 processor, provide a function interface, input is the message length and the pointer of the message, and output is the message digest;

[0067] S102. Use sw5gcc to compile the main core code of the SHA-256 algorithm with the compile options -O2 and -msimd.

[0068] Among them, the -O2 option is the compilation optimization level of the SHA-256 algorithm, which can provide higher performance than other compilation levels, such as -O1 and -O3; the -msimd option provides the main core with the ability to perform vector calculations;

[0069] S103. Use sw5gcc compiler to compile the slave core code with the compile option -O2-msimd-funroll-all-loops-mslave.

[0070] Among them, the -O2 option can provide the highest compilation performance of the SHA-256 algorithm, the -funroll-all-loops option indicates the use of aggressive loop unrolling optimization, and the -mslave option indicates that this is a slave core compilation;

[0071] S104 , using the sw5gcc compiler to link the compiled master core intermediate file and the slave core intermediate file, and linking them with the file for calling the function interface to generate an executable file.

[0072] In order to run the executable file compiled by sw5gcc, you need to modify the run command and add the -sw3runswrun-5a run parameter.

[0073] S2, using the optimized DMA transmission mode to process the slave core code and the master core code obtained in step S1, to obtain the master core code and the slave core code after DMA transmission optimization;

[0074] The present invention uses DMA transmission to replace the direct access main memory method, and optimizes the single data amount of DMA transmission according to the message length, and transmits short messages once and transmits long messages in segments, thereby improving DMA transmission efficiency and improving the overall memory access bandwidth.

[0075] S201, opening up a new static space in the slave core local memory for DMA transmission, and using the DMA interface to replace the slave core interface for directly accessing the main memory;

[0076] Specifically, when the slave core gets a message, it uses the athread_get interface to access data from the main memory to the slave core's local storage. When the slave core writes the message digest back, it uses the athread_put interface to write the data in the slave core's local storage back to the main memory.

[0077] S202. For long messages, multiple blocks of the message (block size is 64 bytes) can be loaded simultaneously in one DMA transfer. Since the size of the slave core local memory is limited, the maximum number of blocks is set to M (to ensure parallel scheduling, M is a multiple of 8). For messages with a block number less than M, only one DMA transfer is performed. For messages with a block number greater than M, multiple DMA transfers are performed.

[0078] Experimental verification shows that when M is 64, the DMA transfer overhead is almost negligible, so there is no need to increase the value of M.

[0079] S3, using the master core code obtained in step S2 to provide a data interface to be transmitted, and according to the slave core code obtained in step S2, performing vector-level parallelization on the scheduling part of the SHA-256 algorithm to obtain the final parallel scheduling result, and then obtaining the slave core code after DMA transmission optimization and parallel scheduling optimization through the scalar register;

[0080] To perform vector-level parallelism, multiple blocks of a message are loaded into the vector register at the same time, and multiple blocks of little-endian data are stored in the vector register for big-endian conversion. After the big-endian conversion and scheduling calculation are completed, the scheduling data is unloaded from the vector register.

[0081] S301, determining the parallel scheduling interface to be called according to the number of message blocks;

[0082] Because the vector register length of the Shenwei 26010 processor is 256 bits, and the operation object of the SHA-256 scheduling part is 32-bit data, up to 8 schedules can be processed simultaneously using the vector registers. However, the number of message blocks is not always a multiple of 8.

[0083] For messages with a length greater than 8 blocks, first use the interface for processing 8 schedules to process them in sequence until the number of remaining blocks is less than 8 blocks. Then, based on the number of remaining blocks, call the interface for processing different numbers of blocks.

[0084] For messages whose length is no more than 8 blocks, just call the corresponding interface.

[0085] S302, performing parallel scheduling design;

[0086] See also Figure 1 First, load the S (S≤8) blocks of the message into the vector register, convert the little-endian data into big-endian data, and then complete the calculation without dependencies to obtain the final message scheduling result; then unload the scheduling result from the vector register and complete the message compression calculation serially. Scheduling describes the process of loading message blocks. The scheduling processes of multiple messages are independent of each other, while the message compression process is dependent and difficult to accelerate in parallel. Therefore, only message scheduling is accelerated in parallel.

[0087] S4, according to the slave core code after DMA transmission optimization and parallel scheduling optimization obtained in step S3, the instructions of Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm are equivalently replaced to obtain the slave core code after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization;

[0088] S401, perform equivalent transformation on the Ch operation defined by the SHA-256 standard;

[0089] The number of instruction cycles is reduced from 4 cycles to 3 cycles through equivalent transformation. The equivalent transformation is as follows:

[0090]

[0091] Where x, y, z are the first, second, and third operands in the Ch operation.

[0092] S402, performing an equivalent transformation on the Maj operation defined in the SHA-256 standard;

[0093] The equivalent transformation reduces the number of instruction cycles from 5 cycles to 4 cycles. The equivalent transformation is as follows:

[0094]

[0095] S403, perform equivalent transformation on the circular shift operation defined in the SHA-256 standard, so as to reduce the number of instruction cycles from 4 cycles to 1 cycle.

[0096] For example, if we rotate data v to the left by n bits, the conversion is as follows:

[0097] (((v)<<(n))|((v)>>(32-(n))))→vrotlw(v,n)

[0098] The '<<' and '>>' operations represent logical left shift and logical right shift operations, and vrotlw represents the circular shift instruction provided by Shenwei 26010.

[0099] Instructions can be called in two ways, one is to use the simd interface to call, the other is to call by calling the assembly language interface, the effects of the two are the same.

[0100] S5. Loop-expand the slave core code obtained in step S4 after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order to make full use of the dual-issue technology, limit the number of registers used, and finally realize the parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0101] By limiting the number of registers used to further improve efficiency, compared to the original SHA-256 algorithm implementation that optimizes memory access efficiency, the running time for hash calculation of 1MB message is reduced from 45.01s to 16.25s, and the running efficiency is improved by 2.77 times.

[0102] S501, perform manual cycle expansion;

[0103] In order to avoid the overhead of loop index, loop unrolling is manually performed in the assembly code, and both the parallel scheduling and message compression parts can be loop unrolled.

[0104] S502, adjusting the instruction sequence;

[0105] In order to make full use of the dual-issue technology, the order of instructions is adjusted to execute the code more efficiently, and finally the performance optimization of the algorithm is completed; the key point of the instruction adjustment scheme is to stagger the memory access instructions, because memory access instructions are usually slower and are completed in one pipeline.

[0106] S503: Reduce the number of registers.

[0107] By reducing the use of temporary registers when designing assembly code, the algorithm execution time can be further reduced.

[0108] In this embodiment, the number of intermediate registers used in message scheduling is 5, and the number of intermediate registers used in message compression is also 5.

[0109] See also Figure 4 In another embodiment of the present invention, a parallel scheduling implementation system based on the SHA-256 algorithm is provided, and the system can be used to implement the above-mentioned parallel scheduling implementation method based on the SHA-256 algorithm. Specifically, the parallel scheduling implementation system based on the SHA-256 algorithm includes a transplantation module, a replacement module, a parallel module, an operation module and an implementation module.

[0110] Among them, the transplantation module is used to transplant the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively;

[0111] A replacement module is used to process the slave core and master core codes obtained by the transplantation module using the optimized DMA transmission mode to obtain the slave core code optimized by DMA transmission;

[0112] A parallel module is used to provide a data interface to be transmitted by using the master core code obtained by the replacement module, perform vector parallelization on the scheduling part of the SHA-256 algorithm according to the slave core code obtained by the replacement module, store multiple blocks of little-endian data in vector registers for big-endian conversion, obtain the final parallel scheduling result, and then obtain the slave core code optimized by DMA transmission and parallel scheduling through the scalar register;

[0113] An operation module is used to perform equivalent replacement of the instructions of the Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm according to the slave core code optimized by DMA transmission and parallel scheduling obtained by the parallel module, so as to obtain the slave core code optimized by DMA transmission, parallel scheduling and instruction operation;

[0114] The implementation module is used to loop expand the slave core code obtained by the operation module after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order, and limit the number of intermediate registers used to achieve parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0115] In another embodiment of the present invention, a terminal device is provided, the terminal device includes a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of the parallel scheduling implementation method based on the SHA-256 algorithm, including:

[0116] The C code of the SHA-256 algorithm is transplanted to the Shenwei SW26010 processor, and the sw5gcc compiler of the Shenwei SW26010 processor is used to compile the codes of the slave core and the master core respectively; the slave core code and the master core code obtained in step S1 are processed using the optimized DMA transmission method to obtain the master core code and the slave core code after DMA transmission optimization; the master core code provides the data interface to be transmitted, and the scheduling part of the SHA-256 algorithm is vectorized and parallelized according to the slave core code, and multiple blocks of little-endian data are stored in the vector register for big-endian conversion to obtain the final parallel scheduling result, and then the master core code and the slave core code are optimized through the optimized DMA transmission method; the master core code provides the data interface to be transmitted, and the scheduling part of the SHA-256 algorithm is vectorized and parallelized according to the slave core code, and multiple blocks of little-endian data are stored in the vector register for big-endian conversion to obtain the final parallel scheduling result. Through the scalar register, the slave core code that has been optimized for DMA transmission and parallel scheduling is obtained; based on the slave core code that has been optimized for DMA transmission and parallel scheduling, the instructions of the Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm are equivalently replaced to obtain the slave core code that has been optimized for DMA transmission, parallel scheduling and instruction operation; the slave core code that has been optimized for DMA transmission, parallel scheduling and instruction operation is loop expanded, the assembly code order is adjusted, and the number of intermediate registers used is limited to realize the parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0117] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory.

[0118] The processor may load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the parallel scheduling implementation method based on the SHA-256 algorithm in the above embodiment; the processor may load and execute the following steps of one or more instructions in the computer-readable storage medium:

[0119] The C code of the SHA-256 algorithm is transplanted to the Shenwei SW26010 processor, and the sw5gcc compiler of the Shenwei SW26010 processor is used to compile the codes of the slave core and the master core respectively; the slave core code and the master core code obtained in step S1 are processed using the optimized DMA transmission method to obtain the master core code and the slave core code after DMA transmission optimization; the master core code provides the data interface to be transmitted, and the scheduling part of the SHA-256 algorithm is vectorized and parallelized according to the slave core code, and multiple blocks of little-endian data are stored in the vector register for big-endian conversion to obtain the final parallel scheduling result, and then the master core code and the slave core code are optimized through the optimized DMA transmission method; the master core code provides the data interface to be transmitted, and the scheduling part of the SHA-256 algorithm is vectorized and parallelized according to the slave core code, and multiple blocks of little-endian data are stored in the vector register for big-endian conversion to obtain the final parallel scheduling result. Through the scalar register, the slave core code that has been optimized for DMA transmission and parallel scheduling is obtained; based on the slave core code that has been optimized for DMA transmission and parallel scheduling, the instructions of the Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm are equivalently replaced to obtain the slave core code that has been optimized for DMA transmission, parallel scheduling and instruction operation; the slave core code that has been optimized for DMA transmission, parallel scheduling and instruction operation is loop expanded, the assembly code order is adjusted, and the number of intermediate registers used is limited to realize the parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

[0120] See also Figure 2As shown, the computer device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, the parallel scheduling implementation method based on the SHA-256 algorithm in the embodiment is implemented. To avoid repetition, it is not described one by one here. Alternatively, when the computer program 63 is executed by the processor 61, the functions of each model / unit in the parallel scheduling implementation system based on the SHA-256 algorithm in the embodiment are implemented. To avoid repetition, it is not described one by one here.

[0121] The computer device 60 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will appreciate that Figure 2 This is only an example of the computer device 60 and does not constitute a limitation of the computer device 60. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.

[0122] The processor 61 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0123] The memory 62 may be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 60.

[0124] Furthermore, the memory 62 may include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.

[0125] In summary, the present invention provides a method, system, medium and device for implementing parallel scheduling based on the SHA-256 algorithm, which achieves efficient parallel scheduling of the SHA-256 algorithm through all-round performance optimization, including compilation option optimization, DMA transmission optimization, parallel scheduling optimization, instruction reduction optimization, and assembly-level loop unrolling and dual-emission technology optimization. Compared with the original SHA-256 algorithm implementation that optimizes over-memory access efficiency, the operating efficiency is improved by 2.77 times.

[0126] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0127] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0128] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0130] The above contents are only for explaining the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A parallel scheduling implementation method based on SHA-256 algorithm, characterized in that: The following steps are involved: S1. Port the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively; S2, using the optimized DMA transmission mode to process the slave core code and the master core code obtained in step S1, to obtain the master core code and the slave core code after DMA transmission optimization; S3, using the master core code obtained in step S2 to provide a data interface to be transmitted, and performing vector parallelization on the scheduling part of the SHA-256 algorithm according to the slave core code obtained in step S2 to obtain the final parallel scheduling result, and then obtaining the slave core code after DMA transmission optimization and parallel scheduling optimization through the scalar register; S4, according to the slave core code after DMA transmission optimization and parallel scheduling optimization obtained in step S3, the instructions of Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm are equivalently replaced to obtain the slave core code after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization; S5. Perform loop expansion on the slave core code obtained in step S4 after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order, and limit the number of intermediate registers used to achieve parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

2. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 1, characterized in that: Step S1 is specifically as follows: S101, transplant the C code of SHA-256 algorithm in OpenSSL to the Shenwei SW26010 processor, provide a function interface, input is the message length and the pointer of the message, and output is the message digest; S102. Use sw5gcc to compile the main core code of the SHA-256 algorithm with the compile options -O2 and -msimd. S103. Use sw5gcc compiler to compile the slave core code with the compile option -O2-msimd-funroll-all-loops-mslave. S104, using the sw5gcc compiler to link the main core intermediate file and the slave core intermediate file generated by the compilation, linking with the file calling the function interface, generating an executable file, and adding -sw3run swrun-5a running parameters to run the executable file generated by the sw5gcc compilation.

3. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 1, characterized in that: Step S2 is specifically as follows: S201, in the Shenwei SW26010 processor, a static space is opened from the local memory of the core for DMA transmission; S202. The data transmission amount used by the DMA transmission interface is adjusted, and the maximum number of blocks is set to M. For messages with a block number less than M, only one DMA transmission is performed. For messages with a block number greater than M, multiple DMA transmissions are performed to obtain a slave core code optimized for DMA transmission.

4. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 1, characterized in that: Step S3 is specifically as follows: S301, judging the parallel dispatch interface to be called according to the number of blocks of the message, for messages with a length greater than 8 blocks, using 8 dispatch interfaces to process them in sequence, when the number of remaining blocks is less than 8 blocks, calling interfaces for processing different numbers of blocks according to the number of remaining blocks; for messages with a length less than or equal to 8 blocks, directly calling the corresponding interface; S302, load the S block of the message into the vector register, S≤8, convert the little-endian data into big-endian data, and obtain the final message scheduling result; The scheduling result is unloaded from the vector register, and the message compression calculation is completed serially to obtain the slave core code that has been optimized for DMA transmission and parallel scheduling.

5. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 1, characterized in that: In step S4, the Ch operation defined in the SHA-256 standard is transformed into an equivalent mathematical operation; the Maj operation defined in the SHA-256 standard is transformed into an equivalent mathematical operation; The circular shift operation defined in the SHA-256 standard is transformed into an equivalent hardware instruction, and the instruction is called using the simd interface or by calling the assembly language interface to obtain the slave core code that has been optimized for DMA transmission, parallel scheduling, and instruction operation.

6. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 5 is characterized in that: Shift the data v to the left by n bits to realize the equivalent transformation of the circular shift operation defined by the SHA-256 standard and the hardware instruction. The conversion is as follows: (((v)<<(n))|((v)>>(32-(n))))→vrotlw(v,n) The '<<' and '>>' operations represent logical left shift and logical right shift operations, and vrotlw represents the circular shift instruction provided by Shenwei 26010.

7. The parallel scheduling implementation method based on the SHA-256 algorithm according to claim 1, characterized in that: Step S5 is specifically as follows: S501, loop expansion is performed on the slave core code in the assembly code, so that the parallel scheduling and message compression parts are loop expanded; S502: After the parallel scheduling and message compression part loops of step S501 are expanded, memory access instructions are staggered to adjust the instruction sequence; S503. Use 8 registers, 3 of which are registers for storing variables and 5 are intermediate registers for message scheduling. Store instructions in the registers to implement parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

8. A parallel scheduling implementation system based on SHA-256 algorithm, characterized in that: It includes transplantation module, replacement module, parallel module, operation module and implementation module; The transplantation module is used to transplant the C code of the SHA-256 algorithm to the Shenwei SW26010 processor, and use the sw5gcc compiler of the Shenwei SW26010 processor to compile the codes of the slave core and the master core respectively; A replacement module is used to process the slave core and master core codes obtained by the transplantation module using the optimized DMA transmission mode to obtain the slave core code optimized by DMA transmission; The parallel module is used to provide a data interface to be transmitted by using the master core code obtained by the replacement module, and to perform vector parallelization on the scheduling part of the SHA-256 algorithm according to the slave core code obtained by the replacement module to obtain the final parallel scheduling result, and then obtain the slave core code optimized by DMA transmission and parallel scheduling through the scalar register; An operation module is used to perform equivalent replacement of the instructions of the Ch operation, Maj operation and circular shift operation of the SHA-256 algorithm according to the slave core code optimized by DMA transmission and parallel scheduling obtained by the parallel module, so as to obtain the slave core code optimized by DMA transmission, parallel scheduling and instruction operation; The implementation module is used to loop expand the slave core code obtained by the operation module after DMA transmission optimization, parallel scheduling optimization and instruction operation optimization, adjust the assembly code order, and limit the number of intermediate registers used to achieve parallel scheduling of the SHA-256 algorithm on the Shenwei 26010 processor.

9. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 7.

10. A computing device, characterized in that include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to claims 1 to 7.