An Optimized Implementation Method of Ascon Algorithm Based on RISC-V Processor

By optimizing the Ascon algorithm on the RISC-V platform, reconstructing the S-box replacement logic, decomposing the cyclic shift and eliminating the cyclic control overhead, the Ascon algorithm's performance in embedded processors is solved, and the execution efficiency and applicability are improved.

CN120090792BActive Publication Date: 2025-07-25BEIJING YOULUE SECURITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510232971.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The implementation of Ascon algorithm on embedded processors faces insufficient performance in resource-constrained environments. Traditional optimization methods fail to make full use of RISC-V architecture features. There is room for optimization for key operations such as S-box operations and permutation functions.

Method used

The S-box replacement logic is reconstructed through the RISC-V assembly instruction, the 64-bit cyclic shift operation is decomposed using bit interleaving technology, and the cyclic control overhead is eliminated through the assembly level full expansion, and the intermediate state and key are directly stored in the register to form the optimized cyclic shift logic.

Benefits of technology

It significantly improves the execution efficiency and performance of Ascon algorithm on the RISC-V platform, reduces redundant bit operations, reduces computing latency and memory access requirements, and is suitable for resource-constrained embedded systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090792B_ABST
    Figure CN120090792B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of cryptography technology, and in particular to an optimized implementation method of the Ascon algorithm based on the RISC-V processor. The method includes: reconstructing the S-box substitution logic through RISC-V assembly; optimizing the 64-bit cyclic shift through bit interleaving technology; fully expanding the assembly code to eliminate loop overhead; storing the intermediate state and the key in registers. The present invention significantly improves the execution efficiency and performance of the Ascon algorithm on the RISC-V platform through various technical means. The reconstruction of the S-box substitution logic effectively reduces redundant bit operations. By optimizing the bit operation sequence and reusing temporary variables, the calculation delay is reduced. The bit interleaving technology is used to decompose the 64-bit cyclic shift operation, making the operation more efficient. And the loop control overhead is eliminated through the fully expanded technology at the assembly level, avoiding unnecessary jump instructions, further improving the execution speed, and effectively solving the problem of insufficient performance and resource waste of the Ascon algorithm due to the low utilization of the RISC-V architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cryptography, and particularly to an optimized implementation method of the Ascon algorithm based on the RISC-V processor. Background Art

[0002] With the rapid development of emerging technologies such as the Internet of Things (IoT), Vehicle-to-Everything (V2X), and Industrial Internet, more and more devices need to achieve efficient and reliable encrypted communication in resource-constrained environments. These devices usually have limited computing power, storage resources, and energy supply. Therefore, implementing lightweight cryptographic algorithms in these environments has become an important topic in cryptography research and applications.

[0003] As an advanced authenticated encryption (AEAD) and hash algorithm, the Ascon algorithm has received extensive attention in the field of lightweight cryptography due to its high efficiency and security. The Ascon algorithm is designed to balance the efficiency of hardware and software implementations and is applicable to a variety of application scenarios, including but not limited to IoT devices, smart cards, embedded systems, etc. However, although the Ascon algorithm itself has high security, in practical applications, how to further optimize its software implementation to improve the operation speed is still an urgent problem to be solved.

[0004] Currently, the implementation of the Ascon algorithm on embedded processors faces the following challenges: the existing implementation schemes have insufficient performance in resource-constrained embedded environments; traditional optimization methods fail to fully utilize the characteristics of the RISC-V architecture; there is room for optimization in key operations such as S-box operations and permutation functions. Summary of the Invention

[0005] To this end, the present invention provides an optimized implementation method of the Ascon algorithm based on the RISC-V processor, which is used to overcome the problems of insufficient performance and resource waste of the Ascon algorithm caused by low utilization of the RISC-V architecture in the prior art through in-depth optimization targeting the characteristics of the RISC-V architecture.

[0006] To achieve the above object, the present invention provides an optimized implementation method of the Ascon algorithm based on the RISC-V processor, including:

[0007] Reconstruct the S-box substitution logic through RISC-V assembly instructions to form a reconstructed S-box substitution logic;

[0008] Based on the 32-bit register architecture of RISC-V, use the bit-interleaving technique to decompose the 64-bit cyclic shift operation in the reconstructed S-box substitution logic to form an optimized cyclic shift logic;

[0009] Eliminate the loop control overhead of the optimized cyclic shift logic through assembly-level full expansion to form a fully expanded cyclic shift logic;

[0010] Directly store the intermediate state and the key in the fully-expanded cyclic shift logic into registers to form the Ascon optimization algorithm.

[0011] Furthermore, reconstruct the S-box substitution logic through RISC-V assembly instructions to form the reconstructed S-box substitution logic, including:

[0012] Analyze the dependency relationships of the original 22-bit operations and identify duplicate operations that can be merged;

[0013] Use temporary variables to store intermediate results and reduce the number of register reads and writes;

[0014] Redesign the bit operation sequence, optimize the number of bit operations from 22 to 17 to form the reconstructed S-box substitution logic;

[0015] The temporary variables include three intermediate variables (tmp0, tmp1, tmp2), which are used to store the results of bit operations during the S-box substitution process, and reduce redundant calculations by reusing the intermediate variables.

[0016] Furthermore, redesign the bit operation sequence, including:

[0017] Decompose the non-linear operations in the original S-box substitution into a combination of linear operations;

[0018] Use the bit operation instructions of RISC-V to replace multi-step logical operations;

[0019] Hide the operation latency through instruction reordering.

[0020] Furthermore, based on the 32-bit register architecture of RISC-V, use the bit interleaving technique to decompose the 64-bit cyclic shift operation in the reconstructed S-box substitution logic to form the optimized cyclic shift logic, including:

[0021] Divide the 64-bit input into the upper 32 bits and the lower 32 bits;

[0022] Perform a logical left shift operation on the upper 32 bits and a logical right shift operation on the lower 32 bits;

[0023] Generate the final 64-bit cyclic shift result by synthesizing the shifted high and low bits through exclusive OR operation to form the optimized cyclic shift logic;

[0024] The displacement amounts of the logical left shift operation and the right shift operation are dynamically adjusted according to the specific requirements of the Ascon algorithm round function, and the synthesis process of the exclusive OR operation is implemented through the XOR instruction of RISC-V.

[0025] Furthermore, the decomposition of the 64-bit cyclic shift operation using the bit interleaving technique also includes:

[0026] Perform bit-interleaving preprocessing on the upper 32 bits and the lower 32 bits before performing the shift operation;

[0027] Ensure bit alignment of the shifted data through masking operations;

[0028] Perform displacement using the shift instructions of RISC-V.

[0029] Furthermore, eliminate the loop control overhead of the optimized cyclic shift logic through full expansion at the assembly level to form a fully expanded cyclic shift logic, including:

[0030] Repeat the original loop body 12 times to cover all rounds;

[0031] Eliminate the loop counter and jump instructions through inline assembly code;

[0032] Independently optimize the S-box operation for each round of permutation to eliminate function call and return instructions;

[0033] Optimize the instruction ordering to reduce pipeline stalls;

[0034] Eliminate conditional judgments through static branch prediction to form a fully expanded cyclic shift logic.

[0035] Furthermore, static branch prediction includes:

[0036] Convert conditional branches to unconditional jumps;

[0037] Eliminate branch prediction failures through pre-computation of branch target addresses;

[0038] Utilize the delay slot technology of RISC-V to fill the invalid instruction cycles.

[0039] Furthermore, directly store the intermediate states and keys in the fully expanded cyclic shift logic in registers to form the Ascon optimization algorithm, including:

[0040] Allocate the message blocks and key states to the RISC-V general registers;

[0041] Replace array and structure accesses with direct register operations;

[0042] Repeat the original loop body 12 times to cover all rounds;

[0043] Eliminate the loop counter and jump instructions through inline assembly code;

[0044] Independently optimize the S-box operation for each round of permutation to form the Ascon optimization algorithm.

[0045] Further, when storing the messages in groups in the register, the little - endian mode is adopted, and data between memories is exchanged through LW instructions and SW instructions.

[0046] Further, reconstructing the S - box substitution logic through RISC - V assembly instructions further includes;

[0047] Counting the number of clock cycles for a single S - box substitution through the RDTIME instruction of RISC - V;

[0048] Calculating the deviation between the number of clock cycles and a preset reference value;

[0049] When the absolute value of the deviation is greater than a preset deviation threshold, advancing the bit - operation step with the longest execution time in the S - box substitution to form a corrected instruction sequence;

[0050] When the deviation is less than the preset deviation threshold, delaying the bit - operation step with the longest execution time in the S - box substitution to form a corrected instruction sequence;

[0051] For the corrected instruction sequence, using the intermediate variable to store intermediate results to reduce the number of register reads and writes, forming a first temporary instruction sequence;

[0052] Reusing the intermediate variable for the first temporary instruction sequence to reduce redundant calculations, forming a second temporary instruction sequence;

[0053] Updating the allocation priority of the intermediate variable for the second temporary instruction sequence to form an updated instruction sequence;

[0054] Allocating the register and memory according to the life cycle and usage frequency of the intermediate variable to form an allocation instruction sequence;

[0055] Inserting FENCE instructions into the allocation instruction sequence to ensure pipeline synchronization, forming an optimized instruction sequence;

[0056] Integrating the optimized instruction sequence into the S - box substitution logic to form the reconstructed S - box substitution logic.

[0057] Compared with the prior art, the beneficial effects of the present invention are that the execution efficiency and performance of the Ascon algorithm on the RISC-V platform are significantly improved through various technical means. First, the reconstruction of the S-box substitution logic effectively reduces redundant bit operations. By optimizing the bit operation sequence and reusing temporary variables, the number of register reads and writes is reduced, and the calculation latency is decreased. Second, the bit interleaving technique is adopted to decompose the 64-bit cyclic shift operation, making the operation more efficient. And through the assembly-level full expansion technique, the overhead of loop control is eliminated, unnecessary jump instructions are avoided, and the execution speed is further improved. By directly storing the intermediate state and the key in the register, frequent memory access is avoided, and the memory bandwidth requirement is significantly reduced. This not only improves the execution speed of the Ascon algorithm but also reduces the energy consumption, especially suitable for resource-constrained embedded systems or high-performance computing demand scenarios, enhancing its applicability and efficiency in practical applications, and effectively solving the problem of insufficient performance and resource waste of the Ascon algorithm due to low utilization of the RISC-V architecture.

[0058] Furthermore, by reducing redundant calculation steps and memory access, the execution efficiency of the S-box substitution is significantly improved. The redesigned bit operation sequence not only reduces the number of operations but also reduces the calculation redundancy by reusing temporary variables, making the entire S-box substitution process more efficient. This optimization can effectively improve the execution speed of the Ascon algorithm in resource-constrained environments, especially suitable for applications such as embedded systems and the Internet of Things, enhancing its performance and practicality.

[0059] Furthermore, through a series of optimization measures, not only are the complex operations in the S-box substitution process reduced, but the execution efficiency of the algorithm is also improved. Decomposing the non-linear operation and replacing it with a linear operation can accelerate the calculation process. At the same time, by using the efficient bit operation instructions of RISC-V, the complexity of multi-step operations is reduced. The use of instruction reordering effectively reduces the latency and further improves the parallelism of instruction execution, thus significantly improving the overall performance and meeting the requirements of embedded processors for fast and efficient execution.

[0060] Furthermore, through the bit interleaving technique and XOR synthesis, the complexity in the 64-bit cyclic shift operation is significantly reduced, and the shift amount is flexibly adjusted to meet the algorithm requirements. This optimization makes the calculation process more efficient, reduces the resource consumption, enhances the implementation ability on embedded systems, and further improves the performance of the Ascon algorithm in resource-constrained environments.

[0061] Furthermore, through bit interleaving preprocessing and masking operations, the complexity of data movement and alignment is effectively reduced, and the accuracy and efficiency of shift operations are optimized. By reasonably decomposing 64-bit operations into 32-bit parts, the computational burden of each operation is reduced, making the whole process more efficient, while improving the parallelism of shift operations. The shift instructions in the RISC-V instruction set are used to further optimize the hardware execution efficiency, achieving faster and more accurate 64-bit cyclic shifts, providing a solid foundation for the efficient implementation of the Ascon algorithm in embedded systems.

[0062] Furthermore, through the fully expanded cyclic shift logic, this embodiment significantly reduces the control instructions and function calls in the program, improving the execution efficiency of the processor. The optimization of inline assembly and instruction ordering avoids pipeline stalls and improves the parallel processing ability of instructions, thus significantly enhancing the execution speed of the algorithm. In addition, static branch prediction reduces the overhead of conditional judgments, making the execution of each round of operations smoother and further optimizing the overall performance of the algorithm. These optimization strategies enable the Ascon algorithm to run more efficiently in resource-constrained embedded systems.

[0063] Furthermore, through static branch prediction optimization, the system can reduce the performance loss caused by conditional judgments and branch operations. Unconditional jumps and pre-computed branch target addresses ensure more efficient pipeline utilization and avoid the additional latency of branch prediction failures. The delay slot technique fills the idle instruction cycles, effectively improving the instruction throughput rate of the processor, enabling the Ascon algorithm to execute more quickly in embedded systems, and thus significantly enhancing the overall performance of encryption operations.

[0064] Furthermore, by reducing memory accesses and instruction jumps and eliminating unnecessary loop control overhead, the execution efficiency of the Ascon algorithm is significantly improved, and pipeline stalls and instruction latency are reduced. At the same time, by directly storing and operating on intermediate states and keys in registers, the data access speed is increased, the register usage efficiency is optimized, and the performance on resource-constrained embedded processors is further enhanced. This optimization can significantly improve its application efficiency in embedded systems such as the Internet of Things and vehicle-to-everything while ensuring the security of the algorithm, and is particularly suitable for scenarios with high requirements for computing speed.

[0065] Furthermore, using little-endian mode storage can provide efficient memory access in most modern processors and systems. At the same time, data exchange is performed using LW instructions and SW instructions, effectively reducing data access latency. By directly storing message blocks in registers, the frequent read and write operations to memory can be reduced, improving the execution efficiency of the encryption algorithm. In addition, this method also improves the flexibility and portability of data exchange, ensuring compatibility across different platforms.

[0066] Furthermore, by dynamically adjusting the operation order and optimizing the timing, the waste of clock cycles can be reduced, the execution time can be decreased, and the performance can be improved. By optimizing the use and reuse of intermediate variables, unnecessary register reads and writes are reduced, further reducing the consumption of hardware resources. By reasonable register and memory allocation and pipeline synchronization, the efficient utilization of computing resources is ensured, thus achieving the overall optimization of the S-box substitution logic and improving the execution efficiency of the algorithm in resource-constrained embedded systems. Description of the Drawings

[0067] Figure 1 This is the flowchart of the optimized implementation method of the Ascon algorithm based on the RISC-V processor in this embodiment;

[0068] Figure 2 This is the flowchart of forming the reconstructed S-box substitution logic in this embodiment;

[0069] Figure 3 This is the flowchart of optimizing the cyclic shift logic in this embodiment;

[0070] Figure 4 This is the flowchart of fully expanding the cyclic shift logic in this embodiment. Detailed Implementation Manner

[0071] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0072] The preferred embodiments of the present invention will be described below with reference to the drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0073] Please refer to Figure 1 as shown, which is the flowchart of the optimized implementation method of the Ascon algorithm based on the RISC-V processor in this embodiment;

[0074] Based on resource-constrained environments such as the Internet of Things and the Internet of Vehicles, this embodiment provides an optimized implementation method of the Ascon algorithm based on the RISC-V processor, including:

[0075] Reconstructing the S-box substitution logic through RISC-V assembly instructions to form a reconstructed S-box substitution logic;

[0076] Based on the 32-bit register architecture of RISC-V, using the bit-interleaving technique to decompose the 64-bit cyclic shift operation in the reconstructed S-box substitution logic to form an optimized cyclic shift logic;

[0077] The loop control overhead of the optimized cyclic shift logic is eliminated by full expansion at the assembly level to form a fully expanded cyclic shift logic;

[0078] The intermediate state and key in the fully expanded cyclic shift logic are directly stored in registers to form the Ascon optimized algorithm.

[0079] RISC-V is an open-source reduced instruction set architecture (RISC) for processor design. It is efficient, flexible, and extensible, and is widely used in embedded systems, Internet of Things (IoT), and other fields. RISC-V assembly instructions are low-level instructions that directly interact with the RISC-V architecture and are used to control hardware and perform computational tasks.

[0080] Assembly instructions are low-level language instructions for computer hardware, through which every operation of the processor can be precisely controlled. Assembly language is close to machine language, so it is efficient.

[0081] The S box (Substitution Box) is a commonly used substitution structure in cryptography, used to replace each part of the data to enhance the security of the encryption algorithm. It is usually used to perform complex substitutions on the input data to increase the difficulty of cracking.

[0082] The S box substitution logic refers to the rules and algorithms for processing input data through the S box. In encryption algorithms such as Ascon, the S box substitution logic is an important step to improve the algorithm's security.

[0083] The bit interleaving technique is a data processing technique that interleaves multiple data streams or multiple bits to process multiple parts of data simultaneously and enhance data parallelism. Through this technique, the latency of data processing can be reduced and the computational efficiency can be improved. In the optimization of the Ascon algorithm, the Bit Interleaving technique can be used to decompose 64-bit data processing into multiple smaller data parts to optimize the processing efficiency.

[0084] A 64-bit cyclic shift refers to a displacement operation on 64-bit data, and the bits that exceed after the shift will reappear at the other end of the data. This operation is often used for data scrambling in encryption algorithms to enhance the complexity and difficulty of encryption.

[0085] Full expansion at the assembly level means expanding the loop body in the assembly code to avoid loop control overhead. By repeatedly expanding the loop operation into fixed code, the performance overhead brought by loop counters, conditional jumps, etc. is eliminated.

[0086] Registers are efficient storage units inside the processor for quickly storing and accessing data. Registers are extremely fast and work directly inside the processor, avoiding the latency of accessing memory. The 32-bit registers in the RISC-V processor are used to store operands, intermediate data, or processing results.

[0087] The intermediate state refers to the temporary result generated after data is processed through multiple steps during the execution of an algorithm. The intermediate state is commonly used in various stages of encryption algorithms.

[0088] A key is the secret data used for encryption and decryption operations. It is usually generated by the encryption algorithm and continuously used during the execution of the algorithm.

[0089] The Ascon algorithm is a modern lightweight authenticated encryption algorithm, commonly used in resource-constrained devices that require efficient encryption processing, such as Internet of Things (IoT) devices. The Ascon algorithm provides authenticated encryption (providing both encryption and authentication) functionality and is suitable for environments that require high security.

[0090] By reconstructing the S-box substitution logic in the Ascon algorithm using RISC-V assembly instructions, the redundant calculations in the original bit operations are reduced and the bit operation sequence is optimized. The bit interleaving technique is used to decompose the 64-bit cyclic shift operation, and the loop control overhead is eliminated through the assembly-level full expansion technique, improving the execution efficiency. At the same time, the intermediate state and the key are directly stored in the registers, avoiding frequent memory accesses. Through these optimization measures, the execution speed and efficiency of the Ascon algorithm on the RISC-V platform are significantly improved.

[0091] The execution efficiency and performance of the Ascon algorithm on the RISC-V platform are significantly improved through various technical means. First, the reconstruction of the S-box substitution logic effectively reduces redundant bit operations. By optimizing the bit operation sequence and reusing temporary variables, the number of register reads and writes is reduced, and the calculation latency is decreased. Second, the bit interleaving technique is used to decompose the 64-bit cyclic shift operation, making the operation more efficient, and the loop control overhead is eliminated through the assembly-level full expansion technique, avoiding unnecessary jump instructions and further improving the execution speed. By directly storing the intermediate state and the key in the registers, frequent memory accesses are avoided, significantly reducing the memory bandwidth requirement. This not only improves the execution speed of the Ascon algorithm but also reduces power consumption, especially suitable for resource-constrained embedded systems or high-performance computing demand scenarios, enhancing its applicability and efficiency in practical applications and effectively solving the problem of insufficient performance and resource waste of the Ascon algorithm due to low utilization of the RISC-V architecture.

[0092] Please continue to participate Figure 2 As shown, it is the flowchart for forming the reconstructed S-box substitution logic in this embodiment;

[0093] Reconstruct the S-box substitution logic through RISC-V assembly instructions to form the reconstructed S-box substitution logic, including:

[0094] Analyze the dependency relationships of the original 22 bit operations and identify duplicate operations that can be merged;

[0095] Use temporary variables to store intermediate results and reduce the number of register reads and writes;

[0096] Redesign the bit operation sequence, optimize the number of bit operations from 22 to 17 to form the reconstructed S-box substitution logic;

[0097] The temporary variables include three intermediate variables (tmp0, tmp1, tmp2), which are used to store the results of bit operations during the S-box substitution process, and reduce redundant calculations by reusing the intermediate variables.

[0098] Dependency relationships are the relationships between various operations during the calculation process, and some operations depend on the results of previous operations. Analyzing the dependency relationships helps identify parts that can be optimized.

[0099] In the original algorithm, there are multiple operations performing similar or identical tasks. Merging duplicate operations improves efficiency by reducing redundant calculation steps.

[0100] Bit operations refer to operations on binary bits, including operations such as AND, OR, XOR, left shift (<<), and right shift (>>). In encryption algorithms, bit operations are often used for data transformation and confusion.

[0101] Reusing intermediate variables means reusing the already stored intermediate results during the calculation process instead of creating new temporary variables. Reusing intermediate variables can reduce redundant calculations and memory operations and improve efficiency.

[0102] In this embodiment, the S-box substitution logic is reconstructed through RISC-V assembly instructions. First, the dependency relationships of the original 22 bit operations are analyzed, duplicate operations are identified and merged, thereby reducing unnecessary calculation steps. Then, temporary variables (such as tmp0, tmp1, tmp2) are used to store the intermediate results generated during the S-box substitution process to reduce frequent reads and writes to registers and optimize memory access. By redesigning the bit operation sequence, the steps that originally required 22 bit operations are optimized to 17, thus improving the operation efficiency and reducing the processing time.

[0103] By reducing redundant computational steps and memory accesses, the execution efficiency of the S-box substitution is significantly improved. The redesigned bit operation sequence not only reduces the number of operations but also decreases computational redundancy by reusing temporary variables, making the entire S-box substitution process more efficient. This optimization can effectively enhance the execution speed of the Ascon algorithm in resource-constrained environments, especially suitable for applications such as embedded systems and the Internet of Things, enhancing its performance and practicality.

[0104] Specifically, the bit operation sequence is redesigned, including:

[0105] Decompose the non-linear operations in the original S-box substitution into a combination of linear operations;

[0106] Use the bit operation instructions of RISC-V to replace multi-step logical operations;

[0107] Hide the operation latency through instruction reordering.

[0108] Non-linear operations refer to operations where there is no direct linear relationship between the input and output. In encryption algorithms, non-linear operations increase the complexity of the algorithm, making it difficult to reverse engineer or crack. For example, S-boxes usually perform non-linear operations, converting the input bit pattern into a completely different output, thereby enhancing the encryption strength.

[0109] Linear operations refer to operations where there is a direct proportional relationship between the input and output, or operations completed through simple operations such as addition and multiplication. Compared with non-linear operations, linear operations are relatively simple and easy to predict. Decomposing non-linear operations into a combination of linear operations can simplify the calculation and improve efficiency.

[0110] Instruction reordering refers to rearranging the execution order of instructions when executing a program, thereby improving performance. In the pipeline design of a processor, reordering instructions helps to hide the execution latency, reduce resource contention, and thus enhance the computational efficiency. Through reordering, some delays caused by dependency relationships can be avoided, optimizing the execution time.

[0111] Operation latency refers to the time required for an operation to start and complete. In a processor, each operation (such as addition, multiplication, or bit operation) will have a certain latency. The instruction reordering technique hides these latencies by adjusting the order of operations, improving the overall efficiency of the system.

[0112] Optimize the performance by redesigning the bit operation sequence in the S-box substitution. First, decompose the non-linear operations in the original S-box substitution and transform them into a combination of a series of linear operations to reduce the computational complexity. Then, utilize the bit operation instructions provided by the RISC-V architecture, such as exclusive OR (XOR) and bit shift, to replace multi-step logical operations, thus simplifying the calculation process. Finally, through instruction reordering, arrange some operations at different stages of the pipeline to hide the operation latency and improve the instruction execution efficiency.

[0113] Through a series of optimization measures, not only are the complex operations in the S-box substitution reduced, but also the execution efficiency of the algorithm is improved. Decomposing non-linear operations and replacing them with linear operations can accelerate the calculation process. At the same time, by utilizing the efficient bit operation instructions of RISC-V, the complexity of multi-step operations is reduced. The use of instruction reordering effectively reduces the latency and further improves the parallelism of instruction execution, thus significantly enhancing the overall performance and meeting the requirements of embedded processors for fast and efficient execution.

[0114] Please continue to participate Figure 3 As shown, it is the flowchart for optimizing the cyclic shift logic in this embodiment;

[0115] Based on the 32-bit register architecture of RISC-V, the bit interleaving technique is used to perform the decomposed 64-bit cyclic shift operation in the reconstructed S-box substitution logic, and the optimized cyclic shift logic is formed as follows:

[0116] Split the 64-bit input into the high 32 bits and the low 32 bits;

[0117] Perform a logical left shift operation on the high 32 bits and a logical right shift operation on the low 32 bits;

[0118] Generate the final 64-bit cyclic shift result by synthesizing the shifted high and low bits through XOR operation, and form the optimized cyclic shift logic;

[0119] The displacement amounts of the logical left shift operation and the right shift operation are dynamically adjusted according to the specific requirements of the Ascon algorithm round function, and the synthesis process of the XOR operation is implemented through the XOR instruction of RISC-V.

[0120] Logical left shift is a displacement operation that means shifting each binary bit of the data to the left, and filling zeros from the right when shifting. For example, shifting left by 1 bit moves each bit in the data one bit to the left, the leftmost bit is lost, and a zero is filled on the rightmost bit.

[0121] Logical right shift is a displacement operation opposite to logical left shift, which shifts each binary bit of the data to the right, and fills zeros from the left when shifting. For example, shifting right by 1 bit moves each bit in the data one bit to the right, the rightmost bit is lost, and a zero is filled on the leftmost bit.

[0122] Exclusive OR operation is a common logical operation, with symbols "^" or "⊕". Its rule is: if the corresponding bits of two operands are the same, the result is 0; if they are different, the result is 1. For example, 1^0 = 1, 0^1 = 1, 1^1 = 0. Exclusive OR operation is often used in encryption and verification because it has some important mathematical properties, such as reflexivity (A⊕A = 0) and commutativity (A⊕B = B⊕A).

[0123] The synthesized shifted data refers to the data obtained by combining two different parts of data through exclusive OR operation after completing separate logical shift operations (left shift and right shift). This synthesis operation can ensure that no information is lost during the shift process through exclusive OR operation.

[0124] In many encryption algorithms, especially symmetric encryption algorithms, the round function is used to process data multiple times to improve the security of encryption. Each round function usually includes various operations, such as displacement, substitution, mixing, etc., to continuously increase the complexity and anti-attack ability of the encryption process.

[0125] The displacement amount refers to how many bits are shifted each time during the displacement operation. The displacement amount can be adjusted according to the requirements of the algorithm, affecting the result and efficiency of the shift operation.

[0126] The bit interleaving technique is used to optimize the 64-bit input data. After splitting it into the upper 32 bits and the lower 32 bits, logical left shift and right shift operations are performed on the two parts of data respectively. The two shifted parts of data are synthesized through exclusive OR operation to form the final 64-bit circular shift result. During this process, the displacement amount of the shift operation is dynamically adjusted according to the requirements of the round function of the Ascon algorithm, and all shift operations are implemented through the XOR instruction of RISC-V to improve the operation efficiency and accuracy.

[0127] Through the bit interleaving technique and exclusive OR synthesis, the complexity in the 64-bit circular shift operation is significantly reduced, and the algorithm requirements are met by flexibly adjusting the displacement amount. This optimization makes the calculation process more efficient, reduces resource consumption, enhances the implementation ability on embedded systems, and further improves the performance of the Ascon algorithm in resource-constrained environments.

[0128] Specifically, the decomposition of the 64-bit circular shift operation using the bit interleaving technique also includes:

[0129] Performing bit interleaving preprocessing on the upper 32 bits and the lower 32 bits before the shift operation;

[0130] Ensuring the alignment of the shifted data bits through mask operations;

[0131] Using the shift instructions of RISC-V for displacement.

[0132] Bit interleaving preprocessing refers to rearranging each bit position in the data before performing an operation, usually by alternately combining the data in a specific way. In this embodiment, it refers to alternately arranging the segmented data of the high 32 bits and the low 32 bits to facilitate the subsequent shift operation.

[0133] Mask operation is a technique for operating on data by using a mask (usually a binary bit pattern). It is often used to extract specific bits from data, clear or set bits, etc. In this embodiment, the purpose of the mask operation is to ensure that after the shift operation, the data bits are correctly aligned to avoid overflow or position disorder.

[0134] Shift instructions are instructions in the computer instruction set used to perform bit shift operations on data, which can shift each bit of the data to the left or right. In the RISC-V processor, shift instructions allow these bit operations to be performed quickly and efficiently. Especially when decomposing 64-bit data and then performing shift operations on the 32-bit parts separately, shift instructions can effectively improve performance.

[0135] When optimizing the 64-bit cyclic shift operation, first, the bit interleaving technique is adopted to divide the 64-bit input data into two parts: the high 32 bits and the low 32 bits. Then, through bit interleaving preprocessing, the data bits of these two 32-bit parts are cross-arranged, so that the subsequent shift operation can be performed more efficiently. Subsequently, the mask operation is used to ensure that the data bits after the shift are correctly aligned, avoiding information loss or position errors caused by data overflow. Finally, the shift instructions in the RISC-V architecture are used to perform independent logical left shift or right shift operations on the high and low 32 bits to complete the 64-bit cyclic shift.

[0136] Through bit interleaving preprocessing and mask operation, the complexity of data movement and alignment is effectively reduced, and the accuracy and efficiency of the shift operation are optimized. By reasonably decomposing the 64-bit operation into 32-bit parts, the computational burden of each operation is reduced, making the whole process more efficient. At the same time, the parallelism of the shift operation is improved. Using the shift instructions in the RISC-V instruction set further optimizes the hardware execution efficiency, realizing a faster and more accurate 64-bit cyclic shift, providing a solid foundation for the efficient implementation of the Ascon algorithm in embedded systems.

[0137] Please continue to participate Figure 4 As shown, it is the flowchart of the fully expanded cyclic shift logic of this embodiment;

[0138] For the optimized cyclic shift logic, the loop control overhead is eliminated through assembly-level full expansion to form a fully expanded cyclic shift logic, including:

[0139] Repeat the original loop body 12 times to cover all rounds;

[0140] Eliminating loop counters and jump instructions through inline assembly code;

[0141] Independently optimize the S-box operations for each round of permutation to eliminate function call and return instructions;

[0142] Optimize instruction scheduling to reduce pipeline stalls;

[0143] Eliminate conditional judgments through static branch prediction to form a fully unrolled cyclic shift logic.

[0144] Full unrolling at the assembly level is a technique that eliminates loop control overhead by repeatedly unrolling the loop body. In embedded systems, loop control (such as judgments, jumps, etc.) consumes precious processing time. The full unrolling technique writes the operations in the loop multiple times, avoiding frequent jumps or judgments during runtime, thus saving instruction execution time, especially in cases where a large amount of repeated calculations are required.

[0145] Repeating the loop body 12 times to cover all rounds means unrolling the operations that originally needed to be processed through a loop multiple times. Each round of operation corresponds to a set of independent instructions, which can avoid the operations on the counter and the execution of jump instructions during each loop. Repeating 12 times is because the Ascon algorithm usually contains 12 rounds of processing. After unrolling, the code no longer depends on the loop counter, but each round of operation is directly executed through the same code multiple times.

[0146] Inline assembly code is a technique of embedding assembly language instructions in high-level language code (such as C language). Through inline assembly, developers can manually optimize specific parts of the code and directly control low-level hardware operations to improve performance. For example, inline assembly can eliminate unnecessary loop counter operations or jump instructions, making the code execution more efficient.

[0147] Eliminating loop counters and jump instructions means that in a traditional loop structure, during each iteration, the program needs to determine whether to continue the loop and jump to the next iteration. This process consumes time and may affect performance, especially in embedded systems. Through full unrolling and inline assembly, these jump instructions and loop counter operations can be removed, thereby reducing unnecessary overhead and improving execution efficiency.

[0148] Eliminating function call and return instructions means that by independently optimizing and inlining the S-box operations in the code, function call and return instructions can be eliminated, avoiding context switching during each encryption operation, thus improving efficiency.

[0149] Static branch prediction is based on the static structure of the program and does not rely on dynamic runtime information. Through this method, the processor can make branch decisions in advance, reducing pipeline stalls caused by waiting for branch results.

[0150] Conditional judgments (such as if statements) usually cause branch jumps, which affect the execution efficiency of the program. By using static branch prediction and optimization techniques, conditional judgments can be converted into unconditional jumps or more efficient instruction execution methods, avoiding pipeline stalls and delays and improving the execution efficiency of the program.

[0151] Eliminate the loop control overhead through assembly-level full expansion technology and optimize the loop shift logic. First, repeat the original loop body 12 times to cover all rounds, thus avoiding the time overhead of loop control. Through inline assembly code, the loop counter and jump instructions are removed, reducing the execution of control instructions. Each round of S-box operation permutation is independently optimized, thus avoiding the overhead of function calls and returns and improving the execution efficiency. In addition, the instruction order is optimized to reduce pipeline stalls and improve the efficiency of the processor pipeline. At the same time, through static branch prediction technology, conditional judgments are eliminated, further optimizing the execution speed.

[0152] By fully expanding the loop shift logic, this embodiment significantly reduces the control instructions and function calls in the program and improves the execution efficiency of the processor. The optimization of inline assembly and instruction order avoids pipeline stalls and improves the parallel processing ability of instructions, thus significantly improving the execution speed of the algorithm. In addition, static branch prediction reduces the overhead of conditional judgments, making the execution of each round of operations more smooth and further optimizing the overall performance of the algorithm. These optimization strategies enable the Ascon algorithm to run more efficiently in resource-constrained embedded systems.

[0153] Specifically, static branch prediction includes:

[0154] Convert conditional branches to unconditional jumps;

[0155] Eliminate branch prediction failures by pre-computing branch target addresses;

[0156] Use the delay slot technology of RISC-V to fill the invalid instruction cycles.

[0157] A conditional branch is a branch instruction in which the program decides the execution path based on the result of a conditional judgment (such as an if statement). If the condition is true, a section of code is executed; otherwise, another section of code is executed. Such branch operations may cause interruptions and delays in the processor pipeline.

[0158] An unconditional jump is a jump instruction that does not depend on any conditions, and the program always jumps to the specified target address. Such jump instructions are not affected by the program state and are therefore easier to predict and execute in the pipeline.

[0159] The branch target address refers to the address to which the branch instruction jumps after execution. When a conditional branch instruction is executed, the processor must determine the jump address to continue executing the subsequent instructions in the program.

[0160] A branch prediction failure means that the branch direction predicted by the processor is inconsistent with the actual executed branch direction, causing the processor to discard the loaded instructions and reload the correct ones, resulting in performance loss. A branch prediction failure causes pipeline stalls and affects the execution efficiency.

[0161] A pipeline stall refers to the interruption or suspension of the processor's instruction pipeline, usually due to issues such as dependencies, branch prediction failures, or resource conflicts. The stall causes a delay in the processor's execution.

[0162] The delayed slot technique refers to inserting idle instruction cycles before or after a branch instruction. These instructions do not affect the execution of the branch, and the processor can use this time to execute instructions that do not affect the program result, reducing the delay caused by the branch jump. The RISC-V architecture uses delayed slots to fill the idle cycles in the pipeline.

[0163] Precomputation means calculating certain values or results before the program runs, avoiding repeated calculations during runtime, and improving the program's execution efficiency. In branch prediction, precomputing the branch target address can reduce the calculation of the target address during runtime, thus avoiding pipeline stalls.

[0164] The jump target address is the address to which the jump instruction jumps when executed, similar to the target address of a branch instruction. When the jump instruction is executed, the program control flow transfers to this address.

[0165] In the optimization process of static branch prediction, first, by converting conditional branches to unconditional jumps, the overhead of performing branch judgments is eliminated, making the program flow smoother and reducing the performance loss caused by branch instructions. Then, by precomputing the branch target address, the jump target is determined in advance, thus eliminating the pipeline stalls caused by branch prediction failures and ensuring the smooth execution of instructions. Finally, using the delayed slot technique of RISC-V, instructions that do not need to be executed are inserted into the idle instruction cycles before and after the branch, making full use of the idle time in the pipeline and further improving the execution efficiency.

[0166] Through static branch prediction optimization, the system can reduce the performance loss caused by conditional judgments and branch operations. Unconditional jumps and precomputed branch target addresses ensure more efficient use of the pipeline, avoiding the additional delay of branch prediction failures. The delayed slot technique fills the idle instruction cycles, effectively improving the instruction throughput rate of the processor, enabling the Ascon algorithm to execute more quickly in embedded systems, and thus significantly improving the overall performance of encryption operations.

[0167] Specifically, directly storing the intermediate state and key in the register in the fully-expanded cyclic shift logic forms the Ascon optimization algorithm, including:

[0168] Allocating the message block and key state to the RISC-V general-purpose registers;

[0169] Replacing array and structure accesses with direct register operations;

[0170] Repeating the original loop body 12 times to cover all rounds;

[0171] Eliminating the loop counter and jump instructions through inline assembly code;

[0172] Independently optimizing the S-box operation for each round of permutation to form the Ascon optimization algorithm.

[0173] General-purpose registers are hardware units inside the processor for temporarily storing data. They are faster than memory (RAM) and can accelerate data reading and writing operations. In the RISC-V architecture, registers are widely used to store intermediate results and states during calculations.

[0174] Message block means splitting the input data (usually a long message) into multiple blocks of fixed size. Each block can be processed independently, which helps accelerate encryption and decryption operations.

[0175] Key state refers to the internal representation of the key used in the encryption algorithm during the calculation process, including the current key value and other information required during the encryption process.

[0176] Replacing array and structure accesses with direct register operations means directly storing data in the registers without loading from memory, which can reduce the overhead of memory access and improve efficiency.

[0177] By directly storing the intermediate state and key in the general-purpose registers of RISC-V in the Ascon algorithm, the speed and efficiency of data processing are optimized. The specific steps include: allocating the message block and key state to the registers, avoiding the overhead of traditional array and structure accesses and reducing the latency of memory access; covering the 12 rounds of the Ascon algorithm by repeating the loop body 12 times to ensure that all rounds are efficiently processed; eliminating the counter and jump instructions required for loop control through inline assembly code, reducing the complexity of processor instructions; in addition, the S-box operation in each round of permutation is also independently optimized, further improving the execution efficiency.

[0178] By reducing memory accesses and instruction jumps, eliminating unnecessary loop control overhead, the execution efficiency of the Ascon algorithm is significantly improved, and pipeline stalls and instruction latencies are reduced. At the same time, by directly storing and operating on intermediate states and keys in registers, the data access speed is increased, the register usage efficiency is optimized, and the performance on resource-constrained embedded processors is further enhanced. This optimization can significantly improve its application efficiency in embedded systems such as the Internet of Things and the Internet of Vehicles while ensuring the security of the algorithm, and is particularly suitable for scenarios with high requirements for computing speed.

[0179] Specifically, when storing the message block in the register, the little-endian mode is adopted, and data between memories is exchanged through LW instructions and SW instructions.

[0180] LW instructions and SW instructions are common memory access instructions in the RISC-V architecture, used to transfer data between registers and memories.

[0181] The little-endian mode is a data storage format used to define the storage order of multi-byte data in computer memory. In the little-endian mode, the least significant byte of the data is stored at the earliest memory address, while the most significant byte is stored at the last memory address.

[0182] In this embodiment, when the message block is stored in the register, the little-endian mode is adopted, that is, the least significant byte is stored at the smallest address position. To ensure efficient data exchange, data is transferred between the register and the memory through LW instructions and SW instructions. The specific operation is to load the grouped message from the memory into the register, and then save the processing result back to the memory through the store instruction. This method makes full use of the basic instruction set of the RISC-V architecture and combines the high-speed access speed of the register to achieve smooth data exchange.

[0183] Storing in the little-endian mode can provide efficient memory access in most modern processors and systems. At the same time, using LW instructions and SW instructions for data exchange effectively reduces data access latency. By directly storing the message block in the register, frequent reads and writes to the memory can be reduced, and the execution efficiency of the encryption algorithm can be improved. In addition, this method also improves the flexibility and portability of data exchange, ensuring compatibility on different platforms.

[0184] Specifically, reconstructing the S-box substitution logic through RISC-V assembly instructions also includes;

[0185] Counting the number of clock cycles for a single S-box substitution through the RDTIME instruction of RISC-V;

[0186] Calculating the deviation between the number of clock cycles and a preset reference value;

[0187] When the absolute value of the deviation is greater than the preset deviation threshold, advance the bit operation step with the longest execution time in the S-box substitution to form a corrected instruction sequence;

[0188] When the deviation is less than the preset deviation threshold, delay the bit operation step with the longest execution time in the S-box substitution to form a corrected instruction sequence;

[0189] For the corrected instruction sequence, use the intermediate variable to store the intermediate results to reduce the number of read and write operations of the register, and form a first temporary instruction sequence;

[0190] For the first temporary instruction sequence, reuse the intermediate variable to reduce redundant calculations and form a second temporary instruction sequence;

[0191] For the second temporary instruction sequence, update the allocation priority of the intermediate variable to form an updated instruction sequence;

[0192] Allocate the register and memory according to the life cycle and usage frequency of the intermediate variable to form an allocation instruction sequence;

[0193] Insert the FENCE instruction into the allocation instruction sequence to ensure pipeline synchronization and form an optimized instruction sequence;

[0194] Integrate the optimized instruction sequence into the S-box substitution logic to form the reconstructed S-box substitution logic.

[0195] The RDTIME instruction is an instruction in the RISC-V instruction set used to obtain the current timestamp counter of the processor (the timestamp counter is a hardware counter used to measure time intervals), which can be used to calculate the execution duration of a code segment and is usually used for performance analysis.

[0196] The number of clock cycles represents the number of clock cycles required for the processor to complete a certain operation. Each clock cycle is the time unit for the processor to perform an operation. By measuring the number of clock cycles, the performance and efficiency of the operation can be understood.

[0197] The life cycle refers to the time range from the creation to the destruction of the intermediate variable during the program execution. During the life cycle, the intermediate variable will be accessed and modified multiple times.

[0198] The usage frequency refers to the number of times the intermediate variable is accessed during the program execution. Variables with high usage frequency usually require better storage and access strategies to improve program performance.

[0199] The FENCE instruction is an instruction in RISC-V used to control the execution order of instructions to ensure that a specific operation order is followed. In pipeline processing, the FENCE instruction can ensure that subsequent operations are performed after the previous operations are completed, avoiding out-of-sync situations.

[0200] Pipeline synchronization means ensuring that instructions are executed in the correct order and timing in the instruction pipeline. Since the RISC-V processor adopts a pipeline architecture, the FENCE instruction can help synchronize the pipeline and avoid conflicts and inconsistencies between instructions.

[0201] The preset reference value refers to the ideal number of clock cycles when the S-box substitution operation is executed under normal circumstances, which is used as a benchmark for performance evaluation. It depends on the performance of the hardware platform, the optimization goal, and the complexity of the algorithm. Usually, it is set between 10 and 100 clock cycles. In this embodiment, it is set to 50 clock cycles. The difference between the actual execution time and the expected execution time can be monitored, and then dynamic adjustment can be made to optimize the execution efficiency.

[0202] The preset deviation threshold refers to the maximum deviation value of the number of clock cycles allowed when the S-box substitution operation is executed, which is used to determine whether the operation steps need to be corrected. It depends on the clock frequency of the hardware, the characteristics of the operation steps, and the performance requirements of the system. Usually, it is set between 5 and 10 clock cycles. In this embodiment, it is set to 8 clock cycles. When the deviation of the number of clock cycles is too large, the operation steps can be adjusted to avoid excessive delay and ensure the accuracy and efficiency of optimization.

[0203] Reconstruct the S-box substitution logic through RISC-V assembly instructions. First, use the RDTIME instruction to count the number of clock cycles of a single S-box substitution and calculate the deviation from the preset reference value. According to the size of the deviation, dynamically adjust the bit operation steps with the longest execution time in the S-box substitution, and optimize the execution order by advancing or delaying the execution respectively. Then, use intermediate variables to store intermediate results, reduce the number of register reads and writes, and reduce redundant calculations by reusing intermediate variables. Subsequently, update the allocation priority of intermediate variables, allocate registers and memory according to the lifetime and usage frequency, and insert the FENCE instruction to ensure pipeline synchronization. Finally, form an optimized instruction sequence and integrate it into the S-box substitution logic.

[0204] By dynamically adjusting the operation order and optimizing the timing, the waste of clock cycles can be reduced, the execution time can be shortened, and the performance can be improved. Through the optimized use and reuse of intermediate variables, unnecessary register reads and writes are reduced, and further consumption of hardware resources is reduced. Through reasonable register and memory allocation and pipeline synchronization, the efficient use of computing resources is ensured, thus realizing the overall optimization of the S-box substitution logic and improving the execution efficiency of the algorithm in resource-constrained embedded systems.

[0205] The method for optimizing the Ascon algorithm based on the RISC-V processor provided by this embodiment has achieved remarkable results in terms of performance improvement, as follows:

[0206] The performance of Ascon-128 is improved by 44.1%, and the execution efficiency is significantly increased to 99 cycles / byte;

[0207] The performance of Ascon-128a is improved by 47.3%, and the execution efficiency is greatly increased to 69 cycles / byte;

[0208] The performance of Ascon-80pq is improved by 44.1%, and the execution efficiency is significantly increased to 99 cycles / byte;

[0209] The performance of Ascon-hash / xof is improved by 38.5%, and the execution efficiency is significantly increased to 184 cycles / byte;

[0210] The performance of Ascon-hasha / xofa is improved by 40.1%, and the execution efficiency is greatly increased to 124 cycles / byte.

[0211] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. An optimized implementation method of the Ascon algorithm based on the RISC-V processor, characterized in that Including: Reconstruct the S-box substitution logic through RISC-V assembly instructions to form the reconstructed S-box substitution logic; Based on the 32-bit register architecture of RISC-V, use bit-interleaving technology to decompose the 64-bit cyclic shift operation in the reconstructed S-box substitution logic to form an optimized cyclic shift logic; Eliminate the loop control overhead of the optimized cyclic shift logic through assembly-level full expansion to form a fully expanded cyclic shift logic; Directly store the intermediate state and key in the register in the fully expanded cyclic shift logic to form the Ascon optimization algorithm.

2. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 1, wherein Reconstruct the S-box substitution logic through RISC-V assembly instructions to form the reconstructed S-box substitution logic, including; Analyze the dependency relationship of the original 22-bit operations and identify duplicate operations that can be merged; Use temporary variables to store intermediate results and reduce the number of register reads and writes; Redesign the bit operation sequence, optimize the number of bit operations from 22 to 17 to form the reconstructed S-box substitution logic; The temporary variables include three intermediate variables (tmp0, tmp1, tmp2), which are used to store the bit operation results during the S-box substitution process, and reduce redundant calculations by reusing the intermediate variables.

3. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 2, characterized in that, Redesign the bit operation sequence, including: Decompose the non-linear operation in the original S-box substitution into a combination of linear operations; Use the bit operation instructions of RISC-V to replace multi-step logical operations; Hide the operation latency through instruction reordering.

4. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 1, characterized in that, Based on the 32-bit register architecture of RISC-V, use bit-interleaving technology to decompose the 64-bit cyclic shift operation in the reconstructed S-box substitution logic to form an optimized cyclic shift logic, including: Divide the 64-bit input into the upper 32 bits and the lower 32 bits; Perform a logical left shift operation on the upper 32 bits and a logical right shift operation on the lower 32 bits; Generate the final 64-bit cyclic shift result by performing an exclusive OR operation on the shifted high and low bits to form the optimized cyclic shift logic; The shift amounts of the logical left shift operation and the right shift operation are dynamically adjusted according to the specific requirements of the Ascon algorithm round function, and the synthesis process of the exclusive OR operation is implemented through the XOR instruction of RISC-V.

5. The method for optimizing the implementation of the Ascon algorithm based on the RISC-V processor according to claim 4, wherein Using bit-interleaving technology to decompose the 64-bit cyclic shift operation also includes: Perform bit-interleaving preprocessing on the upper 32 bits and the lower 32 bits before performing the shift operation; Ensure the alignment of the shifted data bits through mask operations; Use the shift instructions of RISC-V for displacement.

6. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 1, characterized in that Eliminate the loop control overhead of the optimized cyclic shift logic through assembly-level full expansion to form a fully expanded cyclic shift logic, including: Repeat the original loop body 12 times to cover all rounds; Eliminate the loop counter and jump instructions through inline assembly code; Independently optimize the S-box operation for each round of permutation and eliminate the function call and return instructions; Optimize the instruction order to reduce pipeline stalls; Eliminate conditional judgments through static branch prediction to form a fully expanded cyclic shift logic.

7. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 6, wherein Static branch prediction includes: Convert the conditional branch to an unconditional jump; Eliminate branch prediction failures by pre-computing the branch target address; Use the delay slot technology of RISC-V to fill the invalid instruction cycles.

8. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 1, wherein, Directly storing the intermediate state and key in the register in the fully-expanded cyclic shift logic to form the Ascon optimization algorithm, including: Allocating the message block and key state to the RISC-V general-purpose registers; Replacing array and structure accesses with direct register operations; Repeating the original loop body 12 times to cover all rounds; Eliminating the loop counter and jump instructions through inline assembly code; Independently optimizing the S-box operation for each round permutation to form the Ascon optimization algorithm.

9. The optimized implementation method of the Ascon algorithm based on the RISC-V processor according to claim 8, wherein, Storing the message block in the register in little-endian mode and exchanging data between memories through LW and SW instructions.

10. The method for optimizing the implementation of the Ascon algorithm based on the RISC-V processor according to claim 2, wherein The reconstruction of the S-box substitution logic through RISC-V assembly instructions further includes; Counting the number of clock cycles for a single S-box substitution through the RDTIME instruction of RISC-V; Calculating the deviation between the number of clock cycles and a preset reference value; When the absolute value of the deviation is greater than a preset deviation threshold, advancing the bit operation step with the longest execution time in the S-box substitution to form a corrected instruction sequence; When the deviation is less than the preset deviation threshold, delaying the bit operation step with the longest execution time in the S-box substitution to form a corrected instruction sequence; For the corrected instruction sequence, using the intermediate variable to store intermediate results to reduce the number of register reads and writes to form a first temporary instruction sequence; Reusing the intermediate variable for the first temporary instruction sequence to reduce redundant calculations to form a second temporary instruction sequence; Updating the allocation priority of the intermediate variable for the second temporary instruction sequence to form an updated instruction sequence; Allocating the registers and memories according to the life cycle and usage frequency of the intermediate variable to form an allocation instruction sequence; Inserting FENCE instructions into the allocation instruction sequence to ensure pipeline synchronization to form an optimized instruction sequence; Integrating the optimized instruction sequence into the S-box substitution logic to form the reconstructed S-box substitution logic.

Citation Information

Patent Citations

  • Techniques of additional bit freezing for polar codes with rate matching

    CN111034058A

  • Data encryption method and device based on lightweight encryption algorithm, and server

    CN118381602A