A method and apparatus for executing a round function

The described method enhances cryptographic algorithms' security against side-channel attacks by using share-based operations and masking techniques, ensuring efficient and secure execution of round functions in cryptographic applications.

WO2025229346A1PCT designated stage Publication Date: 2025-11-06PQSHIELD LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/050947
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2025-05-02
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing cryptographic algorithms like Keccak are vulnerable to side-channel attacks that exploit physical parameters such as power consumption, leading to potential leakage of internal states and compromising security.

Method used

Implementing a round function using a processing unit that generates shares of the state and applies linear and non-linear operations, including S-box operations with masking values, to protect internal states without requiring online randomness, thus reducing circuit area and latency.

Benefits of technology

The method provides first-order security against side-channel attacks while maintaining high throughput and efficiency, with reduced circuit area and latency, applicable to cryptographic applications like hash functions, message authentication codes, and block ciphers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025050947_06112025_PF_FP_ABST
    Figure GB2025050947_06112025_PF_FP_ABST
Patent Text Reader

Abstract

A method performed by a processing unit for determining a round function is provided. The round function applies one or more linear operations and applies a non-linear S- box operation implemented by a plurality of S-boxes. The method generates a first share and a second share using a current state and an input string. The method separately applies the one or more linear operations to the first share and the second share. The method further applies a plurality of functions, which include a masking value from a paired S-box, to shares of each value input to the non-linear S-box operation to generate S-box output shares. The method compresses the S-box output shares to generate shares of the processed state.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A METHOD AND APPARATUS FOR EXECUTING A ROUND FUNCTION

[0002] Technical Field

[0003] The present invention relates to a method and processing unit for executing a round function. A program for executing a round function is also provided.

[0004] Background

[0005] Side-channel attacks attempt to extract secret information from a hardware system, such as a processor. During such attacks, an attacker measures or analyses physical parameters for the system, such as supplied current, execution time, electromagnetic radiation etc. to try to obtain information from the system. For example, an attacker may attempt to measure power consumption by attaching probes to the system.

[0006] In an ideal circuit including gates, each gate exhibits at most one transition per clock cycle. However, physical factors, such as propagation delays can cause a gate to transition more than once. An unintended double transition may be referred to as a glitch. In some circuits, such as CMOS circuits, these glitches may occur. The presence of glitches affects power consumption and may lead to the leakage of information being processed by the system, or even intensify such leakage.

[0007] Summary

[0008] According to a first aspect of the present invention, there is provided a method, performed by a processing unit, of executing a round function wherein the round function includes a plurality of rounds that take a current state and generate a processed state, wherein during each round the round function applies one or more linear operations and applies a non-linear S-box operation implemented by a plurality of S- boxes, the method comprising: generating a first share and a second share using a current state and an input string, wherein each share has a plurality of values; separately applying the one or more linear operations to the first share and the second share to generate a linearly processed first share and a linearly processed second share; applying a plurality of functions to values of the linearly processed first share and linearly processed second share to apply the non-linear S-box operation, wherein: the functions each take values from the linearly processed first share and the linearly processed second share that form inputs to the S-box operation and generate a respective S-box output value share; each function defines at least a portion of the non-linear S-box operation for a combination of values of the linearly processed first share and the linearly processed second share and each function further includes at least one masking value from a paired S-box of the plurality of S-boxes which masking value cancels during a subsequent compression of the S-box output shares; and compressing the S- box output shares to generate shares of the processed state.

[0009] The shares of the processed state following compression may be used as a shares of a current state for a further round. In some implementations, the shares of the processed state may be subject to a further linear operation after compression and prior to being used as the current state for a further round. The further linear operation may take a round constant as an input. The round constant may vary with each round.

[0010] The round function may be part of a sponge algorithm. In such embodiments, the round function may form at least part of an absorbing phase.

[0011] Generating the first share and the second share may comprise generating shares of a function, x, such that x=f(a,b)=ab+b, where a is a first of a value of the current state and a value of the input string and b is the other of the value of the current state and the value of the input string.

[0012] There may be provided a method of generating at least one of: a hash; a salted hash; a mask; a message authentication code; a pseudorandom number; a cipher, a stream cipher; or a block cipher; wherein the method comprises the method of the first aspect of the invention.

[0013] Methods may comprise storing the linearly processed first share and linearly processed second share in separate storage units after the linear operations have been applied to the first share of the state and the second share of the state. In some implementations, the methods further comprise storing the S-box output shares in separate storage units after the plurality of functions have been applied.

[0014] The non-linear S-box operation may be quadratic. The non-linear S-box operation may be represented in a form f (a, b, c) = ab + b + c.

[0015] In some implementations the functions may be: where x'o, x\, x'2, and x’3are the S-box output values, ao, ai, bo, bi, co and a are shares of values a, b, and c that form inputs to the S-box operation and ko and lo are masking values from the paired S-box that cancel during a subsequent compression of the S-box outputs.

[0016] In some embodiments, the round function may be Keccak, the one or more linear operations are 9, p, and 7t, and the S-box operation is %. In other embodiments, the round function may be one of PRINCE, Midori, LED, or SKINNY.

[0017] The method may comprise combining the shares of the processed state to generate the processed state.

[0018] Following compression of the S-box output shares to generate shares of the processed state, the shares of the processed state may be used as shares of a current state for a further round by passing the shares to circuitry for applying the one or more linear operations without storing the first share and second share of the current state in a storage. The circuitry for applying the one or more linear operations may be one of: circuitry that previously applied one or more linear operations to shares of a current state, and a second circuitry that for applying the one or more linear operations, wherein applying the one or more linear operations was previously performed by a first circuitry for applying one or more linear operations.

[0019] According to a second aspect of the invention there is provided a program that, when executed by a processing unit, causes the processing unit to perform a method of determining a round function according to the first aspect.

[0020] According to a third aspect of the invention there provided a processing unit configured to execute a round function wherein the round function includes a plurality of rounds that take a current state and generate a processed state, wherein during each round the round function applies one or more linear operations and applies a non-linear S-box operation implemented by a plurality of S-boxes, the processing unit comprising: a generating unit configured to generate a first share and a second share using the current state and an input string, wherein each share has a plurality of values; at least one operation unit configured to separately apply the one or more linear operations to the first share and the second share to generate a linearly processed first share and a linearly processed second share; at least one S-box processing unit configured to apply a plurality of functions to values of the linearly processed first share and linearly processed second share to apply the non-linear S-box operation, wherein: the functions each take values from the linearly processed first share and the linearly processed second share that form inputs to the S-box operation and generate a respective S-box output value share; each function defines at least a portion of the non-linear S-box operation for a combination of values of the linearly processed first share and the linearly processed second share and each function further includes at least one masking value from a paired S-box of the plurality of S-boxes which masking value cancels during a subsequent compression of the S-box output shares; and a compressing unit configured to compress the S-box output shares to generate shares of the processed state.

[0021] Further features and advantages of the invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only, which is made with reference to the accompanying drawings.

[0022] Brief Description of the Drawings

[0023] Figure 1 is a schematic diagram showing a sponge function;

[0024] Figure 2 is a diagram showing steps in a round that applies the permutation function of the Keccak - algorithm;

[0025] Figure 3 shows a visualisation of a slice of a state;

[0026] Figure 4 is a schematic diagram of an S-box corresponding to the non-linear operation % in Keccak;

[0027] Figure 5a shows a masked AND function for generating first and second output shares from shares of input values a and b without needing additional (fresh) randomness;

[0028] Figure 5b is a table showing combinations of values that can generate the first and second shares;

[0029] Figure 6 is a schematic diagram showing pairing of S-boxes;

[0030] Figure 7a is a diagram corresponding to Figure 2 illustrating removal of a register; Figure 7b is a diagram showing steps performed by a partially unrolled circuit that applies two rounds of the permutation function of the Keccak algorithm; and

[0031] Figure 8 is a schematic diagram of components of an example information processing apparatus.

[0032] Detailed Description

[0033] An example of a processing unit for implementing a round function, exemplified by Keccak 1600, will now be described.

[0034] A round function is a transformation that is repeated (iterated) multiple times inside an algorithm. Typically, the rounds are performed using the same function, parameterized by a round constant. The round constant may be used to prevent similarity of the function between rounds that could lead to attacks on the algorithm. Round functions are used in cryptographic applications, but the techniques described below are applicable more generally and apply to any case in which a round function is to be applied and it is desired not to leak information about internal states of the round function.

[0035] While the following explanation focuses on Keccak 1600, as will be explained further below, the techniques described below are applicable to algorithms other than Keccak 1600.

[0036] Keccak

[0037] The Keccak algorithm is a family of sponge functions. Figure l is a schematic diagram showing such a sponge function. A sponge function uses an iterated approach for determining an output value. Keccak takes a variable-length input and generates a fixed length output. This makes Keccak useful for implementations such as hashing, message authentication codes, or authenticated encryption as will be described in more detail later.

[0038] The sponge function operates on a state including a fixed number of b bits, where b = r + c. Here r is referred to as the bitrate and c is referred to as the capacity. A binary input string, M, is padded to an appropriate length and split into blocks of r bits. An initial state of b bits, shown generally at 1 in Figure 1, is set with each bit value at zero. In an absorbing phase, the r-bit input blocks are XORed into the first r bits of the current state. The bits of the state are then interleaved with a function f to generate a processed state. This XORing of the current state with a new input block and interleaving with the function f is repeated until all the input string, M, has been absorbed.

[0039] During the squeezing phase, the r-bits of the state are returned as output blocks. The state may be interleaved by applications of function / to generate further output blocks until an output of a desired length is achieved. The last c bits of the state are never directly affected by the input blocks or output during the squeezing phase.

[0040] The permutation function / for Keccak is shown in Figure 2. As described above, the permutation is performed on a current state of b bits. According to variants of the Keccak algorithm, the value b may take the values 25, 50, 100, 200, 400, 800, and 1600. The description below focuses on Keccak- 1600, but the techniques described herein are applicable regardless of the size of the state.

[0041] The state can be visualised as a block of 5 X 5 X 2Zbits with a permutation length, I = 0, 1, 2, 3, 4, 5, or 6 depending on the value of b. Each 5 X 5 portion of the state is referred to as a ‘slice’ and each line of 2ldata values extending in the z direction is referred to as a lane. The 5 X 5 portion of the state may be visualised as a grid in x and j dimensions. Bits in vertical alignment are referred to as ‘columns’ and bits in horizontal alignment are referred to as ‘rows’ in the 5 X 5 slices. A visualisation of a slice is shown in Figure 3.

[0042] The number of rounds in which the function f is applied depends on the permutation length, I. The number of rounds is given by n = 12 + 21. Accordingly, for Keccak-1600, where I = 6, the number of rounds is 24.

[0043] The permutation function / is made up of five operations 9, p, 7t, %, and r. Four of the operations are linear operations and the other, %, operation is non-linear operation that mixes bits within the state. The operations will be now described in order.

[0044] The 9 function makes use of a column parity. The column parity of a column C[x,z] is defined as:

[0045] C [x, z\ = A [x, 0, z] © A [x, 1, z] © A [x, 2, z] © A [x, 3, z] © A [x, 4, z] . That is to say that the column parity is an XOR of each of the bits in the column. The 9 function also makes use of a value D determined based on two columns in the state - one in the same slice and one in a different slice:

[0046] D [x, z] = C [(% — l)mocZ 5, z] © C [(% + l)mocZ 5, (z — l)mocZ I

[0047] To determine 9, an XOR is taken such that output value A' is obtained from input value A, as follows:

[0048] A! [x, y, z] = A [x, y, z] © D [x, z]

[0049] The p function diffuses values between the slices by cyclic shifting of bits within lanes. The bits of each lane are rotated by a length, called the offset, which depends on the fixed x, y coordinates of the lane. Accordingly, for each bit in the lane the z coordinate is modified by adding the offset modulo the lane size. The transformation is:

[0050] 1. For all z such that 9<z< / , let A' [x, y, z] = A [x, y, z]

[0051] 2. Let (x, y) = (1, 9)

[0052] 3. For t from 9 to 23 a. For all z such that 9<z< / , let A' [x, y, z] = A[x,y, (z — (t + l)(t + 2) / 2) mod Z] b. Let (x, y) = (y, (2x+3y) mod 5)

[0053] 4. Return A'

[0054] The 7t function disturbs horizontal / vertical alignment of lanes within a slice. The transformation is:

[0055] A' [x, y, z] = A[(x + 3y)mod5,x, z] where 9<x<5, 9<y<5, and 9<z<Z, An S-box corresponding to the non-linear % function is illustrated in Figure 4. The % operation is performed across each row of the state. The operation is defined as follows:

[0056] 1. For all triples (x, y, z) such that 0 < x < 5, 0 < y < 5, and 0 < z < I

[0057] A' [x, y, z] = A [x, y, z] © ((A [(% + 1) mod 5, y, z] ® 1)

[0058] • A[(x + 2) mod 5,y, z])

[0059] As will be discussed further below, the structure of this S-box is that it combines three values as shown in Figure 4. Accordingly, if the values in the input row are denoted a, b, and c (c.f. x+2, x+1, x mod 5 above) then % takes a form f (a, b, c) = ab + b + c.

[0060] The • symbol indicates integer multiplication, which is equivalent to a Boolean ‘AND’ operation in this case. It is clear from the above, that the % function combines the values from three adjacent bits in a row of the state to generate the non-linear value.

[0061] The r function is parameterized by a round index ir, which counts the number of times that each of the functions has been applied (i.e., the number of rounds). The purpose of the r function is to ensure that each round is different. The method uses round constants which are as follows:

[0062] The i function operates such that, for all z such that 0<z< / ,

[0063] A'[0,0,z] = A[0,0,z] © 7?C[ir]

[0064] Accordingly, the r function modifies bits of Lane (0, 0) in a manner that depends upon the round index, ir. The other 24 lanes are not affected by the r function.

[0065] A concern to be addressed when implementing hardware or software for performing the Keccak algorithm described above is that the internal states should remain secret during processing for security. Knowledge of an internal state may, for example, make the Keccak algorithm to some extent reversible causing potential breakdown of any cryptographic or other algorithm in which it is employed.

[0066] Differential Power Analysis (P. C. Kocher, J. Jaffe, and B. Jun. Differential power analysis. In M. Wiener, editor, Advances in Cryptology - Crypto ’99, volume 1666 of Lecture Notes in Computer Science, pages 388-397. Springer, 1999.) discusses a side-channel attack that exploits dependencies between power consumption of a device and intermediate computation results. A technique to mitigate this side-channel attack is a method in which computations are performed based on shares of the sensitive input variable and the results are then combined at the end to obtain the desired result. For example, when using Boolean masking, the state may be split into a first share and a second share whose addition (binary XOR) results in the same original (unshared) value.

[0067] More generally, a sensitive variable may be split into a number of shares such that the bit-wise sum across the shares equals the sensitive variable. The shares are generated such that any subset of the shares gives no information about the sensitive value. Functions (such as round functions described herein) are computed on the shares resulting in a number of output shares. For the masking scheme to be correct, the sum of the output shares must equal the result of applying the implemented function to the sum of the input shares (i.e., the original sensitive value).

[0068] When designing a circuit to perform Keccak, there is a trade-off between latency and circuit area. Techniques for masking may involve using three or more shares, whereby the Keccak algorithm is performed for each of the three shares in parallel with low latency. However, these circuits may have high initial masking costs to generate the three shares, although they may be able to do this without the use of an online randomness, and have large internal states corresponding to the 1600-bit values for each share of the state. The need to store and process all these values tends to increase the circuit area. On the other hand, the number of shares may be reduced to two shares. However, to maintain security, the two shares may need an online randomness, which requires an increase in circuit area to generate the additional randomness. Further such implementations may need three cycles per round which reduces performance.

[0069] An embodiment will be described below that uses two shares and does not require an online randomness. Some implementations of these techniques may achieve a smaller circuit area than implementations using three or four shares while maintaining a latency of better than three cycles per round.

[0070] As described above and illustrated in Figure 2, each round of the Keccak algorithm during the absorbing phase takes the existing state (or the zero state 1 illustrated in Figure 1 for the first round) and a portion of the padded input message as input. Note that the algorithm only takes a single input, which is the state, during the squeezing phase, but during this phase the values are no longer considered sensitive because the Keccak algorithm has obscured the values of the padded input message and the states are now successively output to generate the output string as described above.

[0071] Figure 5a illustrates a function for generating first and second shares xo, xi from input shares, ao, fli. bo, and b\. The input shares are generated such that:

[0072] Where a and b are the input data values: one from the current state and the other from the block of the input string. The subscript (0 or 1) represents a mask-share domain and values that take a 0 subscript may be processed in a separate set of registers from values that take a 1 subscript.

[0073] As will be described in connection with Figure 5b below, for a given a (respectively b), there are two different possible values of ao and a\ (respectively bo and bi). The shares are generated by selecting a random value drawn from a uniform distribution. By virtue of the relationships above, it is clear that for a binary string, the values of a cannot be retrieved without having both shares ao and a\. The same applies to the other data input values bo and bi which are required to recover b in a case where there are two shares.

[0074] First and second shares to be processed through the Keccak algorithm are denoted xo and xi, respectively. The method for deriving the first and second secret share is shown in Figure 5a. As indicated in Figure 5a, the shares are generated by performing XOR (denoted +) and AND operations. The shares can be generated by the following relations:

[0075] As noted above, there are a limited number of values that a and b can take and accordingly values that the first and second shares can take. These are shown in Figure 5b.

[0076] After the first and second shares are generated, the shares xo and xi are processed in parallel through the 9, p, and it operations shown in Figure 2. These operations are linear and the output from the it operation may be referred to as linearly processed first and second shares. The depth effect in Figure 2 indicates processing through two parallel registers for these operations separated by the domain as described above. The % operator is more complicated to implement because this operation is non-linear and combines values from adjacent values in rows of the state as explained above.

[0077] The % operator is sometimes referred to as an S-box (substitution box) and is the non-linear part of the Keccak algorithm. Accordingly, x may be referred to as a nonlinear S-box operation. The x operator may be implemented in pairs of S-boxes so that randomness from one S-box can be introduced as randomness into the other and vice- versa. As will be explained further below, this allows the use of fewer register (storage) stages. In Keccak each S-box operates on a row of the state, i.e., on five values to generate 5 output values.

[0078] Figure 6 illustrates the concept of pairing S-boxes. In the figure the input vectors, x and y, are the rows of the states of processed shares which are input to each S-box of the x operator.

[0079] In order to perform the permutation operation, x, the relevant operation needs to be performed on combinations of the linearly processed first and second shares to allow the result of applying the x operation to be recovered later by XORing the S-box outputs together. As illustrated in Figure 4 and explained above, the x operation in Keccak takes inputs from three adjacent values in the row of the state.

[0080] For the purposes of generalization, the x operation can be generalized as a quadratic function in the form: f (a, b, c) = ab + b + c, where the details of the actual function are given above. Considering a simplified case (provided for explanation only), with no randomness added from the neighbouring S-box, the functions that need to be determined by the S-box are:

[0081]

[0082] As the % operation requires four functions, the S-box outputs x’o, x\, x'2, and x'3from these functions are stored in four registers in Figure 2. The first and second shares can be recovered as:

[0083] The equations above are simplified because for security the embodiments pair the S-boxes and a value from the paired S-box should be used to introduce randomness into the above functions. Accordingly, we consider the arrangement illustrated in Figure 6 where:

[0084] The values a, A, and c are the three adjacent pixels corresponding to input to the first S-box and the values k, / , and m are the three adjacent pixels corresponding to input to the second S-box. As before the subscripts 0 and 1 denote values from the first and second shares respectively. The functions above are modified to be: In this example, using Keccak, the S-box, %, is a function that takes five input bits (see Figure 4 and associated explanation above). However, each output bit only depends on three of the five input bits. The function takes a form f (a, b, c) = ab + b + c as explained above. Accordingly, the non-linear function is quadratic and breaks down into four functions as indicated above. As there are only two groups that will be eventually XORed together there is some flexibility in how to introduce randomness from the second S-box, which provides random three bits. In other words, the m bit from the paired S-box could be used in place of one of the k and I bits above. Of course, it does not matter, which randomness is added to which function as long as the randomness subsequently cancels out.

[0085] In other implementations, such as when implementing PRINCE or Midori, there may be a different number of input bits (in these examples four) and the functions may be of a different order. For example, if a target function takes a form f (a, b, c, d) = abc + be + d, there will be eight component functions due to the need to mix all the combinations of ao, fli, Z>o, Z>i, co and ci in the cubic term. These eight functions will need subsequent XORing to compress down to the desired two shares. Each pair of functions will be XORed in the first round to generate four intermediate values on the way to creating the pair of S-box output shares. In this example, the paired S-box can provide four random bits, fo, Zo, mo, and no because the function takes four input values. These bits will need to be cyclically permuted to add to a different randomness to each group of functions. In other words, the randomness values will be (ko+lo, lo+mo, mo+no, no+ko). Two random values are used so that the values are still masked after the first XOR operation compresses the output of the eight functions to four intermediate values prior to compression to the desired two shares. Accordingly, in the Keccak example above there is some choice in selection of the randomness values from the paired S- box. However, as explained, in other examples there may not be any choice.

[0086] Now that the operations in the Keccak round function have been described, a completed description of a circuit for performing Keccak can now be given with a comparison between the circuit shown in Figure 2 and the circuit in Figure 7 in which one of the registers has been removed as illustrated by the ‘X’ symbol.

[0087] A multiplexer 20 is provided at the start of a round. The multiplexer 20 is configured to select either the input state to which the function / is to be applied or the state from the previous round during the 24 rounds in which the function / is applied . As described above, the multiplexer receives two shares as described with reference to Figures 5a and 5b. After generation of the shares, they are stored in a first register 21. The shares are subsequently processed by combinatorial circuits that apply the 9, p, and Ti operations and store outputs in parallel registers. That is to say that the operations 9, p, and it are separately applied to each of the first and second share.

[0088] Circuitry for performing the 9 operation 22, p operation 23 and it operation 24 are shown in sequence. As discussed above, these operations are linear functions and no cross processing between those circuits takes place when performing these operations, i.e., they are inner-domain operations.

[0089] A second register 25 is provided prior to the % operation 26. Circuitry for performing the % operation 26 is then shown followed by four registers shown at 27 that store the S-box output shares from performing the % operation 26. As discussed above, the % operation 26 is performed using two S-boxes in which the S-boxes share randomness from the input values, i.e., no additional online randomness is needed.

[0090] A compressor 28 regenerates the state from the processed shares of the state. The state is regenerated by performing an XOR between a subset of the S-box outputs to reduce the number of shares to two. If a further round involving application of the permutation function, / is to be performed, the i function is applied by circuitry 29 to the state to generate a processed state before the processed state is returned to the multiplexer 29 for the next round. If twenty-four rounds have been completed, the i function is applied to the state by circuitry 29 and is passed to the output. The circuit shown in Figure 2 may be called several times during the absorb phase to apply the function / illustrated in Figure 1. The output of circuitry 29 can be XORed with an input string and passed to another Keccak-1699 set of round transformations (in the absorb phase), or can be output and in some cases passed to another Keccak-1699 set of round transformations (in the squeeze phase).

[0091] Turning to Figure 7a, an advantage of some implementations is that the first register 21 can be removed without loss of security. A purpose of the storage in the form of registers 21, 25, and 27 shown in Figure 2 is that glitches cannot propagate beyond those registers because the data values are stored. However, because the four function values stored in registers 27 following application of the % operation 26 are effectively masked due to the use of randomness from the paired S-boxes, glitches from the compressor 28 won’t leak information in the subsequent round transformations and the first register 21 can be removed.

[0092] The removal of the first register 21 reduces the circuit area, but this is to some extent compensated by the introduction of two XOR gates required when pairing the S- boxes. However, removal of the first register does tend to improve throughput ,i.e., reduce latency.

[0093] Accordingly, the techniques above allow the provision of a first order secure Keccak implementation which requires no fresh randomness (apart from an initial randomness discussed above to split the input) with two shares, where each round is executed within two clock cycles. A comparison of known circuits with an embodiment implementing the techniques disclosed above is set out below:

[0094] Designs 1 and 2 are described in Bilgin et al. Efficient and first-order DPA resistant implementations of KECCAK. CARDIS 2014. Design 3 is described in GroB et al. Higher-Order Side-Channel Protected Implementations of KECCAK. DSD 2017. Design 4 is described in Shahmirzadi et al. Re-Consolidating First-Order Masking Schemes. TCHES 2021.

[0095] It is recalled that Keccak 1600 requires 24 rounds. Accordingly, latency of around 24 indicates one round per cycle, 48 indicates two rounds per cycle etc. Figure 7a illustrates a single iteration of the Keccak round which is repeated 24 times as described above. In some embodiments a circuit may be ‘unrolled’ (or ‘pipelined’) so that multiple rounds of the Keccak permutation are performed at the same time. Such implementations may have increased throughput. Figure 7b is a diagram showing steps performed by a partially unrolled circuit for sequentially applying two rounds of the permutation function of the Keccak algorithm. In other embodiments, circuits for sequentially performing four, eight or some other number of rounds may be implemented. It is noted that the multiplexer 20 is only required at the initial input stage in cases where the pipeline is partially unrolled i.e. fewer rounds are implemented in hardware circuits than are required by the Keccak algorithm A fully pipelined implementation is possible and would include circuitry to implement the 24 rounds of the Keccak permutation and does not require the multiplexer 20.

[0096] The above method may be implemented in software or hardware. In hardware implementations, a processing unit includes circuitry for performing each of the operations 0, p, 7t, %, and r as described above. The circuitry may be separated by storage in the form of registers as illustrated in Figures 7a and 7b and described above.

[0097] Implementation of the method described above by software is also possible. Figure 8 is a schematic diagram of components of an example information processing apparatus suitable for use in performing the method described above. The diagram is illustrative and different hardware configurations for information processing apparatus are possible as is well known in the art. The information processing apparatus includes an I / O interface 80, such a USB port, Thunderbolt port, etc. to which an additional device, such as a storage device, could be connected. The information processing apparatus comprises a processing unit in the form of a processor 81, such as a CPU, GPU or TPU, a storage in the form of memory 82, a network module 83, a display 84, and a user interface 85. The network module may allow the information processing apparatus to communicate over a network such as a Wi-Fi network, a mobile telecommunications network, a local area network etc. The user interface may include components such as a keyboard, mouse, camera, etc. The components of the information processing apparatus may communicate with each other over a bus 86. Further components may be provided but are not shown or described. Any of the steps of the methods described above may be performed by computer-readable instructions of one or more programs stored in a storage and executed by a processing unit on one or more information processing apparatuses.

[0098] The above embodiments are to be understood as illustrative examples. Further embodiments are envisaged.

[0099] Other algorithms

[0100] It will be apparent that the techniques described above can be applied regardless of the size of the state in the Keccak or any other round function. Accordingly, the methods described can be applied regardless of the state size.

[0101] The techniques are also applicable to other round functions, such as PRINCE, Midori, LED, or SKINNY. These round functions will have different S-boxes. In the example above, an S-box expressible in the form / (a, b, c) = ab + b + c was discussed. However, the techniques above are applicable to different types of non-linear functions / different S-boxes. For example, an S-box may take four input values (as opposed to three for Keccak). In such an example, the S-box may be representable as f (a, b, c, d) = abc + be + d. Such an S-box is cubic and will require 8 functions to represent the possible combinations of two shares that will allow a processed state to be recovered during compression. Each of the functions will define at least a portion of the non-linear S-box operation for a combination of values of the linearly processed first share and second share. Each function further includes at least one masking value from a paired S-box of the plurality of S-boxes which masking value cancels during a subsequent compression of the S-box outputs in which the processed state is generated. Accordingly, based on the teaching above, adapting the techniques described to different S-boxes and identifying different functions that XOR (or otherwise combine) together to result in the desired non-linear S-Box function is a routine matter that the skilled person can perform without undue burden.

[0102] The techniques above implement Boolean masking in which the first and second shares can be recombined to obtain the processed state by performing a logical XOR on the two shares. However, the technique described above is equally applicable to arithmetic masking. In such an implementation, the randomness from the neighbouring S-box in the pair is added to the function being implemented by the S-box. The method described above may be first-order secure, as indicated in the above table comparing circuit implementations. This means that the method is secure if an attacker can only see a single stream of variables. Second-order security may be obtained by introducing fresh randomness into the pairs of S-boxes described above. Accordingly, the absence of fresh randomness for masking in the pairs of S-boxes, while a desirable property in terms of simplifying circuit design, is not essential and further benefits can be obtained by introducing randomness into the functions that implement the S-box.

[0103] Round functions that may use two shares and the paired S-boxes as described above are useful in many different applications.

[0104] Some round functions, including Keccak, may be used in hash functions. Such hash functions may take a message to be hashed as the input message (with suitable padding as necessary) and the output after squeezing is the hash value. The hash value is unique to the input message and small changes in the input message will result in large changes to the output message. Further due to the non-linear nature of the S-box, reversing the hash to determine the input message is presumed to be very difficult. A salted hash may be easily generated by adding a salt value to the beginning of the padded message prior to hashing.

[0105] Round functions, including Keccak, may be used as mask generation functions. For example, it may be desirable to derive a random output of variable length. In such cases, by repeated application of the permutation function, / in the squeezing phase a variable length output may be obtained. Such functions have use, for example, in key derivation functions, signature schemes, and key encapsulation methods.

[0106] Round functions, including Keccak may be used for generating message authentication codes. The state may be initialized with a key and the message (with suitable padding as necessary) may be processed by the round function to generate a message authentication code, e.g., KMAC.

[0107] Round functions, including Keccak, may be used as a reseedable random number generator in which values (seeds) may be periodically input. The state after one or more applications of the permutation function may be used as the pseudorandom string. Round functions, such as PRINCE, Midori, LED, or SKINNY, may be used as block ciphers that take a key to initialize the state and a message and generate a symmetric cipher that can be decrypted by a party that knows the key.

[0108] Round functions may also be used in stream ciphers. For example, round functions used in a sponge algorithm may be ‘squeezed’ repeatedly to generate a key stream. The key stream may be XORed with plaintext or ciphertext to generate the stream cipher. Performing a second XOR with the key stream allows the stream cipher to be decrypted.

[0109] In some implementations, the message or input value in the round function may be supplemented with linear fault countermeasures. For example, error correcting codes could be added to the message (such as in a symmetric cipher) to improve stability in case of faults during encryption and decryption.

[0110] It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

CLAIMS1. A method, performed by a processing unit, of executing a round function wherein the round function includes a plurality of rounds that take a current state and generate a processed state, wherein during each round the round function applies one or more linear operations and applies a non-linear S-box operation implemented by a plurality of S-boxes, the method comprising: generating a first share and a second share using a current state and an input string, wherein each share has a plurality of values; separately applying the one or more linear operations to the first share and the second share to generate a linearly processed first share and a linearly processed second share; applying a plurality of functions to values of the linearly processed first share and linearly processed second share to apply the non-linear S-box operation, wherein: the functions each take values from the linearly processed first share and the linearly processed second share that form inputs to the S-box operation and generate a respective S-box output value share; each function defines at least a portion of the non-linear S-box operation for a combination of values of the linearly processed first share and the linearly processed second share and each function further includes at least one masking value from a paired S-box of the plurality of S-boxes which masking value cancels during a subsequent compression of the S-box output shares; and compressing the S-box output shares to generate shares of the processed state.

2. A method according to claim 1 wherein the shares of the processed state following compression are used as shares of a current state for a further round.

3. A method according to claim 2, wherein the shares of the processed state are subjected to a further linear operation after the compression and prior to being used as the current state for a further round.

4. A method according to any preceding claim wherein the round function is a part of a sponge algorithm and the round function forms at least part of an absorbing phase.

5. A method according to any preceding claim wherein generating the first share and the second share comprise generating shares of a function, x, such that x = / (a, 6) = ab + b, where a is a first of a value of the current state and a value of the input string and b is the other of the value of the current state and the value of the input string.

6. A method of generating at least one of a hash; a salted hash; a mask; a message authentication code; a pseudorandom number; a cipher; a stream cipher and a block cipher; comprising the method of any preceding claim.

7. A method according to any preceding claim further comprising storing the linearly processed first share and the linearly processed second share in separate storage units after the linear operations have been applied to the first share of the state and the second share of the state.

8. A method according to any preceding claim, further comprising storing the S- box output shares in separate storage units after the plurality of functions have been applied.

9. A method according to any preceding claim wherein the non-linear S-Box operation is quadratic and may be represented in a form f (a, b, c) = ab + b + c.

10. A method according to claim 9, wherein the functions are:where x'o, x\, x'2, and x'3are the S-box output values, ao, ai, bo, bi, co and a are shares of values a, b, and c that form inputs to the S-box operation and ko and lo are masking values from the paired S-box that cancel during a subsequent compression of the S-box outputs.

11. A method according to any preceding claim wherein the round function is Keccak, the one or more linear operations are 9, p, and 7t, and the S-box operation is %.

12. A method according to any of claims 1 to 8, wherein the round function is one of Keccak, PRINCE, Midori, LED, or SKINNY.

13. A method according to any preceding claim further comprising combining the shares of the processed state to generate the processed state.

14. A method according to any preceding claim, wherein following compression of the S-box output shares to generate shares of the processed state, the shares of the processed state are used as shares of a current state for a further round by passing the shares to circuitry for applying the one or more linear operations without storing the first share and second share of the current state in a storage.

15. A method according to claim 14, wherein the circuitry for applying the one or more linear operations is one of: circuitry that previously applied one or more linear operations to shares of a current state, anda second circuitry that for applying the one or more linear operations, wherein applying the one or more linear operations was previously performed by a first circuitry for applying one or more linear operations.

16. A program that, when executed by a processing unit, causes the processing unit to perform a method of determining a round function according to any preceding claim.

17. A processing unit configured to execute a round function wherein the round function includes a plurality of rounds that take a current state and generate a processed state, wherein during each round the round function applies one or more linear operations and applies a non-linear S-box operation implemented by a plurality of S- boxes, the processing unit comprising: a generating unit configured to generate a first share and a second share using the current state and an input string, wherein each share has a plurality of values; at least one operation unit configured to separately apply the one or more linear operations to the first share and the second share to generate a linearly processed first share and a linearly processed second share; at least one S-box processing unit configured to apply a plurality of functions to values of the linearly processed first share and linearly processed second share to apply the non-linear S-box operation, wherein: the functions each take values from the linearly processed first share and the linearly processed second share that form inputs to the S-box operation and generate a respective S-box output value share; each functions defines at least a portion of the non-linear S-box operation for a combination of values of the linearly processed first share and the linearly processed second share and each function further includes at least one masking value from a paired S-box of the plurality of S-boxes which masking value cancels during a subsequent compression of the S-box output shares; and a compressing unit configured to compress the S-box output shares to generate shares of the processed state.