Method and circuit arrangement for constructing a versatile hash function based on a turbo encoder

By constructing a multi-functional hash function based on the overall structural design of the turbo encoder, the shortcomings of existing technologies in resisting collision attacks, preimage attacks, and length attacks are solved, achieving efficient multi-functional support and low-cost security improvement.

CN122394764APending Publication Date: 2026-07-14BEIJING RED & BLUE TREE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING RED & BLUE TREE TECHNOLOGY CO LTD
Filing Date
2026-03-11
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing hash functions are insufficient in resisting collision attacks, preimage attacks, and length attacks, and it is difficult to achieve efficient support for multi-functional metrics such as hash/xof, mac, and duplex mode.

Method used

An overall structural design based on a turbo encoder is adopted. A multifunctional hash function is constructed by using a turbo encoder and a nonlinear logic iterator. The turbo encoder isolates the plaintext block from direct modification, ensuring that the security capacity cl of the controlled unit is not directly modified by the input and is not spied on by the output. Combined with block cipher mode or strong permutation mode, it realizes image attack and collision attack.

Benefits of technology

While ensuring security and performance, the design and implementation costs of nonlinear logic have been reduced, and the anti-image attack, collision attack and length attack capabilities similar to those of sponge structures have been achieved. At the same time, it supports multi-functional indicators and is superior to existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122394764A_ABST
    Figure CN122394764A_ABST
Patent Text Reader

Abstract

The application discloses a multi-functional hash function construction method and circuit device based on a turbo encoder, and proposes a new security strategy and a new overall structure for ICCS new commercial secret hash standard collection.The overall structure is as follows: a master control unit is a turbo encoder, a controlled unit is a nonlinear logic iterator, the controlled unit works in a block cipher mode or a strong permutation mode, and the c.r of the controlled unit is preferably outputted, and the c.r is outputted and updated by the master control unit, and the plaintext block updates the state of the master control unit.The application point is that the c.l of the controlled unit is not directly modified and directly outputted, and the security capacity c.l is used to resist the original image attack and the collision attack.The technical effect is that the maximum code rate + the anti-original image / state capacity is an optimal design, the sponge structure is enabled to increase the code rate, a large state capacity large input large output overall structure is constructed by supporting a plurality of small state capacity overall structures, the controlled units are basically the same, so that one circuit supports a combination mode / decomposition mode, and one set of code covers a family of algorithms.The multi-function is hash / xof, mac or duplex mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer symmetric cryptography design and its industrial applications, and more specifically to a method for constructing multifunctional hash functions based on a turbo encoder, a method for constructing large-state large-output hash functions from small-state small-output hash functions, and a circuit device. Background Technology

[0002] On February 5, 2025, ICCS (Commercial Cryptography Standards Institute) announced a global call for submissions of next-generation commercial cryptography standards. Hash algorithms are one of the three tracks. Submitted hash algorithms must include at least 512-bit and 1024-bit outputs. ICCS expects theoretical and structural innovations and requires self-evaluation, anti-image attack, anti-collision attack, and anti-length attack capabilities.

[0003] NIST (National Institute of Standards and Technology) launched a global competition for the SHA3 standard and the CAESAR competition. The winners were keccak and ASCON, both using the sponge structure. This multi-functional indicator demonstrates that cryptographic hashes or authentication ciphers designed based on the sponge structure can operate in hash / xof, MAC (extended to authentication functions), or full-duplex mode (absorbing and outputting random numbers simultaneously). A significant original work is titled "Permutation-based encryption, authentication, and authenticated encryption."

[0004] The consensus reached by experts at the roundtable forum of the 2025 Xi'an Cryptographic Chip Conference was that "the sponge-based design has won a major victory in the standards competition, and since 2012, no new overall structure has challenged the sponge structure."

[0005] Other overall structures: The MD structure has the advantage of being a standard first-mover, but it has many inherent disadvantages. Typical MD structure standards include SHA2 and SM3. The sponge structure with feedback (feedforward mechanism) was proposed by Sun Siwei (application number CN2025106346957) and Guo Chun. Its main disadvantage is that it does not preserve entropy and is not a reversible permutation compared to the sponge structure. Summary of the Invention

[0006] This instruction manual reading guide states: Items prefixed with ☆☆ represent the highest level of summary, mainly summarizing the technical problems, inventive points, inventive concepts, and core technical effects of the invention; ☆ prefixes and serial numbers indicate the content of the invention.

[0007] In order to invent a new overall structure and construct a new generation of multifunctional cryptographic hash functions, we first studied the sponge structure, MD structure and sponge structure with feedback in the background technology, and extracted the following 5 points as the technical problems to be solved by this invention.

[0008] ##1. The new overall structure must be based on a new security assumption model. For example, the security principle of the MD structure is that if the compression logic is collision-resistant, then the overall structure is collision-resistant. The security principle of the sponge structure is based on the strong permutation assumption and the concept of security capacity. The security capacity c is the invention point, which cannot be directly modified or viewed. It is easy to prove that the ability to resist preimage attacks and collision attacks reaches the birth limit. The total register consumption is r+c, where r is the input code rate.

[0009] ##2. The register consumption metrics of the new hash function theory should be benchmarked against the sponge structure and its implementations. The sponge structure consumes the fewest registers in hash / xof-duplex mode, which is superior to the MD structure.

[0010] ##3. The multi-functionality of the new hash function should be comparable to SHA3 and ASCON. The standard solicitation target mentioned by ICCS is that it should ideally support hash / xof, MAC, or full-duplex mode. Simply put, a single set of code or circuitry should implement hash, authentication, and random number generation functions.

[0011] ##4. Design the optimal embodiment of the overall structure of the new invention, whose performance should be comparable to that of high-performance SHA3 and ASCON.

[0012] ##5. The new overall structure or its embodiments preferably have new characteristics and effects that are not present in the prior art.

[0013] The security assumptions of the new overall structure are mainly inspired by two technologies: Inspiration 1: the zero-forcing design concept of turbo convolutional code encoders in the field of communications; Inspiration 2: the security capacity c of the sponge structure cannot be tampered with by the input, cannot be converged, cannot be predicted, and the output cannot be viewed. The design concept and security assumptions of this invention mainly have two points.

[0014] The preimage attack, collision attack, and second preimage attack of the hash function are the struggle between forcing the state difference of {cl,cr,r} to zero and resisting the forcing of the state difference to zero.

[0015] Our designer's position is to consume the attacker's available bit degrees of freedom in the process of resisting state difference forced to 0 or state convergence, and it is best to use up all of the attacker's optional bit degrees of freedom.

[0016] Based on the above understanding, the safety layout of the new overall structure has three aspects.

[0017] 1. First, the turbo encoder component is used to isolate plaintext blocks from directly modifying {cl and cr}. Simply put, the turbo encoder consumes some of the attacker's degrees of freedom with very low engineering cost. 2.cl represents the safety capacity, which is equivalent to the safety capacity design concept of a sponge structure; 3. The most economical attack route is to force the state difference of the turbo convolutional encoder to zero. Once the state difference is forced to zero, such as the state difference of a master control unit in the intermediate encounter attack mode, it will fall into the "trap" designed in this invention. Assuming that P is a strong random permutation mode or a block cipher mode with better effect, it is easy to prove that the cost of escaping the "trap" is no less than the birthday security boundary specified by the hash competition. The essence of the "trap" is that the attacker cannot control and predict cl.

[0018] To better understand this invention, we will first state the inventive concept ☆☆1 and inventive point ☆☆2, and then correct and supplement the explanation of the two safety design concepts of the new overall structure and the anti-collision principle of anti-image attack.

[0019] ☆☆1 Invention Concept: No new overall structure: The main control unit is a turbo encoder with a code rate of r, and the controlled unit is a nonlinear logic iterator. The controlled unit operates in block cipher mode or strong permutation mode. The output of this invention is preferably the cr of the controlled unit. The control relationship is that the plaintext block first updates the state of the main control unit, and then the states of cr and cl are updated by the output of the main control unit.

[0020] The second inventive point is that the controlled unit's safety capacity cl is not directly modified by the input and is not spied on by the output. cl is used to resist preimage attacks and collision attacks. More specifically, the safety capacity c.l + single output bit is no greater than the capacity of the main control unit. On the other hand, the turbo encoder prevents plaintext blocks from directly modifying the controlled unit; at least two plaintext blocks are needed to force the state difference of the main control unit to 0. The linkage between the main control and the controlled unit achieves a technical effect comparable to that of the sponge structure. Although the safety capacity cl of this invention is not as large as the safety capacity c of the sponge structure, the overall technical effect of c.l + r is comparable to that of the sponge structure. In other words, by limiting the total register capacity, the input bit rate + output bit rate / state register capacity index, the new overall structure is as good as the sponge structure, thus achieving the technical problem ##2.

[0021] ☆☆3 Summary of Invention Effects 1: Under the constraints of ##2, the nonlinear iterative logic of this invention is only the controlled unit, while the sponge structure represents the full capacity. Therefore, it is friendly to both implementation design and engineering implementation. For example, in the duplex mode of the kaccake1600, the maximum bit rate reference value is 512 bits (1600 / 3=533.3). In the duplex mode of the optimal embodiment A circuit of this invention, the total capacity of circuit A is 512*3 bits, of which the turbo encoder capacity is 512 bits and the controlled unit capacity is 1024 bits. Obviously, the cost of nonlinear logic design and implementation is significantly lower than that of kaccake. ##5.1 is derived: under the same capacity and bit rate, the nonlinear logic width of this invention is smaller than that of the sponge structure.

[0022] ☆☆3 Summary of Invention Effects 2: Similar to the sponge structure, the overall structure of this invention can resist preimage attacks, collision attacks, and length attacks, etc. More detailed information will be provided in the specific implementation section. Main conclusions: 1. The principle of resisting preimage attacks is that the intermediate encounter attack forces the state difference of the main control unit to zero. The side effect of forcing zero is that cl is uncontrollable. The collision cost of the cl region guarantees the safety boundary. 2. The principle of resisting collision attacks is that cl is unviewable and unmodifiable. 3. Similar to the sponge structure, because of the existence of a safety capacity that is unviewable and unmodifiable, it resists length attacks.

[0023] ☆☆3 Summary of Invention Effects 3: Directly empowers multifunctional cryptographic hash functions based on sponge structures (such as industry standards). By using the strong permutation module of the sponge structure as the main control unit, multifunctional cryptographic hash functions with higher code rates and / or higher security against image attacks can be constructed with almost no increase in software and hardware overhead. The Keccak1600-24 rounds are used as the main control unit, and the main control unit capacity can be selected from 768, 1024, and 1536.

[0024] Hash mode input bitrate output bitrate encounter attack index birthday boundary Turbo1024 ->keccak1600 1024 512 800 512 Turbo1024 ->keccak1600 1024 768 800 768 Turbo1024 ->keccak1600 1024 512+512xof 800 1024 ☆1. A method for constructing a multifunctional hash function based on a turbo encoder, wherein the function includes an initialization phase, an absorption phase, and an output phase. Its features include the following overall structure: the main control unit is a turbo (convolutional code) encoder, and the controlled unit is a nonlinear logic iterator. The main function of the nonlinear logic iterator is to obtain the block cipher mode effect or strong permutation mode effect through multiple sub-wheel iterations. The control relationship is as follows: the plaintext block drives the movement of the main control unit, and the output of the main control unit drives the movement of the controlled unit with linear operations (preferably the modulo-2 addition or fixed-point addition operation with the lowest implementation cost). State capacity: The number of bits in the main control unit register is r, where r is referred to as the (maximum) input code rate. The number of bits in the controlled unit register is cr and cl. cr and r are equal, where the control relationship between r and cr is the linear operation described above; cl is the safety capacity, used to resist preimage attacks and collision attacks; Absorption phase: The plaintext block controls the movement of the master control unit, and the synchronous output of the master control unit controls the movement of the controlled unit; Output stage: Extract {cl, cr} as output or compress {cl, cr} with main control unit information for output, prioritizing the extraction of bits that do not belong to the safe capacity.

[0025] Similar to the sponge structure, truncating {cl, cr} as the output is recommended. This is the simplest form with an engineering cost of zero, because cl represents the safe capacity, so bits not belonging to the safe capacity are truncated first. Increasing computational logic for higher safety is also a common design philosophy, such as compressing the contents of the master and controlled unit registers for output. In short, the design of the output stage is open.

[0026] The invention relates to a circuit-based algorithm modeled after a sponge structure. The algorithm has a fixed input bitrate (r + (c.r + cl)). r + (c.r + cl) is equivalent to the r + c index of the sponge structure, where r is the maximum supported input bitrate. The actual bitrate can be less than r. Obviously, the smaller the actual bitrate, the larger the corresponding security capacity, meaning a better anti-image attack index. Similarly, this applies to output bitrate selection. It is recommended to use cr directly as the output. If security is prioritized, choose an output bitrate less than cr; if bitrate is prioritized, choose an output bitrate greater than cr. In short, like a sponge structure, the output must reserve security capacity. For the issue of interchangeable input / output bitrate and security, please refer to the Turbo->keccak series main control unit implementation examples.

[0027] Additional notes on the initialization phase: This is a necessary step using standard methods. Simply put, it involves initializing the contents of the capacity register, referring to the OpenSSL hash function initialization function and the ukey technical standards SKF_DigestInit and SKF_MacInit.

[0028] Additional explanation for the output phase: Referring to ASCON and SHA3 (kaccake1600) based on the sponge structure, besides outputting the hash, like the sponge structure, continuous empty runs are XOF mode or random number encryption mode. More importantly, similar to the ASCON duplex state, it is equivalent to merging the "absorption phase:" and "output phase:" into one round to achieve absorption and output (the sponge structure calls this "squeezing"). That is, it absorbs the latest plaintext block while outputting the new random number. The duplex mode has MAC code functionality. In short, it can solve technical problem ##3.

[0029] Supplementary explanation during the absorption phase: Plaintext blocks correspond to different entities in different modes. In hash mode, it is a plaintext message; in duplex mode (corresponding to ASCON), the plaintext blocks are arranged in the following iteration order: IV -> Key -> Plaintext. The essential technical features of this invention are strongly correlated with its technical effects, corresponding to the maximum input code rate and the capacity of the main control unit.

[0030] Further explanation of the design concept: 1) The design purpose of cl is to directly benchmark the security capacity C of ascon and SHA3, which is an unobservable and untamperable security entropy pool.

[0031] 2) The P-permutation directly corresponds to the ideal reversible random permutation of ASCON and SHA3. 3) In engineering, P-permutations are often constructed using SP structures with multiple iterations. As long as there are enough differential active S-boxes, the safety is guaranteed in practice.

[0032] ☆2. According to the multifunctional hash function construction method described in ☆1, the working mode of the controlled unit is block cipher mode, not strong permutation mode. {cl, cr} is truncated as output. Note that the bits used for output are not part of the secure capacity. The specific steps are as follows: In the first sub-wheel, the plaintext block drives the turbo encoder to move; in the other sub-wheels, the turbo encoder runs idle, that is, the plaintext block input is turned off. {cl,cr} is equivalent to the plaintext of the block cipher, and the output of the turbo encoder is used as the block cipher wheel key; The turbo encoder encrypts one sub-wheel with a block cipher for each sub-wheel that moves.

[0033] Below is the preferred block cipher operating mode. The rationale for this preference is that, through multiple sub-round iterations of the block cipher, the security capacity of the controlled unit increases from cl to c({cl,cr}), and the maximum input code rate increases from r to c. The example of turbo-keccak demonstrates that its input code rate and / or output code rate are superior to keccak, and its 512-bit security throughput is also superior. Compared to engineering implementation, the block cipher mode only requires an additional empty module "Ti+2=Turbo(Ti+1,PAR_X)" compared to the strong permutation mode, which can be considered as having no hardware overhead and very little software overhead.

[0034] Ti+1=Turbo(Mi⊕Ti, PAR_X); {ci+1.l,ci+1.r}=Enc_1round({ci.l,ci.r},Ti+1,*); Ti+2=Turbo(Ti+1, PAR_X); / / Idle mode {ci+2.l,ci+2.r}=Enc_1round({ci+1.l,ci+1.r},Ti+2,*); ... Ti+t=Turbo(Ti+t-1, PAR_X); {ci+tl,ci+tr}=Enc_1round({ci+t-1.l,ci+t-1.r},Ti+t,*); Below is the strong permutation mode, where t is equivalent to ASCON's 12 sub-wheels and keccak's 24 sub-wheels. Strong permutation is the earliest designed mode with significant theoretical value. It can be considered a degenerate version of the Enc_1round round-logic block cipher mode, where the wheel keys are not updated.

[0035] Tj+1=Turbo(Mj⊕Tj, PAR_X); {ci+1.l,ci+1.r}=Enc_1round({ci.l,ci.r+Tj+1}, *); {ci+2.l,ci+2.r}=Enc_1round({ci+1.l,ci+1.r},Ti+2,*); ... {ci+tl,ci+tr}=Enc_1round({ci+t-1.l,ci+t-1.r},Ti+t,*).

[0036] For design implementation, bitrate and security interchangeability, and actual performance testing, please refer to the Turbo->keccak series main control unit example.

[0037] ☆3. Based on the multifunctional hash function construction method described in ☆2, cr and cl are equal The bit width of cr is 128, 256, 384, 512 or a multiple of 512.

[0038] ☆3 Design objective 1: In duplex mode, cl can ensure that the output bit rate meets the target threshold for image attack. The optimal implementation circuit A has an input bit rate of 512 bits, an output bit rate of 512 bits, and cl is 512 bits.

[0039] ☆3 Design Purpose 2: When the width of the controlled unit {cl,cr} is a multiple of 256, 384, 512 or 1024, the nonlinear logic iterator is particularly suitable for AVX, AVX512 instruction set design and instruction optimization implementation.

[0040] The core operation of the main control unit is Ti+1=Turbo(Mi⊕Ti, PAR_X), etc.; the most economical industrial design for Turbo() is circular shift, especially byte circular shift (at which point the 8-bit machine and avx2 code quality are highest); considering that the hash function constructed in this invention also needs to be compatible with xof mode and duplex mode, it is expected that the continuous idle state of Turbo() is a large cycle, and Turbo() is preferably an m-sequence recursion. Research found that the design of ax=ax byte circular shift i⊕(2*ax&MASK) (or ax=ax byte circular shift i⊕(ax / 2&MASK)) has excellent performance for both 8-bit machines and avx2 code, and Turbo() can be preferably selected as an m-sequence recursion MASK through computer programming. Compared to the linear recursive modules of SONN-3g, ZUC, SONN-5g, and LOTO, the SIMD code overhead of the optimal implementation ☆4 is smaller for each new state generated; specifically, each state update requires 4 avx512 instructions to achieve a full update of 4 128-bit states.

[0041] ###043 ☆4. Based on the multifunctional hash function construction method described in ☆3, The main control unit is defined as follows: ax = ax bytes circular shift i ⊕ (2 * ax & MASK), where ax is the master control unit register; It should be noted that only the hash function is considered, while the XO function is ignored. Security is sacrificed for maximum performance, and a MASK value of 0 is allowed.

[0042] Below is the actual parameter table of the main control units X, Y, Z, and W of circuit A, and the implementation of 4 AVX512 instructions. The main control units generate 4 sets of polynomial m sequences with different characteristics during idle operation. The inputs of the main control units are the same. Because the cyclic shift bits are coprime, the cost of forcing 0 to 0 for all 4 main control units simultaneously is the maximum.

[0043] const uint64_t PAR_X_1[8]={ 0x32922002011fdf7eL,0x0000000000000000L,0x32a60c2a0041b37eL,0x0000000000000000L, 0x404a040a00fbff7eL,0x0000000000000000L,0x2ed2102a00434f96L,0x0000000000000000L, }; const uint64_t PAR_X_2[8]={ 0x0807060504030201L,0x000f0e0d0c0b0a09L,0x0a09080706050403L,0x0201000f0e0d0c0bL, 0x0c0b0a0908070605L,0x04030201000f0e0dL,0x0e0d0c0b0a090807L,0x060504030201000fL, }; __asm("vpshufb PAR_X_1,%zmm26, %zmm6"); __asm("vpaddq %zmm26,%zmm26,%zmm26"); __asm("vandpd PAR_X_2,%zmm26,%zmm26"); __asm("vxorpd %zmm6,%zmm26,%zmm26"); The consensus and conventional approach in the field of block ciphers and cryptographic hash design is to construct a theoretically strong permutation or a key-controlled strong permutation through multiple iterations of round transformations with weak diffusion and confusion. This strong permutation is one that cannot be distinguished from a randomly selected substitution table. The controlled unit design inherits this consensus and provides a recursive calculation formula. The core logic of the main control unit can be described as follows: {ci+1.l,ci+1.r}=Enc_1round({ci.l,ci.r},Ti+1,*). Enc_1round can directly adopt the design results of ASCON or Keccak, or refer to existing block cipher designs, such as the SP structure, Festl structure, and ARX components. More embodiments and technical summaries are provided in the specific implementation section.

[0044] The SP structure exhibits superior low-latency performance and parallel optimization, significantly outperforming the Festel and Lai-mass structures. Therefore, it is currently the mainstream design choice. The inventors found that the SP structures of ASCON or Keccak can be applied to the controlled cells of this invention without any problems. However, the minimum number of S-boxes activated in the first round is 1. It would be better if more S-boxes could be activated in the first round. The inventors researched the EWES block cipher key diffusion algorithm. The main idea is to use the core logic of the odd-numbered rounds of EWES as the core logic of EWES key diffusion, which can achieve an extremely high number of differential branches. The specific recursive method is "S_odd(t) -> bit permutation -> S_odd(t) -> bit permutation S_odd(t) -> bit permutation…", where S_odd(t) = P -1 (X(P(t))), for detailed definition please refer to the invention patent.

[0045] The inventors discovered that by setting t equal to {cl, cr}, Enc_1round activates at least 6 S-boxes, and the output of Enc_1round is associated with all S-box outputs. In other words, the output of Enc_1round can cover all activated S-box outputs; this design performance is clearly superior to the SP and PS structures. This result can be easily generalized to a nonlinear component: P1->S-box layer->P2, where P1 and P2 are linear diffusion layers. P1 contributes that the input from cr to the S-box layer is linear encoding; information inputs to cr, and the encoded output is the S-box layer, with a higher minimum code weight being preferable. P2 contributes that the cr portion of Enc_1round ideally associates with all S-box layer outputs. In summary, P1->S-box layer->P2 is a design with better diffusion performance than the SP or PS structures.

[0046] ☆5. Based on the multifunctional hash function construction method described in ☆3, The controlled unit is defined as follows: The logic transformation of the controlled unit sub-wheel is a nonlinear component P1->S box layer->P2, where P1 and P2 are linear diffusion layers.

[0047] ☆6. Based on the multifunctional hash function construction method described in ☆5, Its characteristic is that the P1->S box layer->P2 is The odd-round logic of the block cipher EWES is a bit permutation, where the bit permutation is {l,r}->{r,l<<<15}, where <<< is a 64-bit left circular shift, denoted as ADJUSTEMENT_BIT({l,r})={r,l<<<15}; The EWES odd-round logic definition refers to claim 7 of the invention patent "A Self-Reversible Block Cipher" (application number 2024114403213); an excerpt is as follows: S_odd(t) = P -1 (X(P(t))), linear diffusion component P, defined based on 64-bit modulo-2 addition and circular shift instructions, input {l,r}, the intermediate steps of the calculation are: {l⊕r,r} ->{l,r<<<14} ->{l,r⊕l} ->{l<<<44,r} ->{l⊕r,r} ->{l,r<<<8} ->{l,r⊕l} ->{l<<<24,r} ->{l⊕r,r}.

[0048] ☆6's main technical effects: 1) The program demonstrates that the component's diffusion efficiency is extremely high. The reason for choosing 15 is to maximize the number of active S-boxes in the S-box layer, and the linear diffusion capability approximates the effect of the MDS matrix. Only cr can be modified. The evaluation results below prove that the diffusion effect is superior to the SP structure design. (Removing additional data...) Figure 3 The two modulo-2 addition logics only calculate the lower bound of the differential S-box from rounds 1 to 4.

[0049] 6-->14-->20-->28 Lower bounds of the 1st to 4th round differential S-boxes between S-box layers: 1-->13-->16(15)-->27 2) Because the nonlinear logic is equivalent to the odd-round logic of the block cipher EWES, the block ciphers EWES and ☆6 share the same software code and hardware code or physical hardware circuit (FPGA / ASIC), which is beneficial for engineering implementation.

[0050] ☆7. According to the multifunctional hash function construction method described in ☆6, the sub-wheel of the block cipher mode further includes count and RC parameters, where count is a counter for the number of times the controlled unit moves, and RC is a wheel constant.

[0051] ☆7 Purpose and technical effects of the invention: 1) Count can better resist encounter-in-the-middle attacks, which are only effective for equal-length attack models; 2) RC is a mode degeneration prevention mechanism for block ciphers. When multiple small-state-capacity hash functions are linked together to design a large-state-capacity hash function, the purpose of RC is to construct different block ciphers.

[0052] The following describes a method for constructing a large-state-capacity multi-function hash function from a small-state-capacity multi-function hash function. Both the small-state-capacity and large-state-capacity multi-function hash functions conform to the overall structure of this invention, i.e., they both conform to the definitions in ☆1-3 of this invention. ☆☆4 The inventive point is that the main control units of all small-state-capacity multi-function hash functions are basically the same, which is inherently friendly for clock-to-area engineering implementation. More importantly, the characteristic polynomials of the turbo encoder are preferably coprime, which results in the strong permutation or block ciphers generated by the Enc_1round corresponding to each small-state-capacity multi-function hash function after multiple iterations of the sub-rounds being mutually independent logical transformations. Even if the input plaintext blocks of all small-state-capacity multi-function hash functions are the same, the {cl,cr} of each function are also mutually independent.

[0053] ☆8. A method for constructing a large-capacity, multi-functional hash function, wherein the function includes an initialization phase, an absorption phase, and an output phase. The technical approach is to first define multiple small-state-capacity multi-functional hash functions, and then construct a large-state-capacity multi-functional hash function through circuit integration and clock synchronization. The small-state-capacity hash function conforms to the definition of ☆1-7. Wherein, the input of the large-state capacity hash function is equal to the bit concatenation of the input of the small-state capacity hash function. The output of the large-state capacity hash function is equal to the bit concatenation of the outputs of the small-state capacity hash function. The safety capacity cl of the large state capacity hash function is equal to the sum of the small state capacity hash functions cl; In order to reuse the nonlinear computation logic of the controlled unit, the controlled units of all small-state-capacity multifunctional hash functions are basically the same; To resist the zero-forcing of the turbo encoder, all small-state-capacity multifunctional hash functions of the turbo encoder feature polynomials that are coprime. Initialization phase: Synchronously execute the initialization functions of all small-capacity multi-functional hash functions; Absorption Phase: The absorption phase function synchronously executes all small-capacity multi-functional hash functions; Output phase: The output phase function synchronously executes all small-state-capacity multi-functional hash functions.

[0054] Clearly, the constructed large-state multifunctional hash function conforms to the definition in ☆1-3.

[0055] ☆8 Technical Effects: 1) Because the controlled units are basically the same, the software and hardware implementations can reuse a set of computational logic, and the clock and area interchangeability is particularly effective. 2) Designing and publicly reviewing a strong permutation or block cipher with a single large state is very costly. Theoretically, this invention can construct a multifunctional hash function with a larger or even infinitely high code rate based on a single strong permutation or block cipher design. More specifically, because {cl,cr} of each function are independent of each other, the input code rate is equal to the code rate of the small-state capacity hash function, and the output code rate can be infinitely expanded synchronously through circuit integration and clock synchronization.

[0056] "The controlled units are basically the same". The nonlinear logic of the main control unit is exactly the same, but the RC is different or the entropy value of other units is introduced. For example, the maximum code rate of the A->B->C->D->A->B->C->D circuit is designed with a code rate of 4096 bits. The content of the main control unit of the intermediate A circuit is also driven by the content of the main control unit of the D circuit. The content of the main control unit of the two A circuit sub-units can be considered to be independent of each other.

[0057] "Coprime characteristic polynomials" means that the characteristic polynomials of a turbo encoder should be as different and independent as possible. When multiple small-state-capacity multifunction hash functions are given the same input, the cost of forcing all small-state-capacity multifunction hash functions to zero in the master control unit register should be as high as possible. A good design has cyclic shift bits for circuits A and B, while a poor design is that A, B, C, and D share a single cyclic shift bit with different masks. Ideally, the random number pattern should generate m-sequences with coprime characteristic polynomials.

[0058] ☆9 The method for constructing large-state-capacity hash functions according to ☆8 is characterized in that, The input code rate and output code rate of the large state capacity hash function are 512 bits, denoted as circuit A; The input code rate of the small-state capacity hash function is 128 bits. The number of small-state capacity hash functions is four, denoted as A_X, A_Y, A_Z, and A_W. The master control unit parameters and controlled unit RC parameters of A_X, A_Y, A_Z, and A_W are as follows: Parameter names: right circular shifter, MASK, RC table PAR_A_X 1, 0x32922002011fdf7eL,0x0703 PAR_A_Y 3, 0x32a60c2a0041b37eL,0x0e06 PAR_A_Z 5, 0x404a040a00fbff7eL,0x1509 PAR_A_W 7, 0x2ed2102a00434f96L,0x1c0c Each sub-wheel of the main control unit is defined as follows: ax = right circular shift of the ax byte i⊕(2*ax&MASK), where ax is the 128-bit register of the master control unit; Referring to the accompanying drawings in the specification for a clearer understanding, each sub-wheel of the controlled unit is defined as follows: {cl,cr}->{c2.l,c2.r} {l0,r0}= { cl[63:0] ,cr[63:0]}; {l1,r1}= { cl[127: 64] ,cr[127: 64]}; {l2,r2}=ADJUSTEMENT_BIT(S odd({l0,r0})⊕{l1⊕RC,r1}); {l3,r3}=ADJUSTEMENT_BIT(S odd({l1,r1})⊕{l2⊕RC,r2}); { c2.r[127: 64] ,c2.r[63:0]}={r3,r2}; { c2.l[127: 64] , c2.l[63:0]}={l3,l2}; Where ADJUSTEMENT_BIT() equals {l,r}->{r,l<<<15}, and S_odd is the odd-numbered round logic of EWES.

[0059] Each child wheel adds entropy exchange calculation logic: XC `=WC⊕XC⊕YC; YC `=XC⊕YC⊕ZC; ZC `=YC⊕ZC⊕WC; WC `=ZC⊕WC⊕XC.

[0060] Define the number of sub-round iterations: Industrial grade / 4 to 6 rounds, Top secret grade / 6 to 8 rounds.

[0061] The newly constructed hash function is referred to as Circuit A. Circuit A's main competitor is the Keccake1600. Circuit A has a state capacity of 1536 bits and a maximum code rate of 512 bits. In terms of peak performance during the absorption phase, the C code performance of Circuit A / 4-child wheel mode is slightly better than that of Keccake1600. The processor used is an AMD 9700X@3.8GHz with a turbo boost of 5.5GHz, a gcc compiler, 64-bit code, 512-bit security, and 4 child wheels.

[0062] SHA3-512-72 byte code rate: 0.4553 GB / s; Circuit A - 512 - 72 bytes bitrate - 4 sub-wheels 0.4653 GB / s.

[0063] The effect of area and clock swapping in circuit A is demonstrated: This proves that the area and clock swapping effects of ☆9 and ☆8 are excellent, resulting in a superior area throughput. The hardware platform is EP3C25Q240C8, and the code is simulated using Modelsim 10.1 and synthesized using Quartus II 9.0. Considering that circuit A has 1600 bits of registers (including a 64-bit count), the 4289-cell performance is already very good. The low-area performance synthesis results of SM3 prove that on the EP3C platform, the area throughput of circuit A is significantly better than that of the SM3 circuit.

[0064] Area per beat of clock Circuit A, 1 sub-coil = 1 beat 41.8MHz 14.268 cells Circuit A, 1 sub-wheel = 2 cycles (controlled unit 1 / 2 multiplexed) 50.44mze 11.257 cells Circuit A, 1 sub-wheel = 6 cycles (controlled unit 1 / 4 multiplexed) 73.53MHz 6786 cells Circuit A, 1 sub-wheel = 10 beats (controlled unit 1 / 8 multiplexed) 75.47MHz 4289 cells 64-sub-column Sm3 circuit: 1 sub-column = 1 beat 74.58MHz 3544 cells The combined results of Xlinx devices and TMSC 90nm-ASIC are basically consistent with the above results.

[0065] ☆10. Based on the large state capacity hash function construction method described in ☆8 and ☆9, a larger state capacity hash function is constructed from circuit A and circuit B, denoted as circuit A->B; The difference between circuit B and circuit A lies in the different parameters of the main control unit and the RC parameters of the controlled unit. The other operational logic and control logic are exactly the same. The parameters of the four small state capacity hash functions corresponding to circuit B are as follows; Parameter names: right circular shifter, MASK, RC table PAR_B_X 9, 0x329220020045effeL,0x230f PAR_B_Y 11, 0x32a60c2a0069ddfaL,0x2a12 PAR_B_Z 13, 0x404a040a002bb9fcL,0x3115 PAR_B_W 15, 0x2ed2102a0005dff8L,0x3818 The control relationship between circuit A and circuit B is as follows: each sub-wheel adds entropy exchange calculation logic from A to B. BC = BC⊕AC.

[0066] The meaning of "->" in this manual is that after each iteration of the child wheel, the entropy exchange calculation logic from the main control circuit to the controlled circuit is added. A->B is equivalent to executing BC=BC⊕AC, and so on to the cascaded control mode of A, B, C, and D.

[0067] The overall structure and definition of A, B, C, and D are exactly the same except for the PAR_*_* parameters. It can be assumed that the software and hardware code share a common template, and the critical path and performance of each platform are exactly the same.

[0068] "->" represents the ingenuity of the engineers in the implementation example. When the state capacity of circuits A, B, C, and D reaches the level of 1536 bits, on the one hand, unidirectional entropy exchange is particularly friendly to circuit routing and low latency implementation. Simply put, bidirectional entropy exchange FPGAs often result in poor clock speeds even if routing is successful, because bidirectional entropy exchange is extremely unfriendly to line delay control. On the other hand, the total state capacity is very large, and the image attack capability has reached 512 bits. The security of hashes and random numbers is sufficient. Even without entropy exchange, the security of multi-functional application scenarios is sufficient.

[0069] The A->B circuit has a maximum code rate of 1024 bits, which can cover ICCS hash function tracks 512, 768, and 1024 output targets.

[0070] The A->B->C->D circuit has a maximum code rate of 2048 bits and is mainly used for random number encryption or authentication with extremely high throughput. If the code rate or throughput is insufficient, two or even an infinite number of A->B->C->D units can be cascaded. Theoretically, even if an infinite number of A->B->C->D units are cascaded, the critical path will not be corrupted.

[0071] The A->B->C->D structure is naturally friendly to pipeline hardware code; one sub-wheel divides the 4-stage pipeline, P1#S box layer#P2#“->”, EP3C device synthesis results, 130.24 MHz, 19.199 cells, and the throughput ratio is better than (41.8Mhz / 14.268 cells).

[0072] The difference between the A->B->C->D circuit, two A->B circuits, and four A circuits lies in the "->" operator, the main control unit parameters, and the controlled unit parameters. Therefore, a single set of software or hardware code can encode these circuits. Because AMD's AVX512 is a typical throughput-first design, meaning that while AMD's AVX512 instruction pairing capability is relatively poor, it offers 4-8 instruction parallelism, meaning that parallel resources are always utilized by AMD processors. Programming tests revealed that the time overhead of two parallel A circuits and one parallel A circuit is almost the same, as detailed below.

[0073] AMD 9700X @ 3.8G turbo boost up to 5.5G, gcc compiler, 64-bit code, 512-bit security / 4 child wheels / single thread; total data during absorption phase 64GB.

[0074] Circuit / Hash Algorithm Measured Time Inference Throughput Input Code Rate Output Code Rate A->B circuit 67.03 S 0.955 GB / S 512 1024 Two A circuits: 66.56 S, 0.962 G / S, +0.962 G / S, 512+512, 512+512 Single-channel A circuit 61.41 S 1.042 GB / S 512 512 X-unit circuit 8 channels 58.91 S 1.086GB / S (8 channels) 128*8 channels 128*8 channels ☆11. A multifunctional hash function calculation circuit (computation circuit device), which, when executed by physical execution or simulation software (Modelsim, etc.), can realize the functions of ☆1 and ☆8, with clk and control code as inputs, and includes registers and computational combinational circuits within the module, characterized in that, It also includes a plaintext input and a hash / xof output. The inputs to the computing circuit include clk, state update control code, and plaintext. The computing circuitry includes registers for state retention and combinational logic circuitry for computation. The computing circuit consists of a main control unit and a controlled unit. The main control unit is a turbo encoder, and the controlled unit consists of nonlinear logic and register C. The main function of the nonlinear logic is to enable the controlled unit to obtain the block cipher mode effect or strong permutation mode effect through multiple sub-wheel iterations. State capacity: The master control unit register has a bit count of r, where r equals the number of bits in the input plaintext. The number of bits in the controlled unit register is cr and cl. Wherein, the control relationship between r and cr is the linear operation described above; Reserved CL safety capacity to resist preimage attacks and collision attacks; The output of hash / xof is equal to cr or a complex logical function transformation including cr; The state update control code, in conjunction with clk, performs three functions. 1) Initialize the contents of the registers of the master control unit and the controlled unit; 2) The plaintext drives the turbo encoder to move, and the output of the turbo encoder synchronously drives the controlled unit to move; 3) Turn off the plaintext turbo encoder to run idle, and the output of the turbo encoder synchronously drives the movement of the controlled unit.

[0075] The computing circuit in ☆11 refers to "circuit device", including but not limited to, in addition to FPGA entity, ASIC entity, IP soft core results that can be instantiated in industry are also included, such as FPGA executable code, and ASIC industrial ecosystem functional use hardware code or wiring diagram that clearly conforms to the definition of ☆11 and ☆1.

[0076] ☆11 is the definition and implementation of ☆1 at the hardware circuit layer; each plaintext block needs to be absorbed and processed through multiple sub-rounds, so processing one plaintext block often involves executing a sequence of 233**3 function codes; the hash function constructed by ☆8 supports clock-to-area tradeoffs, so each function code obviously needs to be split into more clock cycles for execution.

[0077] ☆12. The multifunctional hash function timing calculation circuit described in ☆11, The output of hash / xof is equal to cr; The number of bits at the output of hash / xof is the same as the number of bits at the plaintext input. cr and cl have the same capacity.

[0078] The design purpose and effects of ☆12 can be referenced in ☆3. The main effects are: 1. In duplex mode, cl can ensure that the output bit rate meets the target threshold for image attack, reaching the limit cl. The optimal implementation circuit A has an input bit rate of 512 bits, an output bit rate of 512 bits, and cl of 512 bits; 2. It is as excellent as ASCON and keccak1600 based on sponge structures, with the computational logic overhead of the output (extraction stage) being equivalent to 0.

[0079] ☆13. The multifunctional hash function timing calculation circuit described in ☆11, The feature is that, in order to reuse the nonlinear calculation logic of the controlled unit, a programmable control port is added, which specifies the combined working mode and the decomposed working mode. The combined working mode corresponds to the calculation circuit of a large-state capacity multifunctional hash function, and the decomposed working mode corresponds to the calculation circuit of several small-state multifunctional hash functions. The relationship between a large-capacity multi-function hash function and several small-capacity multi-function hash functions is as follows: Wherein, the input of the large-state capacity hash function is equal to the bit concatenation of the input of the small-state capacity hash function. The output of the large-state capacity hash function is equal to the bit concatenation of the outputs of the small-state capacity hash function. The safety capacity cl of the large state capacity hash function is equal to the sum of the small state capacity hash functions cl; In order to reuse the nonlinear computation logic of the controlled unit, the controlled units of all small-state-capacity multifunctional hash functions are basically the same; To resist the zero-forcing of the turbo encoder, all small-state-capacity multifunctional hash functions of the turbo encoder feature polynomials that are coprime. The synchronization constraints for the combined mode are as follows; Initialization phase: Synchronously execute the initialization functions of all small-capacity multi-functional hash functions; Absorption Phase: The absorption phase function synchronously executes all small-capacity multi-functional hash functions; Output phase: The output phase function synchronously executes all small-state-capacity multi-functional hash functions.

[0080] Technical advantages: ☆13 is a circuit implementation of ☆8, inheriting all the technical advantages of ☆8. ☆14. The multifunctional hash function timing calculation circuit according to ☆13 The combined working mode is as follows: Mode 0: A->B->C->D, equals one 2048-bit code rate circuit. The decomposition working mode is one of the following modes: Mode 1: A->B, A->B; This is equivalent to two 1024 bitrate circuits; Mode 2: A,A,A,A; This is equivalent to four 512 bitrate circuits. Mode 3: {A_X,A_X,A_X,A_X} 4 groups, which is equivalent to a 128 bit rate circuit.

[0081] ☆14 is a description of the preferred embodiment. In practice, any two of these two can be selected as the combined working mode and the disassembled working mode, which fall within the scope of protection of this patent.

[0082] The code was simulated using Modelsim 10.1, and the synthesis results for the XCZU2CG-2I device using Vivado 2022 are as follows. The control port selects one of three circuits; the increase in area and clock delay is minimal, demonstrating significant industrial value.

[0083] It can be proven that a single hardware codebase can support both combined and multiple disassembled modes. Circuit / hash algorithm constrained clock (ns), inferred throughput (512 bits), area (LUT# FF) A->B circuit 5.5 + 0.446 2.69 GB / s 24681 / 52.25% #3272 / 3.46% Two-way A circuits: 5.5 + 0.012 ns, 2.9 GB / s, +2.9 GB / s, 23864 / 50.53%, #3272 / 3.46%. X-unit circuit, 8 channels, 5.5-0.262ns, 3.05GB / s (8 channels), 22903 / 48.49%#3272 / 3.46% ☆14 Control Port 6.0 +0.366ns Empty 27854 / 58.97%#3273 / 3.46% In industrial scenarios ☆14 and ☆13, an ASIC chip supports a family of standards, and the control port selects the circuit mode, which is the chip's selling point: The control circuit mode can theoretically be integrated into Intel or AMD CPUs, communication processing chips, and commercial cryptographic coprocessors (units).

[0084] The A->B->C->D circuit pattern can be used for high-speed duplex encryption and authentication functions, with main competitors being RORA and SNOW-V. Parallel 4-way A-circuit mode, parallel 4-way Merkel tree computation, suitable for accelerating blockchain and hash-based digital signatures; The parallel 16-channel X-circuit mode enables 16 channels of medium-low speed duplex encryption and authentication functions in parallel, making it suitable for base stations and communication backbone nodes. The main competitors of the X-circuit are ASCON, SNOW-3G, and ZUC.

[0085] ☆☆3 Summary of Invention Effects 4: Based on ☆8, ☆13, and ☆14, we can summarize the new features and effects of ##5.3. With the controlled unit remaining essentially unchanged, it is feasible to construct a multi-functional hash function with larger state capacity and higher code rate through circuit integration and clock synchronization. The newly constructed hash function inherently possesses the advantages of ☆13 and ☆14. Compared to the background technology, it is unique and possesses unique advantages. Attached Figure Description

[0086] Figure 1 A structural diagram illustrating the working principle of this invention; Figure 2 An embodiment based on a turbo encoder and block cipher mode, with the blue main control unit corresponding to ☆10 and ☆10; Figure 3 This is the optimal embodiment of the controlled unit, and more specifically, the Enc_1round structure diagram of the x, y, z, w controlled units;

[0087] Figure 4 Diagram illustrating the principle of a mid-interval attack under strong permutation mode; The embodiments and accompanying drawings are used to explain the present invention and do not constitute an improper limitation of the present invention. It should be noted that blue corresponds to the main control unit, red and green correspond to the controlled units, and green represents the safety capacity of the main control unit. Detailed Implementation

[0088] Terminology Explanation

[0089] turbo encoder The most important component of this invention, also known as the main control unit, is essentially a linear iterative encoder, constructed with reference to knowledge of linear feedback shift registers, congruence, and mixed congruence in multiplication and addition. The key innovation is that, for attackers, the turbo encoder is the first, cost-effective, and efficient firewall. By manipulating the state difference or forcing the state to zero of the turbo encoder, at least two plaintext blocks are consumed, and the last block has no degree of freedom for modification. Through linkage with the main control unit, a higher bit rate or throughput can be achieved. The movement is driven by plaintext blocks or idle movement, and its output drives the movement of the main control unit. Preferably, the capacity of the main control unit is no greater than the maximum capacity of the main control unit. Compared to the nonlinear logic transformations of the main control unit, the low implementation cost refers to minimal software and hardware costs and minimal critical path latency.

[0090] Block cipher mode The invention points of this embodiment show that, compared to the strong permutation mode, both compression and diffusion effects are better. Advantage 2: The first output of the block cipher mode does not require a no-load run. The theoretical assumption is that it is equivalent to a large substitution table controlled by the master control unit's state information. The input of the substitution table is the current state of the master control unit, and the output of the substitution table is the next state of the master control unit. The logic of the first sub-round of the block cipher is denoted as Enc_1round(). The strong permutation mode, referencing the sponge structure and its embodiments, is theoretically equivalent to a large substitution table without any control factors. The technical advantage compared to the strong permutation mode is that the plaintext block or turbo encoder output cannot directly control the state of the master control unit.

[0091] Round 1 The security assumption is strongly related to the overall structural design concept of this invention. It corresponds to one absorption or output, like a strong permutation of a sponge structure or a block cipher encryption of the turbo-keccak1600 embodiment. One round corresponds to the ideal strong permutation and ideal block cipher based on the random oracle model. Simply put, the output state is unpredictable and is regarded as a pure random number output in the security assessment.

[0092] 1st wheel The most important concept in implementing this invention is comparable to SHA3's 24 sub-rounds and ASCON's 12 sub-rounds. The number of rounds required to iterate through one sub-round logic to achieve the desired strong permutation or ideal block cipher effect is determined; at least three sub-rounds are needed to achieve the desired effect. Turbo-keccak uses 24 sub-rounds, while the FFHASH example uses 4-8 sub-rounds, equivalent to the number of times one absorption or output iteration Enc_1round() is performed.

[0093] (Input) Bitrate One of the most important indicators strongly correlated with the technical effect of this invention is the input bit rate and the output bit rate. The input bit rate is less than or equal to the capacity of the main control unit, equivalent to the bit rate *r* of a sponge structure. The output bit rate, preferably *cr*, is used as the output, while the other bits of the main control unit have a safe capacity *cl*. It should be noted that the maximum input bit rate does not exceed the capacity of the main control unit register and the controlled register, and the maximum output bit rate cannot exceed the capacity of the controlled register. Similar to the sponge structure, the hash function of this invention is successfully constructed, and its maximum input and maximum output bit rates are fixed. Furthermore, the actual input and output bit rates must be selected based on industrial safety requirements. The overall structure of this invention has a similar conclusion to the sponge structure (input / output): the actual bit rate indicator and the safety indicator cannot be simultaneously achieved but are interchangeable.

[0094] X, Y, Z, W Also known as the X, Y, Z, W (computation) circuit, the optimal embodiment of ☆1-7. This specification defaults to using circuits to represent the constructed multi-functional hash algorithm. The inventors believe that circuit or hardware language descriptions are algorithm descriptions dominated by information control flow, and are most suitable for describing the overall structural technical features of this invention. Compared with common step language descriptions and software algorithm language descriptions, they are simpler, clearer, and have a wider scope of protection. The main control unit is 128 bits, the controlled unit is 256 bits, and the maximum input / output bit rate is 128 bits. It belongs to the small-state, small-bit-rate algorithm design. The core logic of the controlled unit is exactly the same. The controlled unit is a P1->S box->p2 component design. Figure 3 With an input / output code rate of 128 and a register of 384, the X circuit is designed for optimal duplex performance. Its competitors are the Snow-3G and ZUC.

[0095] A, B, D, C In the optimal implementation of ☆8-9, the main control unit is 512 bits, the controlled unit is 1024 bits, and the maximum input / output bit rate is 512 bits. It is constructed by combining X, Y, Z, and W circuits through circuit integration and clock synchronization. Its entropy exchange component is in a mutual control mode. The core logic of the controlled unit is exactly the same. The input / output bit rate of 512 / register 1536 is the optimal design for duplex performance. The competitor of circuit A is Keccak1600 and Snow-5G. On the other hand, ☆8 is continued to be used to construct circuits with larger states and larger bit rates, A->B, A->B->C->D.

[0096] Entropy exchange component The logic transformations related to ☆8 are executed by the runtime for each child round, used to exchange the controlled unit information of the small state capacity hash function. Classified according to the information exchange type of the master control unit: 1. No information exchange, 2. Mutual exchange of state information, 3. Unidirectional exchange of state information. The ABCD circuit is constructed to exchange state information; the equations refer to ☆9. Modulo-2 addition logic transformations are preferred, as are reversible logic transformations (envelope entropy).

[0097] Example 1 of the main control unit (TTK series) Example TTK: "Turbo code Encoder Turbocharging Keccak1600", a turbo encoder that accelerates the Keccak1600. TTK demonstrates that this invention can empower industrial standards or designs based on sponge structures, using strong permutations as controlled units, to increase the input and / or output bit rates with little or no increase in computational overhead. The output bit rate's resistance to meet-in-the-middle attacks is no less than the birth limit. The Keccak1600-24 sub-coils and ASCON are world-leading competitors, having undergone self-design and self-evaluation -> public review -> industrial standardization, and their security and performance are widely recognized as excellent. This invention can increase the maximum bit rate and meet-in-the-middle attack resistance of SHA3 to 800 bits, and ASCON to 160 bits. The following describes the empowerment of SHA3.

[0098] Main control unit: Turbo encoder is a 64-bit block unit cyclic shifter with a capacity of 512, 768 or 1024; Controlled unit: Block cipher mode, 6 (or 4) sub-wheels, where each sub-wheel is a p of a keccak1600. 4 (or p) 6 ), p=ι(χ(π(ρ(θ())))); That is, the 24-wheel Keccak strong permutation is transformed into a 6-wheel block cipher, and the security assumption becomes a huge substitution table controlled by the content of the turbo encoder; The invention point of the embodiment is the block cipher mode, which has greater variation than the strong permutation, the preimage security index is easier to evaluate, and the first hash / xof output does not need to run once.

[0099] Referring to the design concept of Keccak1600, which does not output the safety capacity C, cr or its truncated bit is used as the output.

[0100] Hash mode (compatible with mac mode) Input bitrate Output bitrate Meeting attack indicator Birthday boundary Turbo1536 ->keccak1600 1536 512 800 512 Turbo1536 ->keccak1600 1536 768 800 768 Turbo1536 ->keccak1600 1536 512+512xof 800 1024 Hash mode (compatible with mac mode) Input bitrate Output bitrate Meeting attack indicator Birthday boundary Turbo1024 ->keccak1600 1024 512 800 512 Turbo1024 ->keccak1600 1024 768 800 768 Turbo1024 ->keccak1600 1024 512+512xof 800 1024 The 64-bit code uses the keccak1600.c source code of openssl-3.3, an AMD 9700X processor and a gcc 13.1.0 compiler, and the test accuracy is 0.1G times of keccak-24 sub-round iterations; it can be proven that it significantly improves the input code rate.

[0101] SHA3-512-72 byte input bitrate, throughput 0.455 GByte / S, 12.1 C / B; Turbo1024->keccak1600-128 byte bitrate, throughput 0.755 GByte / s, 7.3 C / B; Turbo1536->keccak1600 - 192 byte bitrate, throughput 1.0 GByte / S, 5.5 C / B.

[0102] The following is a duplex mode design under the condition of minimizing register usage, highlighting the mutual constraints between security and maximum bit rate. Industry-standard designs should consider hardware reuse of software modules, and it is recommended to share Turbo1536->keccak1600 or Turbo1024->keccak1600.

[0103] Full-duplex mode (compatible with random number mode) Input bitrate Output bitrate Encounter attack indicator Birthday boundary Turbo1024 ->keccak1600 1024 1024 800 512 Turbo768 ->keccak1600 768 768 800 768 Turbo512->keccak1600 512 512 800 1024.

[0104] Turbo-merkel # Turbo1024 ->keccak1600 Input bitrate Output bitrate Original image Birthday boundary Turbo_merkelroor512 1024 512 512 Turbo_merkelroor768 768->768 768 768 Turbo_merkelroor1024 512->512 512+512xof 1024.

[0105] Evaluate whether the meet-in-the-middle attack reaches the birth bound of the preimage attack. Block ciphers are reversible, and the attack model is described below: HashInit: -fwd1->-fwd2->intermediate meeting point <-inv-{preimage, cl} <-inv-{preimage xof, cl} Intermediate meeting point: Construct plaintext blocks in the fwd region corresponding to "-fwd1->-fwd2->" and "<-inv-". First, ensure the master control units at the meeting point are identical (convergence). Based on the zero-forcing property of the turbo encoder state, at least two plaintext blocks need to be constructed. According to the block cipher modular assumption, the controlled units at the meeting point are treated as random oracles and cannot be distinguished. Block cipher p 24 It is an ideal block cipher, meaning the plaintext or ciphertext is essentially random and unpredictable. Therefore, it meets the estimation of the number of random collisions. Constructing '-fwd1->-fwd2->' 'a' sequences and '<--inv--' sequences 'b' sequences, only a*b=2. 1600 Only when an average of one collision occurs can the optimal solution for an encounter attack be a=b=2. 800 .

[0106] The technical results, in two aspects, demonstrate that technical issues ##2 and ##4 have been resolved.

[0107] 1) “Superior to others”: 512-bit secure hash / xof, keccak1600 has a maximum bitrate of 576, and after being enabled, the maximum bitrate is 1024, resulting in a performance improvement of 1024 / 576. For example, processing a 1315-bit plaintext message, sha3-512 requires 2.29 (1315 / 576) plaintext blocks, while Turbo1024 -> keccak1600 requires 1.28 (1315 / 1024) plaintext blocks.

[0108] 2) "Unique Features": The maximum bitrate in duplex mode is 800 bits / round, while the Keccak1600 has a bitrate of 1600 / 3 bits. After being enabled, the maximum bitrate is 768 to 800 bits, which is 1.5 times the performance.

[0109] 3) "Unique Features" 2: Hash mode supports 1024-bit output. First, output 512 bits, then output 512 x of bits after one round of block cipher. An attack in the middle consumes 800 degrees of freedom. Based on the assumption of strong permutation or block cipher, 512 -> 512 x of consumes another 512 degrees of freedom. Therefore, it is inferred that 512 + 800 is greater than the birth limit of 1024, and the original image index reaches the birth limit.

[0110] Following TTK's design philosophy, this invention can also empower ASCON, turbo256->ASCON320, and the 256-bit hash value equals 128->128xof.

[0111] Example 2 of the main control unit By referring to the research results and construction demonstrations of the overall structure and nonlinear components of block ciphers, strong permutation or block cipher modes of the master control unit can be constructed. Simply put, for fast diffusion and good parallel performance, a SP structure is adopted; for even better diffusion performance than the SP structure, the component follows the invention point P1->S layer->P2 as described in the embodiment. To save transistors or prioritize area optimization, a Festel structure, a generalized Festel structure, a Lai-mass structure, etc., can be adopted. Component design includes, but is not limited to, S-box + linear diffusion layer, ARX instructions, keccake, or ASCON-type sparse tapped SP structure components.

[0112] The 256-bit block cipher with any block width can be replaced. The main control unit of the X, Y, Z, and W circuits is shown below as a transistor-saving design demonstration.

[0113] The Festel structure has two branches and 128 bits of logic in each child wheel. 8-10 wheels are recommended.

[0114] The 4-branch generalized feistl structure has a 64-bit sub-wheel logic. The overall structure adopts SM4, and the components adopt Camellia. 16-20 wheels are recommended.

[0115] Embodiments of the Invention - FFHASH Series

[0116] Based on state capacity and maximum bit rate, the system is divided into three levels. The core logic of the main control unit is exactly the same for each level, and the construction process uses ☆8 technology twice. The state capacity of X, Y, Z, and W is 384 (bits), with a maximum bit rate of 128, comparable to ASCON and SNOW-3G; the state capacity of A, B, C, and D is 384*4, with a maximum bit rate of 512, comparable to Keccak1600; the maximum bit rates of the A->B circuit and the A->B->C->D circuit are 1024 and 2048, respectively, designed for extremely high or infinite throughput.

[0117] The circuit definitions, structures, and parameter tables for A_X, A_Y, A_Z, A_W, A, and B have been described in sections ☆4, ☆6, and ☆9, with related figures attached. Figure 2 and Figure 3 For a detailed definition of the EWES odd-round logic, please refer to the relevant patent "A Self-Reversible Block Cipher". The byte cyclic shift of the main control unit is coprime, and the m-sequence feature polynomial corresponding to the idle run of the main control unit is coprime. The design purpose is to maximize the cost of forcing the total state to zero when the code rate is equal to the maximum code rate of the smallest unit.

[0118] Feature and Advantage 1: Duplex-friendly, ensuring security while maintaining a maximum code rate / state capacity of 1 / 3. Compared to sponge structures, SP-type nonlinear iterators require less space, which is beneficial for instruction set design and implementation. The actual main control unit reuses the odd-round logic of the EWES block cipher. The purpose of using this component is to implement block encryption / decryption functions and multi-functional hash functions in a single hardware setup at high speed.

[0119] Feature and Advantage 2: As mentioned earlier, the main control unit of the 3rd stage is equivalent to a copy of the main control unit of the X circuit, so the area-for-clock tradeoff is inherently favorable. The synthesis results of EP3C25Q240C8 can corroborate this, showing that circuit A can be implemented in one round with one clock cycle, one round with multiple clock cycles with 1 / 2 logic, one round with multiple clock cycles with 1 / 4 logic, and one round with multiple clock cycles with 1 / 8 logic.

[0120] Feature and Advantage 3: As described in section ☆14, the programmable control port is added to specify the combined working mode and the decomposed working mode, realizing a set of physical circuits covering the 3-level circuit standard of FFHASH.

[0121] Feature and Advantage 4: Supports bitrate-swapping sub-wheel mode. For example, the security index of a 128-bit input bitrate (each XYZW unit copies 1 copy) / 1 sub-wheel in circuit A is equivalent to the security index of a 512-bit input bitrate / 4 sub-wheels. Therefore, it is optimized for short frame data and low latency. Design concept: Because the cyclic shift bits are coprime, forcing 0 to 0 simultaneously on all 4 XYZW main control units incurs the highest cost, which is equivalent to the concept of coprime.

[0122] Preimage safety boundary analysis of the strong permutation mode of the main control unit Purpose of proof: Invention point safety capacity 2 c.l Preimage safety boundary 2 in duplex mode c.l The following argument demonstrates that the X-circuit operates in a strong substitution mode + duplex mode, and its anti-image safety margin is not less than 2. c.l .

[0123] The X circuit's cr, cl, and r are all 128 bits, proving that the antigen imaging safety margin is no less than 2. c.l , i.e., birthday world 2 128 Please refer to the patent drawings for a schematic diagram of an attack occurring in the middle of an encounter. Figure 4 HashInit: -fwd1->-fwd2->intermediate meeting point <-inv-{preimage, cl} 1) The most economical attack route is to first let r2 converge the state or state difference to 0 at the "intermediate meeting point", and then let {c2.l,c2.r} and {c2`.l,c`2.r} meet in the middle. Let's assume that the forward r2 and the reverse "<-r3" state both converge to 0. 2) According to the zero-forcing property of turbo codes, the plaintext block used for zero-forcing has no degrees of freedom, so c2.r and c2.l are random states; 3) The states of {cl,cr} are 256 bits, so the most economical collision overhead is 2 bits for forward production of "-fwd1->-fwd2->". 128 One state, "<-inv-" reverse production 2 128 There are several states.

[0124] In summary, given Y, the cost of using the intermediate encounter attack of the state is 2. 128 Intermediate encounter attack is the most economical. To clarify, if the intermediate encounter point is not used to force zero on the main control unit, the attack complexity is equivalent to a birthday attack.

[0125] To clarify, inserting more blocks between "plaintext block 1" and "forcing 0 plaintext block" does not modify the degrees of freedom of the final plaintext block; the final child wheel's c2.r and c2.l are random states. Similarly, inserting more blocks from the preimage to "intermediate encounter" does not affect the above properties.

[0126] The above conclusions can be generalized to two points: the non-duplex mode (no output) is more effective at resisting attacks, while the block cipher mode is more effective at resisting attacks.

[0127] Immunity to length extension attacks: Because the cl security capacity cannot be viewed or predicted, sponge-like structures are also immune to length extension attacks on the overall structure of this invention.

[0128] Collision attack is equivalent to birthday bound: because cl cannot be viewed, converged, or predicted, it is equivalent to cl being a random number. Even if cr and r are controllable, the result of the forward calculation of the hash value of "-fwd->" is unpredictable and is treated as a random number output; therefore, it is equivalent to birthday bound.

[0129] Design results of entropy exchange component:

[0130] The three types of information exchange each have their own advantages and disadvantages. Engineering friendliness and security friendliness cannot be achieved at the same time. The following are engineering design suggestions.

[0131] When the state capacity is small, mutual exchange of state information is chosen. After one or more child rounds of iteration, the information of the master control units can diffuse to each other. At this time, the wiring and critical path overhead is not large, and the contents of all master control units are indivisible, significantly improving security. The entropy exchange component of Example ☆9 is a design with 3 modulo-2 plus circuits and 4 branches. The diffusion effect is extremely fast but slightly inferior to the MDS code design; the wiring delay loss is calculated as (5.5 + 0.012) - (5.5 - 0.262) based on the synthesis results of the xczu2cg device. The following is the calculation logic definition of the entropy exchange component of ☆9.

[0132] XC `=WC⊕XC⊕YC;YC `=XC⊕YC⊕ZC;ZC `=YC⊕ZC⊕WC; WC `=ZC⊕WC⊕XC.

[0133] Recommended design for exchanging two units: XCr `=XCr⊕YCr; YCl `=XCl⊕YCl.

[0134] When the state capacity of this invention reaches the 1536-bit level, unidirectional exchange of state information is preferred, making MAC code extraction clearly possible. On the one hand, according to convention, 1536 bits of random number or hash security is at least 768 bits, ensuring security even without information exchange; on the other hand, unidirectional circuit implementation is user-friendly, with a small critical path and pipeline-friendly implementation. In comparison, it was found that the line delay of circuits implementing mutual information exchange is extremely large, and tests showed that FPGA routing for mutual information exchange in large state designs often fails, and even if successful, the critical path becomes very poor. As described in ☆10, the entropy exchange component is BC=BC⊕AC, and based on the comprehensive results, the line delay loss of A->B is calculated to be (5.5+0.446)-(5.5-0.262).

[0135] In special cases, a design without information exchange, i.e., with an empty entropy exchange component, can be chosen. This has significant theoretical value, with advantages such as no critical path loss or software loss, and inherently favorable hash patterns. However, the authentication and encryption modes are susceptible to segmentation. Utilizing the ☆8 theory, an infinite output bitrate hash function can be designed. Because the master control unit outputs independently, segmentation attacks are ineffective against the hash pattern.

[0136] The specific embodiments of the present invention disclosed above are intended to aid in understanding the content of the present invention and in its implementation. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the inventive points of "safe capacity" and "master-controlled -> controlled." For example, utilizing knowledge of linear shift registers, the capacity of a newly constructed master control unit can be increased by a factor of N by cascading (series) and / or paralleling multiple coprime master control units, where N equals the number of master control units. Based on the knowledge of linear shift registers, the cost of forcing zero is simultaneously increased from 1 plaintext block to N plaintext blocks. In conclusion, the scope of protection of the present invention should not be limited to the content disclosed in the embodiments of this specification.

Claims

1. A method for constructing a multifunctional hash function based on a turbo encoder, wherein the function includes an initialization phase, an absorption phase, and an output phase, characterized in that, Overall structure: The main control unit is a turbo encoder, and the controlled unit is a nonlinear logic iterator. The main function of the nonlinear logic iterator is to obtain the block cipher mode effect or strong permutation mode effect through multiple sub-wheel iterations. Control relationship: Plaintext blocks drive the movement of the master control unit, and the output of the master control unit drives the movement of the controlled unit through linear operations; State capacity: The number of bits in the main control unit register is r, where r is referred to as the input code rate. The number of bits in the controlled unit register is cr and cl. cr and r are equal, where the control relationship between r and cr is the linear operation described above; cl is the safety capacity, used to resist preimage attacks and collision attacks; Absorption phase: The plaintext block controls the movement of the main control unit, and the main control unit outputs synchronous control to move the controlled unit; Output stage: Extract {cl, cr} as output or compress {cl, cr} with information from the main control unit for output.

2. The method for constructing a multifunctional hash function according to claim 1, characterized in that, The controlled unit operates in block cipher mode, extracting {cl, cr} as output; The specific steps are as follows: In the first sub-wheel, the plaintext block drives the turbo encoder to move; in the other sub-wheels, the turbo encoder runs idle, that is, the plaintext block input is turned off. {cl,cr} is equivalent to the plaintext of the block cipher, and the output of the turbo encoder is used as the block cipher wheel key; For every movement of one sub-wheel by the turbo encoder, the block cipher encrypts one sub-wheel.

3. The method for constructing a multifunctional hash function according to claim 2, characterized in that, cr and cl are equal The bit width of cr is 128, 256, 384, 512 or a multiple of 512.

4. The method for constructing a multifunctional hash function according to claim 3, characterized in that, The main control unit is defined as follows: ax = ax bytes cyclic shift i ⊕ (2 * ax & MASK), where ax is the master control unit register.

5. The method for constructing a multifunctional hash function according to claim 3, characterized in that, The controlled unit is defined as follows: The logic transformation of the controlled unit sub-wheel is a nonlinear component P1->S-box layer->P2, where P1 and P2 are linear diffusion layers.

6. The method for constructing a multifunctional hash function according to claim 5, characterized in that, The P1->S box layer->P2 is EWES block cipher odd-round logic -> bit permutation, where the bit permutation is {l,r}->{r,l<<<15}.

7. The method for constructing a multifunctional hash function according to claim 6, characterized in that, The sub-wheel of the block cipher mode also includes count and RC parameters, where count is a counter for the number of times the controlled unit moves, and RC is a wheel constant.

8. A method for constructing a large-state-capacity, multi-functional hash function, wherein the function includes an initialization phase, an absorption phase, and an output phase, characterized in that, First, define multiple small-state-capacity multi-functional hash functions, and then construct a large-state-capacity multi-functional hash function through circuit integration and clock synchronization. The small-state-capacity hash function conforms to the definition of claims 1-7. in ; The input to the large-state capacity hash function is equal to the bit concatenation of the input to the small-state capacity hash function. The output of the large-state capacity hash function is equal to the bit concatenation of the outputs of the small-state capacity hash function. The safety capacity cl of the large state capacity hash function is equal to the sum of the small state capacity hash functions cl; In order to reuse the nonlinear computation logic of the controlled unit, the controlled units of all small-state-capacity multifunctional hash functions are basically the same; To resist the zero-forcing of the turbo encoder, all small-state-capacity multifunctional hash functions of the turbo encoder feature polynomials that are coprime. Initialization phase: Synchronously execute the initialization functions of all small-capacity multi-functional hash functions; Absorption Phase: The absorption phase function synchronously executes all small-capacity multi-functional hash functions; Output phase: The output phase function synchronously executes all small-state-capacity multi-functional hash functions.

9. The method for constructing a large-state-capacity hash function according to claim 8, characterized in that, The input code rate and output code rate of the large state capacity hash function are 512 bits, denoted as circuit A; The input code rate of the small-state capacity hash function is 128 bits. The number of small-state capacity hash functions is four, denoted as A_X, A_Y, A_Z, and A_W. The master control unit parameters and controlled unit RC parameters of A_X, A_Y, A_Z, and A_W are as follows: Parameter names: right circular shifter, MASK, RC table PAR_A_X 1, 0x32922002011fdf7eL,0x0703 PAR_A_Y 3, 0x32a60c2a0041b37eL,0x0e06 PAR_A_Z 5, 0x404a040a00fbff7eL,0x1509 PAR_A_W 7, 0x2ed2102a00434f96L,0x1c0c Each sub-wheel of the main control unit is defined as follows: ax = right circular shift of the ax byte i⊕(2*ax&MASK), where ax is the 128-bit register of the master control unit; Each sub-wheel of the controlled unit is defined as follows: {cl,cr}->{c2.l,c2.r} {l0,r0}= { cl[63:0] ,cr[63:0]}; {l1,r1}= { cl[127: 64] ,cr[127: 64]}; {l2,r2}=ADJUSTEMENT_BIT(S odd({l0,r0})⊕{l1⊕RC,r1}); {l3,r3}=ADJUSTEMENT_BIT(S odd({l1,r1})⊕{l2⊕RC,r2}); { c2.r[127: 64] ,c2.r[63:0]}={r3,r2}; { c2.l[127: 64] , c2.l[63:0]}={l3,l2}; Where ADJUSTEMENT_BIT() equals {l,r}->{r,l<<<15}, and S_odd represents the odd-numbered round logic of EWES; Each child wheel adds entropy exchange calculation logic: XC `=WC⊕XC⊕YC; YC `=XC⊕YC⊕ZC; ZC `=YC⊕ZC⊕WC; WC `=ZC⊕WC⊕XC; Define the number of iterations for the child rounds: 4 to 8 rounds.

10. The method for constructing a large-state-capacity hash function according to claims 8 and 9, characterized in that, A hash function with a larger state capacity is constructed from the A circuit and the B circuit, denoted as the A->B circuit; The difference between circuit B and circuit A is that the parameters of the main control unit and the RC parameters of the controlled unit are different. The other operation logic and control logic are exactly the same. The parameters of the four small state capacity hash functions of circuit B are as follows. Parameter names: right circular shifter, MASK, RC table PAR_B_X 9, 0x329220020045effeL,0x230f PAR_B_Y 11, 0x32a60c2a0069ddfaL,0x2a12 PAR_B_Z 13, 0x404a040a002bb9fcL,0x3115 PAR_B_W 15, 0x2ed2102a0005dff8L,0x3818 The control relationship between circuit A and circuit B is as follows: each sub-wheel adds entropy exchange calculation logic from A to B. BC = BC⊕AC.

11. A multifunctional hash function calculation circuit device, with clk and control code as inputs, and including registers and calculation combinational circuits within the module, characterized in that, It also includes plaintext input and hash / xof output. The inputs to the computing circuit include clk, state update control code, and plaintext. The computing circuitry includes registers for state retention and combinational logic circuitry for computation. Its features are, The computing circuit consists of a main control unit and a controlled unit. The main control unit is a turbo encoder, and the controlled unit consists of nonlinear logic and register C. The main function of the nonlinear logic is to enable the controlled unit to obtain the block cipher mode effect or strong permutation mode effect through multiple sub-wheel iterations. State capacity: The master control unit register has a bit count of r, where r equals the number of bits in the input plaintext. The controlled unit register has two bits: cr and cl. Wherein, the control relationship between r and cr is the linear operation described above; cl is the safety capacity, used to resist preimage attacks and collision attacks; The output of hash / xof is equal to cr or a complex function transformation including cr; The state update control code, in conjunction with clk, performs three functions. 1) Initialize the contents of the registers of the master control unit and the controlled unit; 2) The plaintext drives the turbo encoder to move, and the output of the turbo encoder synchronously drives the controlled unit to move; 3) With plaintext disabled, the turbo encoder runs idle, and the output of the turbo encoder synchronously drives the movement of the controlled unit.

12. The multifunctional hash function timing calculation circuit device according to claim 11, characterized in that, The output of hash / xof is equal to cr; The number of bits at the output of hash / xof is the same as the number of bits at the plaintext input. cr and cl have the same capacity.

13. The multifunctional hash function timing calculation circuit device according to claim 11, characterized in that, In order to reuse the nonlinear calculation logic of the controlled unit, a programmable control port is added. The programmable control port specifies the combined working mode and the decomposed working mode. The combined working mode corresponds to the calculation circuit of a large-state capacity multifunctional hash function, and the decomposed working mode corresponds to the calculation circuit of several small-state multifunctional hash functions. The association between a large-capacity multifunctional hash function and several small-capacity multifunctional hash functions is defined as follows: in; The input to the large-state capacity hash function is equal to the bit concatenation of the input to the small-state capacity hash function. The output of the large-state capacity hash function is equal to the bit concatenation of the outputs of the small-state capacity hash function. The safety capacity cl of the large state capacity hash function is equal to the sum of the small state capacity hash functions cl; In order to reuse the nonlinear computation logic of the controlled unit, the controlled units of all small-state-capacity multifunctional hash functions are basically the same; To resist the zero-forcing of the turbo encoder, all small-state-capacity multifunctional hash functions of the turbo encoder feature polynomials that are coprime. The synchronization constraints for the combined mode are as follows; Initialization phase: Synchronously execute the initialization functions of all small-capacity multi-functional hash functions; Absorption Phase: The absorption phase function synchronously executes all small-capacity multi-functional hash functions; Output phase: The output phase function synchronously executes all small-state-capacity multi-functional hash functions.

14. The multifunctional hash function timing calculation circuit device according to claim 13, characterized in that, The combined working mode and the decomposed working mode are combinations of any two of the following four modes; Mode 0: A->B->C->D, equals one 2048-bit code rate circuit. Mode 1: A->B, A->B; equals two independent 1024 bitrate circuits; Mode 2: A,A,A,A; equals 4 independent 512 bitrate circuits. Pattern 3: {A_X,A_X,A_X,A_X} 4 groups, equal to 16 independent A_X.