Operation circuit, chip and computing device for executing hash algorithm

By adopting multiple operation stages with pipeline structure in the computing circuit, and using delay modules and supplementary delay modules to make the calculation delays of each operation stage basically equal, the problem of difficult to improve the calculation frequency and throughput rate of the computing circuit in the prior art is solved, and the effect of reducing the power consumption and computing power ratio is achieved.

CN111813452BActive Publication Date: 2025-05-06SHENZHEN MICROBT ELECTRONICS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010837928.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-19
Publication Date
2025-05-06
Estimated Expiration
2040-08-19

AI Technical Summary

Technical Problem

The prior art is difficult to improve the calculation frequency and throughput of the computing circuit without reducing the calculation delay of the operational-level combined logic module, which leads to difficulty in reducing the power consumption and computing power ratio.

Method used

Multiple operation stages arranged in a pipeline structure are adopted, each operation stage includes a combined logic module, a delay module and a supplementary delay module. The delay module and a supplementary delay module are formed through the same delay unit connected in series, ensuring that the calculation delays of each operation stage are basically equal.

Benefits of technology

It is realized that without reducing the calculation delay of the operational-level combined logic module, the calculation frequency and throughput of the operational circuit are improved, thereby reducing the power consumption and computing power ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111813452B_ABST
    Figure CN111813452B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an operation circuit, a chip and a computing device for executing a hash algorithm. The operation circuit for executing the hash algorithm includes a plurality of operation stages arranged in a pipeline structure, each operation stage includes: a group of inputs and a group of outputs, the inputs are correspondingly coupled to the outputs of the previous operation stage, and the outputs are correspondingly coupled to the inputs of the next operation stage; a plurality of combinational logic modules, each of which has an input coupled to at least a part of a group of inputs; a plurality of delay modules, each of which has an input coupled to one of a group of inputs, and an output coupled to one of a group of outputs that is not coupled to the combinational logic module, so that such outputs are each coupled to a delay module; a plurality of supplementary delay modules, each of which has an input coupled to the output of the corresponding combinational logic module, and an output coupled to one of a group of outputs, wherein each delay module and the supplementary delay module are composed of identical delay units connected in series, so that the computational delay from the input of each operation stage to each of the outputs is substantially equal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an operation circuit for executing a hash algorithm, and a chip and a computing device including the operation circuit. Background Art

[0002] Chip size, chip speed and chip power consumption are three crucial factors that determine performance. Among them, chip size determines chip cost, chip speed determines computing power, and chip power consumption determines power consumption. In practical applications, the most important performance indicator is the power consumption per unit computing power, that is, the power consumption to computing power ratio.

[0003] Figure 1 The prior art operation circuit 100 is shown. The operation circuit 100 uses a pipeline structure to implement the SHA-256 algorithm.

[0004] like Figure 1 As shown, the operation circuit 100 includes N operation stages arranged in a pipeline structure, wherein each operation stage has a set of inputs 101 and a set of outputs 102, a set of inputs of each operation stage is correspondingly coupled to a set of outputs of the previous operation stage, and a set of outputs of each operation stage is correspondingly coupled to a set of inputs of the next operation stage.

[0005] Each operation stage includes a plurality of combinational logic modules 111 , 112 , 113 for performing combinational logic operations based on data input to the operation stage.

[0006] In addition, each operation stage also includes a set of registers for storing data. Figure 1 As shown, each group of registers includes 8 cache registers A, B, C, D, E, F, G, H and 16 extension registers W0, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14, W15.

[0007] It should be noted that, for ease of understanding, Figure 1 The number of each group of registers in the SHA-256 algorithm is compiled corresponding to the SHA-256 algorithm, and the connection relationship between each register and each combinational logic module 111, 112, 113 is also schematically drawn corresponding to the SHA-256 algorithm. For the sake of clarity, the connection relationship between the registers and each combinational logic module 111, 112, 113 is drawn only in the first operation level.

[0008] Each group of registers is controlled by a clock, passing data along each operation stage in sequence. In each clock cycle, each group of registers is triggered, passing a group of data stored therein to the next operation stage for calculation. At the same time, a new group of input data is input to the input 101 of the operation circuit 100, and is passed to the first operation stage via the first group of registers to start calculation; and a new group of output data is output from the output 102 of the operation circuit 100 via the last group of registers. That is, the clock is used to trigger registers, feed input data, and extract output data.

[0009] When the register is triggered, the signal at its input should be stable and can be passed back by the register. Therefore, the period of the clock is limited by the computational delay of each operation level, that is, the clock period should be greater than or equal to the computational delay of each operation level. Generally speaking, the clock period is selected to be substantially equal to the computational delay of each operation level.

[0010] For the operation circuit 100, the register delay (e.g., Ck2q delay when the register is a latch), the clock tree delay, etc. are generally much smaller than the calculation delay of the combinational logic module. Therefore, the clock period can be selected to be substantially equal to the calculation delay of the combinational logic module of each operation stage.

[0011] Therefore, the throughput and computing power of the operation circuit 100 for executing the hash algorithm are determined by the clock frequency used for the registers, that is, by the computational delay of the combinatorial logic module of each operation stage.

[0012] However, it is desired to improve the calculation frequency and throughput of the operation circuit 100 without reducing the calculation delay of the combinational logic module of each operation level, thereby reducing the power consumption and computing power ratio. Therefore, there is a demand for new technologies. Summary of the invention

[0013] One of the objectives of the present disclosure is to provide an operation circuit for executing a hash algorithm.

[0014] According to one aspect of the present disclosure, there is provided an operation circuit for executing a hash algorithm, characterized in that the operation circuit comprises a plurality of operation stages arranged in a pipeline structure, wherein each operation stage comprises: a group of inputs and a group of outputs, wherein the group of inputs is correspondingly coupled to a group of outputs of a previous operation stage, and the group of outputs is correspondingly coupled to a group of inputs of a subsequent operation stage; a plurality of combinational logic modules, wherein the input of each combinational logic module is coupled to at least a portion of the group of inputs; a plurality of delay modules, wherein the input of each delay module is coupled to one of the group of inputs, and the output is coupled to one of the group of outputs that is not coupled to the combinational logic module, so that the outputs of the group of outputs that are not coupled to the combinational logic module are each coupled to a delay module; and a plurality of supplementary delay modules, wherein the input of each supplementary delay module is coupled to the output of the corresponding combinational logic module, and the output is coupled to one of the group of outputs, wherein each of the delay modules and the supplementary delay modules of each operation stage is composed of the same delay units connected in series, and is configured so that the computational delay from the group of inputs to each of the group of outputs of each operation stage is substantially equal.

[0015] According to another aspect of the present disclosure, a chip is provided, which includes the computing circuit as described above.

[0016] According to yet another aspect of the present disclosure, a computing device is provided, which includes the chip as described above.

[0017] Other features and advantages of the present disclosure will become more apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0019] The present disclosure may be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:

[0020] Figure 1 A schematic diagram of an operation circuit for executing a hash algorithm according to the prior art is shown.

[0021] Figure 2 A schematic diagram of an operation circuit for executing a hash algorithm according to one or more exemplary embodiments of the present disclosure is shown.

[0022] Figure 3 Shows Figure 2 A schematic diagram of an operational stage in the operational circuit shown.

[0023] Figure 4 Shows Figure 2 The timing diagram of the operation circuit shown executing the hash algorithm.

[0024] Note that in the embodiments described below, sometimes the same reference numerals are used in common between different drawings to represent the same parts or parts with the same functions, and their repeated descriptions are omitted. In some cases, similar numbers and letters are used to represent similar items, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0025] For ease of understanding, the position, size, range, etc. of each structure shown in the drawings and the like may not represent the actual position, size, range, etc. Therefore, the present disclosure is not limited to the position, size, range, etc. disclosed in the drawings and the like. DETAILED DESCRIPTION

[0026] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0027] The following description of at least one exemplary embodiment is in fact merely illustrative and is in no way intended to limit the present disclosure and its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in the present disclosure. However, those skilled in the art will appreciate that they merely illustrate exemplary ways of the present disclosure that can be implemented, rather than exhaustive ways. In addition, the drawings need not be drawn to scale, and some features may be enlarged to illustrate the details of specific components.

[0028] Technologies, methods, and apparatus known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, such technologies, methods, and apparatus should be considered part of the authorization specification.

[0029] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0030] Figure 2 FIG. 2 is a schematic diagram of an operation circuit 200 for executing a hash algorithm according to one or more exemplary embodiments of the present disclosure. The operation circuit 200 may be used to execute a SHA-256 algorithm.

[0031] like Figure 2As shown, the operation circuit 200 includes N operation stages (N is a positive integer) arranged in a pipeline structure, wherein each operation stage includes: a group of inputs and a group of outputs, multiple combinational logic modules 211, 212, 213, multiple delay modules 230, and multiple supplementary delay modules 221, 222, 223.

[0032] For ease of understanding, Figure 2 The connection relationship between each input, output and each combination logic module 211, 212, 213 of each operation stage in FIG. 2 is schematically drawn corresponding to the SHA-256 algorithm. For the sake of clarity, the connection relationship between each input, output and each combination logic module 211, 212, 213 is drawn only in the first operation stage.

[0033] For example, the first operation stage includes a set of inputs 201-1 and a set of outputs 202-1, wherein the inputs 201-1 and the outputs 202-1 each include 24 data, which correspond to the Figure 1 Data stored in the eight cache registers A, B, C, D, E, F, G, H and the 16 extension registers W0, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14, W15 in the operation circuit 100 shown. For ease of understanding, the numbers of the registers corresponding to the various data in the prior art are schematically indicated at each group of inputs and outputs.

[0034] The first operation stage further includes a plurality of combinational logic modules 211, 212, 213, each of which has an input coupled to at least a portion of the set of inputs 201-1. For example, the input of the combinational logic module 213 is coupled to inputs labeled W0, W1, W9, and W14 in the set of inputs 201-1. The configuration and function of the combinational logic modules 211, 212, and 213 in the operation circuit 200 are similar to those in FIG. Figure 1 The configurations and functions of the combinational logic modules 111 , 112 , and 113 in the operation circuit 100 shown correspond to each other.

[0035] In addition, the first operation stage further includes a plurality of delay modules 230 and a plurality of supplementary delay modules 221 , 222 , 223 .

[0036] The input of each delay module 230 is coupled to one of the group of inputs 201-1, and the output is coupled to one of the group of outputs 202-1 that is not coupled to the combinational logic module, so that the outputs of the group of outputs 202-1 that are not coupled to the combinational logic module are each coupled to a delay module. For example, Figure 2The input of the top delay module 230 is coupled to the input labeled A in the input 201-1, and the output is coupled to the output labeled B in the output 202-1. Figure 2 In the output 202-1, the outputs labeled B, C, D, F, G, H, W0, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, and W14 are not coupled to the combinational logic module, and each is coupled to a delay module 230.

[0037] The input of each supplementary delay module 221, 222, 223 is coupled to the output of the corresponding combinational logic module 211, 212, 213, and the output is coupled to one of a group of outputs 202-1. For example, the input of the supplementary delay modules 221, 222, 223 is respectively coupled to the output of the combinational logic modules 211, 212, 213, and the output is respectively coupled to the outputs labeled A, E and W15 in the output 202-1.

[0038] exist Figure 2 In the embodiment shown, preferably, the number of supplementary delay modules of each operation stage is equal to the number of combinational logic modules, so that each of each group of outputs is coupled to one of the delay module and the supplementary delay module. In other embodiments, the number of supplementary delay modules of each operation stage may be less than the number of combinational logic modules.

[0039] Figure 3 Shows Figure 2 FIG. 1 is a schematic diagram of an operation stage 300 in the operation circuit 200 .

[0040] like Figure 3 As shown, the operation stage 300 includes: a group of inputs 301 and a group of outputs 302 , a plurality of combinational logic modules 311 , 312 , 313 , a delay module 330 , and supplementary delay modules 322 , 323 .

[0041] The delay module 330 and the supplementary delay modules 322 and 323 are all composed of the same delay unit 340 connected in series. Figure 3 In the illustrated embodiment, the supplementary delay modules 322 and 323 are respectively composed of one delay unit 340 and three delay units 340 connected in series, and each of the delay modules 330 is composed of M delay units 340 connected in series (M is a positive integer).

[0042] By using the same delay units 340 connected in series to form the delay module 330 and the supplementary delay modules 322 and 323, the delay errors between the delay units 340 can be appropriately offset, so that the delays of the delay modules 330 and the supplementary delay modules 322 and 323 are more accurate. The delay errors between the delay units 340 are caused by various factors (e.g., process, temperature, etc.) during the manufacturing, installation, and operation of the delay units 340.

[0043] In a preferred embodiment, each delay unit 340 may be composed of a buffer or a pair of inverters. In other embodiments, the delay unit 340 may be composed of one or more elements capable of implementing a delay function.

[0044] The delay module and the supplementary delay module of each operation stage should be configured so that the computational delay from a set of inputs to each of a set of outputs of each operation stage is substantially equal. That is, the delay module 330 and the supplementary delay modules 322 and 323 in the operation stage 300 should be configured so that the computational delays from a set of inputs 301 to a set of outputs 302 labeled A, B, C, D, E, F, G, H, WO, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14, and W15 are substantially equal.

[0045] As mentioned above, the register delay, clock tree delay, etc. are much smaller than the calculation delay of the combinational logic module. In other words, the delay module 330 and the supplementary delay modules 322 and 323 in the operation stage 300 should be configured so that the following are substantially equal:

[0046] 1. The computation delay from input 301 to the output labeled A in output 302, i.e., the sum of the computation delays of combinatorial logic modules 311 and 312;

[0047] 2. The computation delay from input 301 to the output labeled E in output 302, i.e., the sum of the computation delays of the combinational logic module 312 and one delay unit 340;

[0048] 3. The computation delay from input 301 to the output labeled W15 in output 302, i.e., the sum of the computation delays of the combinational logic module 313 and the three delay units 340;

[0049] 4. The computational delay of the outputs from input 301 to output 302 labeled as others (B, C, D, F, G, H, W0, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14), that is, the sum of the computational delays of the M delay units 340.

[0050] Those skilled in the art should understand that Figure 3 The number and configuration of the delay module 330 and the supplementary delay modules 322 and 323 are exemplary and can be adjusted accordingly according to the hash algorithm executed by the operation circuit 300 and the specific configuration of the chip.

[0051] exist Figure 3 In the illustrated embodiment, the number of supplementary delay modules 322 and 323 is less than the number of combinatorial logic modules, and the output labeled A is not coupled to the supplementary delay module, but is directly coupled to the combinatorial logic module 311. In some embodiments, an output with the longest computational delay of the corresponding combinatorial logic module in a group of outputs may not be coupled to the supplementary delay module, but may be directly coupled to the corresponding combinatorial logic module. In other words, the computational delay of the output (A) with the longest computational delay from the input 301 to the corresponding combinatorial logic module in the output 302 is directly determined as the computational delay of the operation stage 300, and the computational delays to other outputs (B, C, ..., W15) in the output 302 are completed by the delay module 330 and the supplementary delay modules 322 and 323. The advantage of such an embodiment is that no additional computational delay is introduced, so that the overall computational delay of the operation stage 300 is minimized.

[0052] In such an embodiment, the number of delay units 340 included in the delay module 330 and the supplementary delay modules 322 and 323 can be determined according to the need to complete the calculation delay. Figure 3 In the illustrated embodiment, in order to make up for the difference between the computational delay from input 301 to the output labeled E in output 302 and the computational delay from input 301 to the output labeled A in output 302, that is, in order to make up for the computational delay of the combinational logic module 311, the supplementary delay module 322 is set to be composed of one delay unit 340.

[0053] In other embodiments, the number of supplementary delay modules may be equal to the number of combinational logic modules, and the number of delay units 340 included in the delay module 330 and the supplementary delay modules 322 and 323 may also be determined in combination with other factors. For example, in order to better offset the delay errors between the delay units 340, the number of delay units 340 included in the delay module 330 and the supplementary delay modules 322 and 323 may be appropriately increased. However, considering the manufacturing cost and power consumption of the chip, the number of delay units 340 should not be too large.

[0054] In a preferred embodiment, the number M of the delay units 340 included in the delay module 330 may be greater than or equal to 10 and less than or equal to 20. In a further preferred embodiment, M may be greater than or equal to 12 and less than or equal to 18.

[0055] It should be noted that the expression "substantially equal" herein means that the two are roughly equal within a certain error, but not necessarily strictly and precisely equal. For example, "substantially equal" means that the two are roughly equal within a 2% error. Preferably, the two are roughly equal within a 1% error. In some contexts, the error may be about 5%. Those skilled in the art should understand that this is in line with technical principles and engineering practice.

[0056] The computation delay from a set of inputs to each of a set of outputs of each operation stage is substantially equal, which enables data to be passed in sequence along each operation stage in a timely manner without the need for register triggering. In other words, the operation circuit 200 of the present disclosure does not require the cache registers and extension registers (i.e., Figure 1 The operation circuit 100 shown includes 8 cache registers A, B, C, D, E, F, G, H and 16 extension registers W0, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14, W15).

[0057] In addition, as described above, the cycle of the clock used to trigger the register, feed the input data, and extract the output data in the prior art should be greater than or equal to the calculation delay of each operation level. However, the clock cycle used to feed the input data and extract the output data in the operation circuit 200 of the present disclosure does not need to be greater than or equal to the calculation delay of the combinational logic module of each operation level. Therefore, the calculation frequency and throughput of the operation circuit 200 of the present disclosure are not limited by the calculation delay of the combinational logic module of each operation level.

[0058] Figure 4 Shows Figure 2 The timing diagram of the operation circuit 200 executing the hash algorithm is shown.

[0059] like Figure 4 As shown, the clock CLK is used to feed input data to the input 201 - 1 of the operation circuit 200 . The period of the clock CLK is T. At each rising edge of the clock CLK, a new set of input data is fed to the input 201 - 1 of the operation circuit 200 .

[0060] Those skilled in the art will appreciate that, by way of example only, Figure 4 Each set of input data in is fed to the input 201-1 of the operation circuit 200 at the rising edge of the clock CLK. In other embodiments, the input data may also be fed to the input 201-1 of the operation circuit 200 at the falling edge of the clock CLK.

[0061] As described above, the period T of the clock CLK of the operation circuit 200 does not need to be greater than or equal to the calculation delay of each operation stage. Alternatively, the period T of the clock CLK can be less than the calculation delay of each operation stage, so that the calculation frequency and throughput of the operation circuit 200 are increased, thereby improving the computing power of the operation circuit 200 and reducing the power consumption to computing power ratio.

[0062] In a preferred embodiment, the calculation delay of each operation stage may be substantially equal to k times the period T of the clock CLK, where k is an integer greater than or equal to 2. This allows each operation stage to accommodate exactly k sets of data when the operation circuit 200 is operating.

[0063] On the basis that the computation delay of each operation level is basically determined, increasing the value of k is beneficial to improving the throughput of the operation circuit 200 and reducing its power consumption to computing power ratio. However, when the value of k is large, the negative impact of the delay error between each delay unit 340 will also become larger, which increases the risk of delay misalignment and data confusion in each operation level. Preferably, k can be selected as 2 or 3.

[0064] In order to control the negative impact of the delay error between the delay units 340, M can be preferably selected to be 3 to 10 times of k. Further preferably, M can be selected to be 4 to 8 times of k. Further preferably, M can be selected to be 5 to 7 times of k.

[0065] Figure 4 The timing diagram of the operation circuit 200 executing the hash algorithm when k is 2 is exemplarily shown.

[0066] exist Figure 4 In the illustrated embodiment, the computation delay of each operation stage is 2T. In other words, the computation delay from a set of inputs to each of a set of outputs of each operation stage of the operation circuit 200 is 2T.

[0067] That is, in each operation stage of the operation circuit 200, the sum of the calculation delays of the combinational logic modules 211, 212 and the supplementary delay module 221 (i.e., the calculation delay from a set of inputs of each operation stage to an output labeled A in a set of outputs), the sum of the calculation delays of the combinational logic module 212 and the supplementary delay module 222 (i.e., the calculation delay from a set of inputs of each operation stage to an output labeled E in a set of outputs), and the sum of the calculation delays of the combinational logic module 213 and the supplementary delay module 224 (i.e., the calculation delay from a set of inputs of each operation stage to an output labeled E in a set of outputs). The sum of the computational delays of 3 (i.e., the computational delay from a set of inputs of each operation stage to a set of outputs labeled W15), and the computational delay of delay module 230 (i.e., the computational delay from a set of inputs of each operation stage to a set of outputs labeled others (B, C, D, F, G, H, WO, W1, W2, W3, W4, W5, W6, W7, W8, W9, W10, W11, W12, W13, W14)) are both 2T.

[0068] like Figure 4 As shown, at t=0, at the first rising edge of the clock CLK, the first set of data (data 1) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then is transferred to the combinational logic modules 211, 212, 213 of the first operation stage, the delay module 230, and the supplementary delay modules 221, 222, 223. After a calculation delay of 2T, at t=2T, data 1 arrives at the output 202-1 of the first operation stage, and is further transferred to the input 201-2 of the second operation stage.

[0069] Afterwards, data 1 is passed to the combinational logic modules 211, 212, 213 and the delay module 230 and the supplementary delay modules 221, 222, 223 of the second operation stage, and also undergoes a calculation delay of 2T. At t=4T, data 1 arrives at the output 202-2 of the second operation stage and is further passed to the input 201-3 of the third operation stage.

[0070] After that, after the same calculation delay of 2T, at t=6T, data 1 arrives at the output 202-3 of the third operation stage and is further transmitted to the input 201-4 of the fourth operation stage.

[0071] In addition, at t=T, at the second rising edge of the clock CLK, the second set of data (data 2) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then transferred to the combinational logic modules 211, 212, 213 of the first operation stage, the delay module 230, and the supplementary delay modules 221, 222, 223. Between t=T and t=2T, both data 1 and data 2 are accommodated in the first operation stage of the operation circuit 200. After a calculation delay of 2T, at t=3T, data 2 arrives at the output 202-1 of the first operation stage, and is further transferred to the input 201-2 of the second operation stage.

[0072] Afterwards, data 2 is transferred to the combinational logic modules 211, 212, 213, delay module 230, and supplementary delay modules 221, 222, 223 of the second operation stage. Between t=3T and t=4T, data 1 and data 2 are both accommodated in the second operation stage of the operation circuit 200. After a calculation delay of 2T, at t=5T, data 2 arrives at the output 202-2 of the second operation stage and is further transferred to the input 201-3 of the third operation stage. Between t=5T and t=6T, data 1 and data 2 are both accommodated in the third operation stage of the operation circuit 200.

[0073] In addition, at t=2T, at the third rising edge of the clock CLK, the third set of data (data 3) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then transferred to the combinational logic modules 211, 212, 213 of the first operation stage, the delay module 230, and the supplementary delay modules 221, 222, 223. Between t=2T and t=3T, both data 2 and data 3 are accommodated in the first operation stage of the operation circuit 200. After a calculation delay of 2T, at t=4T, data 3 arrives at the output 202-1 of the first operation stage, and is further transferred to the input 201-2 of the second operation stage.

[0074] Afterwards, data 3 is transferred to the combinational logic modules 211, 212, 213, delay module 230, and supplementary delay modules 221, 222, 223 of the second operation stage. Between t=4T and t=5T, data 2 and data 3 are both accommodated in the second operation stage of the operation circuit 200. After a calculation delay of 2T, at t=6T, data 3 arrives at the output 202-2 of the second operation stage and is further transferred to the input 201-3 of the third operation stage.

[0075] In addition, at t=3T, at the fourth rising edge of the clock CLK, the fourth group of data (data 4) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then transferred to the combinational logic modules 211, 212, 213 and the delay module 230 and the supplementary delay modules 221, 222, 223 of the first operation stage. Between t=3T and t=4T, both data 3 and data 4 are accommodated in the first operation stage of the operation circuit 200. After a calculation delay of 2T, at t=5T, data 4 arrives at the output 202-1 of the first operation stage and is further transferred to the input 201-2 of the second operation stage.

[0076] Afterwards, data 4 is transmitted to the combinatorial logic modules 211, 212, 213 and the delay module 230 and the supplementary delay modules 221, 222, 223 of the second operation stage. Between t=5T and t=6T, data 3 and data 4 are both accommodated in the second operation stage of the operation circuit 200.

[0077] In addition, at t=4T, at the fifth rising edge of the clock CLK, the fifth set of data (data 5) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then transferred to the combinational logic modules 211, 212, 213 of the first operation stage, the delay module 230, and the supplementary delay modules 221, 222, 223. Between t=4T and t=5T, both data 4 and data 5 are accommodated in the first operation stage of the operation circuit 200. After a calculation delay of 2T, at t=6T, data 5 arrives at the output 202-1 of the first operation stage, and is further transferred to the input 201-2 of the second operation stage.

[0078] In addition, at t=5T, at the sixth rising edge of the clock CLK, the sixth group of data (data 6) is fed to the input 201-1 of the first operation stage of the operation circuit 200, and then transferred to the combinational logic modules 211, 212, 213 and the delay module 230 and the supplementary delay modules 221, 222, 223 of the first operation stage. Between t=5T and t=6T, both data 5 and data 6 are accommodated in the first operation stage of the operation circuit 200.

[0079] It can be seen that when the operation circuit 200 works normally, each operation stage can accommodate k groups of data, that is, N operation stages can simultaneously calculate k*N groups of data. In contrast, the operation circuit 100 including N operation stages in the prior art can only simultaneously calculate N groups of data. This is one of the significant advantages of the present invention over the prior art.

[0080] The operation circuit according to the present disclosure may be implemented in various appropriate ways, such as software, hardware, a combination of software and hardware, etc. In one implementation, a chip may include the operation circuit as described above, and the chip may also be included in a computing device.

[0081] The words "front", "rear", "top", "bottom", "over", "under", etc., if any, in the specification and claims are used for descriptive purposes and are not necessarily used to describe invariant relative positions. It is understood that the words so used are interchangeable under appropriate circumstances such that the embodiments of the disclosure described herein, for example, are capable of operation in other orientations than those illustrated or otherwise described herein.

[0082] As used herein, the word "exemplary" means "serving as an example, instance, or illustration," rather than as a "model" to be exactly copied. Any implementation described as an example herein is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, the present disclosure is not limited by any expressed or implied theory given in the above technical field, background technology, summary of the invention, or detailed description.

[0083] As used herein, the term "substantially" is meant to include any minor variations due to design or manufacturing imperfections, device or component tolerances, environmental influences, and / or other factors. The term "substantially" also allows for deviations from a perfect or ideal condition due to parasitic effects, noise, and other practical considerations that may be present in actual implementations.

[0084] Additionally, the foregoing description may have referred to elements or nodes or features being "connected" or "coupled" together. As used herein, unless expressly stated otherwise, "connected" means that one element / node / feature is directly connected (or directly communicates) with another element / node / feature, electrically, mechanically, logically, or otherwise. Similarly, unless expressly stated otherwise, "coupled" means that one element / node / feature can be mechanically, electrically, logically, or otherwise connected to another element / node / feature in a direct or indirect manner to allow interaction, even though the two features may not be directly connected. That is, "coupled" is intended to encompass both direct and indirect connections of elements or other features, including connections utilizing one or more intermediate elements.

[0085] In addition, the terms "first", "second", and the like may also be used herein for reference purposes only, and thus are not intended to be limiting. For example, the terms "first", "second", and other such numerical terms referring to structures or elements do not imply a sequence or order unless the context clearly indicates otherwise.

[0086] It should also be understood that when the term “include / comprises” is used in this document, it indicates the presence of the specified features, integers, steps, operations, units and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, units and / or components and / or their combinations.

[0087] In this disclosure, the term "provide" is used in a broad sense to cover all ways of obtaining an object, and thus "providing an object" includes but is not limited to "purchasing", "preparing / manufacturing", "arranging / setting up", "installing / assembling", and / or "ordering" an object, etc.

[0088] Those skilled in the art will appreciate that the boundaries between the above operations are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed in additional operations, and operations can be performed at least partially overlapping in time. Moreover, alternative embodiments may include multiple instances of specific operations, and the order of operations may be changed in other various embodiments. However, other modifications, variations, and replacements are equally possible. Therefore, this specification and accompanying drawings should be considered illustrative, not restrictive.

[0089] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. The various embodiments disclosed herein may be combined in any manner without departing from the spirit and scope of the present disclosure. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A computing circuit for executing a hash algorithm, characterized in that: The operation circuit comprises a plurality of operation stages arranged in a pipeline structure, wherein each operation stage comprises: a set of inputs and a set of outputs, the set of inputs being correspondingly coupled to a set of outputs of a preceding computing stage, and the set of outputs being correspondingly coupled to a set of inputs of a succeeding computing stage; a plurality of combinatorial logic modules, each combinatorial logic module having an input coupled to at least a portion of the set of inputs; a plurality of delay modules, each having an input coupled to one of the set of inputs and an output coupled to one of the set of outputs that is not coupled to the combinatorial logic module, such that the outputs of the set of outputs that are not coupled to the combinatorial logic module are each coupled to one delay module; and A plurality of supplementary delay modules, each of which has an input coupled to the output of a corresponding combinational logic module and an output coupled to one of the set of outputs, wherein: Each of the delay modules and the supplementary delay modules of each operation stage is composed of identical delay units connected in series and is configured to make the computational delays from the set of inputs to each of the set of outputs of each operation stage substantially equal.

2. The operation circuit according to claim 1, characterized in that: The computational delay of each operation stage is substantially equal to k times the period of a clock used to feed input data to the set of inputs, where k is an integer greater than or equal to 2.

3. The operation circuit according to claim 2, characterized in that: Each delay module is composed of M delay units connected in series, where M is a multiple of k.

4. The operation circuit according to claim 2, characterized in that: k is 2 or 3.

5. The operation circuit according to claim 3, characterized in that: M is greater than or equal to 10 and less than or equal to 20.

6. The operation circuit according to claim 3, characterized in that: M is 3 to 10 times greater than k.

7. The operation circuit according to any one of claims 1 to 6, characterized in that: Each delay unit is composed of a buffer or a pair of inverters.

8. The operation circuit according to any one of claims 1 to 6, characterized in that: The number of the supplementary delay modules of each operation stage is equal to the number of the combinatorial logic modules, so that each of the set of outputs is coupled to one of the delay module and the supplementary delay module.

9. The operation circuit according to any one of claims 1 to 6, characterized in that: The operation circuit is used to execute the SHA256 algorithm.

10. A chip, characterized in that: The chip comprises the operation circuit according to any one of claims 1-9.

11. A computing device, characterized in that: The computing device comprises the chip according to claim 10.

Citation Information

Patent Citations

  • Operating circuit, chip and computing device for executing hash algorithm

    CN212411183U