Arithmetic core circuit and computing chip

By introducing an asynchronous FIFO module and clock signals of different frequencies into the hash operation circuit, the problem of excessive clock signal propagation distance is solved, improving the operation performance and processing speed. It is suitable for operation core circuits with horizontal and vertical structures.

CN114528247BActive Publication Date: 2025-10-24SHENZHEN MICROBT ELECTRONICS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011322084.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-23
Publication Date
2025-10-24
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

In existing hash operation circuits, the long propagation distance of the clock signal causes the signal shape to degrade, affecting the operation performance. In addition, the long data transmission time in the vertical structure limits the processing speed.

Method used

An asynchronous first-in-first-out (FIFO) module is used to transfer data between the computing levels, and two clock signals of different frequencies are used to provide clock signals to the asynchronous FIFO module and the computing level respectively, thereby reducing the clock signal propagation path and improving the clock signal shape.

Benefits of technology

The performance and processing speed of the operation core circuit are improved, the distortion of the clock signal at each operation level is reduced, and the design is suitable for the operation core circuit of horizontal and vertical structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528247B_ABST
    Figure CN114528247B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an operation core circuit and a computing chip. An operation core circuit includes: an input module configured to receive a data block; an operation module configured to perform a hash operation on the data block, including a plurality of operation stages arranged in a pipeline structure, such that a data signal based on the data block is sequentially transmitted along the plurality of operation stages; an asynchronous FIFO module arranged between adjacent first and second operation stages, configured to receive the data signal output from the first operation stage using a first clock signal and output the data signal to the second operation stage using a second clock signal, the first operation stage preceding the second operation stage; a first clock module configured to provide the asynchronous FIFO module and the first operation stage and its preceding operation stages with the first clock signal; and a second clock module configured to provide the asynchronous FIFO module and the second operation stage and its subsequent operation stages with the second clock signal, wherein the first clock signal and the second clock signal have the same frequency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an operation core circuit for performing a hash operation, and more particularly, to an operation core circuit and a computing chip. BACKGROUND

[0002] SHA (Secure Hash Algorithm) series algorithms are published by the National Institute of Standards and Technology of the United States, in which the SHA-256 algorithm is a secure hash algorithm with a hash length of 256 bits. SUMMARY

[0003] According to a first aspect of the present disclosure, there is provided an operation core circuit, comprising: an input module configured to receive a data block; an operation module configured to perform a hash operation on the received data block, the operation module comprising a plurality of operation stages arranged in a pipeline structure such that a data signal based on the data block is sequentially passed along the plurality of operation stages, each operation stage of the plurality of operation stages performing an operation on a data signal received from a previous operation stage and providing a data signal operated by the operation stage to a next operation stage; an asynchronous first-in-first-out (FIFO) module provided between a first operation stage and a second operation stage of the plurality of operation stages, the first operation stage preceding the second operation stage, the asynchronous FIFO module being configured to receive a data signal output from the first operation stage using a first clock signal and output the received data signal to the second operation stage using a second clock signal different from the first clock signal; a first clock module configured to provide the first clock signal to the asynchronous FIFO module and to the first operation stage and operation stages of the plurality of operation stages preceding the first operation stage; and a second clock module configured to provide the second clock signal to the asynchronous FIFO module and to the second operation stage and operation stages of the plurality of operation stages succeeding the second operation stage, wherein a frequency of the first clock signal is the same as a frequency of the second clock signal.

[0004] In some embodiments, a passing direction of both the first clock signal and the second clock signal is the same as a passing direction of the data signal.

[0005] In some embodiments, a passing direction of both the first clock signal and the second clock signal is opposite to a passing direction of the data signal.

[0006] According to a second aspect of the present disclosure, there is provided a computing chip comprising one or more operation core circuits as described above.

[0007] According to a third aspect of the present disclosure, there is provided a computing chip comprising a plurality of operation core circuits as previously described, the plurality of the operation core circuits being arranged in a plurality of columns, the first clock module of each column of operation core circuits receiving a first clock signal via a common clock channel, and the second clock module of each column of operation core circuits receiving a second clock signal via a common clock channel.

[0008] According to a fourth aspect of the present disclosure, there is provided a computing chip comprising a plurality of operation core circuits arranged in a plurality of columns, each operation core circuit comprising: an input module configured to receive a data block; an operation module configured to perform a hash operation on the received data block, the operation module comprising a plurality of operation stages arranged in a pipeline structure such that a data signal based on the data block is sequentially passed along the plurality of operation stages, each operation stage of the plurality of operation stages performing an operation on a data signal received from a previous operation stage and providing a data signal operated by the operation stage to a subsequent operation stage; and a clock module configured to provide a clock signal to the plurality of operation stages, wherein the plurality of columns comprises a first column of operation core circuits and a second column of operation core circuits arranged adjacent to each other and in the stated order, the clock module of the first column of operation core circuits and the clock module of the second column of operation core circuits receiving a clock signal via a common clock channel.

[0009] Other features of the present disclosure, its nature and advantages will become more apparent from the detailed description of exemplary embodiments of the present disclosure below. BRIEF DESCRIPTION OF DRAWINGS

[0010] The included drawings are for illustrative purposes and are in no way limiting on this disclosure. These drawings are merely examples of possible structures and arrangements for the inventive devices and methods disclosed herein. These drawings are in no way limiting of any change in form and detail of the application that can come to be made during the pendency of this application, nor do they serve to limit the present disclosure to precisely as shown and described. The embodiments described and illustrated herein are best understood with the drawings in conjunction with the detailed description below.

[0011] Figures 1-3 is a schematic diagram of an operation core circuit according to some embodiments of the present disclosure.

[0012] Figures 4-6 is a schematic diagram of an operation core circuit comprising a hash engine for performing SHA-256 algorithm (hereinafter, can be referred to as SHA-256 hash engine) according to some embodiments of the present disclosure.

[0013] Figure 7A and Figure 7B is a schematic diagram of an operation core circuit having a vertical structure according to embodiments of the present disclosure.

[0014] Figure 8 is a schematic diagram of a computing chip according to some embodiments of the present disclosure.

[0015] Figures 9A-9D is a schematic layout for distributing clock signals to arithmetic core circuits in a computing chip according to some embodiments of the present disclosure.

[0016] Figures 10A-10D is a schematic layout for distributing clock signals to arithmetic core circuits having a vertical structure in a computing chip according to some embodiments of the present disclosure.

[0017] Figure 11A and Figure 11B is a schematic diagram of a computing chip according to some embodiments of the present disclosure.

[0018] Figure 12 is a schematic diagram of an exemplary pipeline structure for performing the SHA-256 algorithm.

[0019] Note that, in the following embodiments, the same reference numbers are sometimes used in different drawings to denote the same or similar parts or parts having the same function, and repeated description thereof is omitted. In this specification, like numbers and letters designate like items, and once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0020] For ease of understanding, the positions, sizes, ranges, and the like of the structures shown in the drawings and the like are sometimes not actual positions, sizes, ranges, and the like. Therefore, the disclosed application is not limited to the positions, sizes, ranges, and the like disclosed in the drawings and the like. Further, the drawings are not necessarily drawn to scale, and some features can be exaggerated to show details of specific components. DETAILED DESCRIPTION

[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the drawings. It should be noted that the relative arrangement, numerical expressions, and numerical values of the components and steps set forth in these embodiments are not limiting to the scope of the present disclosure unless specifically stated otherwise.

[0022] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way limiting to the scope of the disclosure and its applications or uses. That is, the hash engine herein is shown in an exemplary manner to illustrate different embodiments of circuits in the present disclosure, and is not intended to be limiting. Those skilled in the art will appreciate that they merely explain illustrative ways in which the present application can be implemented, and are not exhaustive.

[0023] Techniques, methods, and apparatus known to those of ordinary skill in the relevant art(s) can not be discussed in any detail since the techniques, methods, and apparatus should be considered part of the state of the art.

[0024] A computing chip generally includes a top module and an operation core circuit. The top module is used to perform communication functions, control functions, input / output (IO) functions, clock PLL functions, and the like. The operation core circuit is used for core computing operations. The operation core circuit obtains operation tasks from the top module and feeds back operation results to the top module. For double hash, a complete calculation generally requires two rounds of 64 cycles (performing two SHA-256 algorithms), i.e., 128-tap operations. Some optimization methods can reduce several taps (e.g., 6 taps) of operations. In embodiments according to the present disclosure, the operation of performing two SHA-256 algorithms (i.e., 128-tap operations) by the operation core circuit is mainly taken as an example for illustration, but those skilled in the art can understand that the present disclosure is not limited thereto and can be applied to any number of taps of operations. The SHA-256 algorithm mentioned herein includes any version of the SHA-256 algorithm known in the art and variants and modifications thereof.

[0025] In the present disclosure, in order to improve the operation throughput, the operation core circuit can be configured to have a plurality of operation stages arranged in a pipeline structure. Figure 12 An exemplary pipeline structure for performing the SHA-256 algorithm is schematically shown, which includes 64 operation stages, each having 8 compression registers A-H and 16 extension registers 0-15. The 1st operation stage can receive an input data block and divide it into 8 32-bit data to be stored in the compression registers A-H, and then perform operation processing and provide it to the 2nd operation stage. Thereafter, each operation stage performs operation on the operation result of the previous operation stage received thereby and provides its own operation result to the next operation stage. Finally, after operation by the 64 operation stages, the operation core circuit can output the hash operation result of performing the SHA-256 algorithm once on the input data block. In this way, when all the operation stages in the pipeline structure are fully loaded (i.e., all the operation stages receive data and perform operation processing), the operation core circuit can output one operation result per tap. For example, for 128-tap operations, the operation core circuit of the present disclosure can include a pipeline structure having a total of 128 operation stages, thereby greatly improving the operation throughput.

[0026] When designing the operations of the operation core circuit in a pipeline structure, a clock signal needs to be provided to each operation stage in the pipeline structure. In one case, the transmission direction of the clock signal can be the same as the transmission direction of the data signal in the pipeline structure, i.e., from the first operation stage to the last operation stage in the pipeline structure, in which case the clock period can be smaller, and accordingly the chip frequency can be faster, achieving higher performance. In another case, the transmission direction of the clock signal can be opposite to the transmission direction of the data signal in the pipeline structure, i.e., from the last operation stage to the first operation stage in the pipeline structure, in which case it is easier to meet the hold time of the registers at each operation stage in the pipeline structure, so that the data can be stably punched into the registers. In both cases, the clock signal needs to pass through each operation stage of the pipeline structure, and the number of transmission stages of the clock signal is usually as many as 128 stages. However, the farther the clock signal propagates, the greater the distortion of the rising edge and / or the falling edge of the clock signal, resulting in the deterioration of the shape of the clock signal and the worse duty cycle. When the clock signal propagates along the transmission direction of the data signal to the operation stage located downstream of the pipeline structure (e.g., the 128th operation stage) or propagates in the opposite direction to the operation stage located upstream of the pipeline structure (e.g., the 1st operation stage), the level of the clock signal can not meet the minimum pulse requirement of the registers of the current operation stage, thereby affecting the operation performance of the entire operation core circuit.

[0027] In the operation core circuit according to the embodiments of the present disclosure, by introducing an asynchronous first-in-first-out (FIFO) module between the operation stages, the number of operation stages that the clock signal needs to pass through can be greatly reduced, thereby the shape of the clock signal at each operation stage can be significantly improved, and thus the performance of the operation core circuit is advantageously improved. Asynchronous FIFO refers to a FIFO design in which data values are written to a FIFO buffer from one clock domain and data values are read from the same FIFO buffer from another clock domain, and the two clock domains are asynchronous to each other. Asynchronous FIFO can be used to safely transfer data from one clock domain to another clock domain.

[0028] The operation core circuit according to the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In these drawings, the dashed arrows are used to indicate the transmission direction of the data signal, and the solid arrows are used to indicate the transmission direction of the clock signal. It should be noted that the actual operation core circuit can also include additional components, and in order to avoid obscuring the points of the present disclosure, these additional components are not shown in the drawings and are not discussed in the present disclosure.

[0029] Figure 1The schematic diagram shows a computing core circuit 100A according to an embodiment of the present disclosure. The computing core circuit 100A may include an input module 110, a computing module 120, an asynchronous FIFO module 130, a first clock module 141, and a second clock module 142. The input module 110 may be configured to receive a data block. The computing module 120 may be configured to perform a hash operation on the received data block. The clock modules 141 and 142 may be configured to provide the required clock signals for the computing module 120.

[0030] like Figure 1 As shown, the operation module 120 may include a plurality of operation stages arranged in a pipeline structure so that data signals based on received data blocks are sequentially transmitted along the operation stages. Each operation stage may operate on the data signal received from the previous operation stage and provide the data signal operated by the operation stage to the next operation stage. In some examples, reference may be made to Figure 12 The operation stages in the operation module 120 are configured according to the exemplary pipeline structure of FIG. 3 . The operation stages in the operation module 120 may also be configured according to other pipeline structures known in the art or developed later.

[0031] The asynchronous FIFO module 130 may be provided between adjacent first and second operation stages 121-a, 121-b, of the plurality of operation stages of the operation module 120, with the first operation stage 121-a preceding the second operation stage 121-b. The asynchronous FIFO module may be configured to receive a data signal output from the first operation stage 121-a using a first clock signal and output the received data signal to the second operation stage 121-b using a second clock signal different from the first clock signal.

[0032] The first clock module 141 can be configured to provide a first clock signal to the asynchronous FIFO module 130 and to the first operation stage 121-a and the operation stages preceding the first operation stage 121-a among the multiple operation stages of the operation module 120. The second clock module 142 can be configured to provide a second clock signal to the asynchronous FIFO module 130 and to the second operation stage 121-b and the operation stages following the second operation stage 121-b among the multiple operation stages of the operation module 120. The frequency of the first clock signal and the frequency of the second clock signal are the same. As shown in the figure, the data signal propagates from left to right through all the operation stages of the operation module 120, while the first clock signal propagates from right to left from the first operation stage 121-a to the first operation stage, and the second clock signal propagates from right to left from the last operation stage to the second operation stage 121-b.

[0033] Thus, by introducing the asynchronous FIFO module 130 among the plurality of operation stages of the operation module 120, the data signal transmission between two adjacent operation stages can be performed by the asynchronous FIFO module 130, while two clock signals can be introduced, each of which only needs to propagate through a part of the operation stages of the operation module 120 without traversing all the operation stages of the operation module 120, so that the shape of the clock signal at each operation stage can be significantly improved, and the performance of the operation core circuit is advantageously improved. Moreover, the introduction of the asynchronous FIFO module 130 does not affect the processing speed and throughput of the entire operation core circuit, because the transmission time of the data signal between the asynchronous FIFO module 130 and the operation stage does not exceed the transmission time between the operation stages.

[0034] In Figure 1 In the illustrated embodiment, the transmission directions of both the first clock signal and the second clock signal are opposite to the transmission direction of the data signal. In other embodiments, the transmission directions of both the first clock signal and the second clock signal can be the same as the transmission direction of the data signal.

[0035] Figure 2 An operation core circuit 100B according to other embodiments of the present disclosure is schematically illustrated. The operation core circuit 100B differs from the operation core circuit 100A in that, in the operation core circuit 100B, the transmission directions of both the first clock signal and the second clock signal are the same as the transmission direction of the data signal. As Figure 2 As shown, the asynchronous FIFO module 130 is disposed between a first operation stage 121-a and a second operation stage 121-b adjacent to each other among the plurality of operation stages of the operation module 120, and receives the data signal output from the first operation stage 121-a by the first clock signal provided by the first clock module 141 and outputs the received data signal to the second operation stage 121-b by the second clock signal different from the first clock signal provided by the second clock module 142. As shown, the data signal propagates through all the operation stages of the operation module 120 in the direction from left to right, while the first clock signal propagates from the first operation stage 121-a in the direction from left to right, and the second clock signal propagates from the second operation stage 121-b in the direction from left to right. Thus, the operation core circuit design containing the asynchronous FIFO module according to the present disclosure can be applied to any transmission direction of the clock signal.

[0036] In some embodiments, the first clock module 141 and the second clock module 142 can be configured to receive the clock signal from the same clock source located outside the operation core circuit. The clock source can be used to provide a basic clock signal. That is, the first clock signal and the second clock signal can be of the same origin, but experience different paths from the clock source to the corresponding clock module.

[0037] In some embodiments, the arithmetic core circuit can include one or more asynchronous FIFO modules, which can be interposed between the arithmetic stages. In this way, the number of arithmetic stages that each clock signal needs to pass through can be further reduced. The interposition of these asynchronous FIFO modules can cause the arithmetic stages to be divided into multiple groups, and in some embodiments, the number of arithmetic stages included in each group can be the same.

[0038] For example, Figure 3 An arithmetic core circuit 100C according to the present disclosure is schematically shown. In some embodiments, the asynchronous FIFO module is a first asynchronous FIFO module 130, and the arithmetic core circuit 100C can further include a second asynchronous FIFO module 131. The second asynchronous FIFO module 131 is disposed between a third arithmetic stage 121-c and a fourth arithmetic stage 121-d of the plurality of arithmetic stages of the arithmetic module 120, the third arithmetic stage 121-c being before the fourth arithmetic stage 121-d and after the second arithmetic stage 121-b. The second asynchronous FIFO module 131 can be configured to receive a data signal output from the third arithmetic stage 121-c using a second clock signal and output the received data signal to the fourth arithmetic stage 121-d using a third clock signal that is different from the second clock signal. The arithmetic core circuit 100C can further include a third clock module 143. The third clock module 143 is configured to provide the third clock signal to the second asynchronous FIFO module 131 and to the fourth arithmetic stage 121-d and to the arithmetic stages of the plurality of arithmetic stages of the arithmetic module 120 that are after the fourth arithmetic stage 121-d. The second clock module 142 is configured to provide the second clock signal to the first asynchronous FIFO module 130 and the second asynchronous FIFO module 131 and to the second arithmetic stage 121-b and the third arithmetic stage 121-c and the arithmetic stages therebetween.

[0039] Further as Figure 3As shown, in some embodiments, the operation core circuit 100C can additionally or alternatively include a third asynchronous FIFO module 132. The third asynchronous FIFO module 132 is disposed between a fifth operation stage 121-e and a sixth operation stage 121-f of the plurality of operation stages of the operation module 120, the sixth operation stage 121-f being after the fifth operation stage 121-e and before the first operation stage 121-a. The third asynchronous FIFO module 132 is configured to receive a data signal output from the fifth operation stage 121-e with a fourth clock signal different from the first clock signal and output the received data signal to the sixth operation stage 121-f with the first clock signal. The operation core circuit 100C further includes a fourth clock module 144. The fourth clock module 144 is configured to provide the fourth clock signal to the third asynchronous FIFO module 132 and to the fifth operation stage 121-e and to the operation stages of the plurality of operation stages of the operation module 120 that are before the fifth operation stage 121-e. The first clock module 141 is configured to provide the first clock signal to the first asynchronous FIFO module 130 and the third asynchronous FIFO module 132 and to the sixth operation stage 121-f and the first operation stage 121-a and the operation stages therebetween.

[0040] Those skilled in the art can understand that, although Figure 3 Although the operation core circuit 100C is shown to include three asynchronous FIFO modules, this is merely exemplary and not limiting, and the number and location of asynchronous FIFO modules in the operation core circuit can be reasonably arranged according to actual needs. In addition, although Figure 3 The data signal and the clock signal are shown to be transmitted in opposite directions, but the same applies to the case where the data signal and the clock signal are transmitted in the same direction. It should also be understood that, according to the number and location of asynchronous FIFO modules in the hash engine, the corresponding clock module can be reasonably arranged to provide clock signals to each operation stage and each asynchronous FIFO module, as long as the clock signal direction is consistent and the asynchronous FIFO module is provided with different clock signals, Figure 3 Only an example arrangement is shown and is not intended to limit the present disclosure.

[0041] In some embodiments, the arithmetic module 120 can include a first hash engine including a first plurality of arithmetic stages of the plurality of arithmetic stages of the arithmetic module 120 and a second hash engine including a second plurality of arithmetic stages of the plurality of arithmetic stages of the arithmetic module 120 following the first plurality of arithmetic stages. The first hash engine and the second hash engine are configured to perform a hash algorithm on the data block sequentially. The first plurality of arithmetic stages of the first hash engine are arranged in a pipelined structure such that a data signal based on the data block is passed along the first plurality of arithmetic stages sequentially, and the second plurality of arithmetic stages of the second hash engine are arranged in a pipelined structure such that the data signal received from the first hash engine is passed along the second plurality of arithmetic stages sequentially. In some embodiments, the first arithmetic stage described above is the last arithmetic stage of the first plurality of arithmetic stages of the first hash engine, and the second arithmetic stage described above is the first arithmetic stage of the second plurality of arithmetic stages of the second hash engine.

[0042] As a non-limiting example, for a computing chip arithmetic core circuit that needs to perform the SHA-256 algorithm twice, a total of 128 arithmetic stages are needed. The arithmetic module of the arithmetic core circuit can include two hash engines, each of which includes 64 arithmetic stages and is configured to perform the SHA-256 algorithm. Each hash engine may, for example, have a configuration as shown in Figure 12 It can be understood that the present disclosure does not particularly limit the hash algorithm performed by the hash engine, and the hash engine of the arithmetic core circuit can actually be used to perform any hash algorithm (not limited to the SHA series algorithm) known at present or developed later, and accordingly can include a corresponding number of arithmetic stages.

[0043] Figures 4-6 is an embodiment of the arithmetic core circuit including a SHA-256 hash engine corresponding to Figures 1-3 respectively.

[0044] As shown in Figure 4As shown, the operation module 120 of the operation core circuit 100A' includes a first hash engine 121 including 64 operation stages 121-1,..., 121-i,..., 121-64 and a second hash engine 122 including 64 operation stages 122-1,..., 122-i,..., 122-64. The operation stages 121-1,..., 121-i,..., 121-64, 122-1,..., 122-i,..., 122-64 are arranged in a pipeline structure such that the data signals are sequentially passed along the operation stages 121-1,..., 121-i,..., 121-64, 122-1,..., 122-i,..., 122-64. The asynchronous FIFO module 130 is preferably arranged between the first hash engine 121 and the second hash engine 122, i.e. between the operation stage 121-64 and the operation stage 122-1. Such an advantage is that, compared with the case that the asynchronous FIFO module 130 is arranged inside the first hash engine 121 or the second hash engine 122, the asynchronous FIFO module 130 arranged between the first hash engine 121 and the second hash engine 122 only needs to store and pass the values in the 8 compressed registers A-H of the operation stage 121-64 to the operation stage 122-1, but does not need to store the values of the 16 expanded registers of the operation stage 121-64, in such a case the asynchronous FIFO module 130 can be designed to have a smaller size, thereby saving the occupation area of the asynchronous FIFO module 130 on the chip.

[0045] Similarly to Figure 1 , in Figure 4 the example, the data signals are passed from each operation stage of the first hash engine 121 to each operation stage of the second hash engine 122 in a direction from left to right, while the first clock signal provided by the first clock module 141 to the first hash engine 121 propagates inside the first hash engine 121 from right to left, and the second clock signal provided by the second clock module 142 to the second hash engine 122 propagates inside the second hash engine 122 from right to left. The passing direction of the data signals in the operation core circuit 100A' is opposite to the passing direction of the clock signals.

[0046] Similarly to Figure 2 , in Figure 5In the example shown in FIG. 1 , data signals are transmitted from each operation stage of first hash engine 121 to each operation stage of second hash engine 122 in a left-to-right direction. The first clock signal provided by first clock module 141 to first hash engine 121 propagates from left to right within first hash engine 121, and the second clock signal provided by second clock module 142 to second hash engine 122 propagates from left to right within second hash engine 122. The transmission direction of data signals and clock signals in operation core circuit 100B′ is the same.

[0047] Similar to Figure 3 ,exist Figure 6 In the example shown in FIG1 , in addition to the asynchronous FIFO module 130 disposed between the first hash engine 121 and the second hash engine 122, the operation core circuit 100C′ further includes a second asynchronous FIFO module 131 disposed within the second hash engine 122 and a third asynchronous FIFO module 132 disposed within the first hash engine 121. It can be seen that the insertion of the additional asynchronous FIFO modules does not affect the relative relationship between the transmission direction of the data signal and the transmission direction of the clock signal. Instead, it further reduces the number of operation stages that each clock signal needs to pass through, thereby further optimizing the shape of the clock signal at each operation stage.

[0048] Typically, the operation core circuit can be implemented on a semiconductor chip (e.g., a silicon chip). All the operation stages of the pipeline structure are typically arranged in the same row and adjacent to each other in the horizontal direction. The horizontal direction referred to here may refer to the extension direction of the pipeline structure, that is, the transmission direction of the data signal. In some embodiments, the multiple operation stages of the pipeline structure can be divided into multiple sub-combinations, and these sub-combinations can be arranged in different rows along the surface of the semiconductor chip, so that each sub-combination is adjacent to each other in the vertical direction perpendicular to the horizontal direction. For example, for the above-mentioned operation core circuit including the first hash engine and the second hash engine, the first hash engine and the second hash engine can be arranged to be adjacent to each other in the vertical direction along the surface of the semiconductor chip. Such an operation core circuit can be referred to as an operation core circuit with a vertical structure in this article. The operation core circuit with a vertical structure can have a more suitable (e.g., closer to a square) aspect ratio, thereby facilitating the flexible arrangement of such an operation core circuit on the computing chip. In such a case, it is more convenient to cut out more chips that are generally rectangular from a silicon wafer that is generally circular. However, for a computation core circuit having a vertical structure, the distance a data signal must travel between the last computation stage of a sub-combination (e.g., the first hash engine) and the first computation stage of a downstream sub-combination (e.g., the second hash engine) is greater than the distance a data signal must travel between two adjacent computation stages within a sub-combination (e.g., the first hash engine and the second hash engine). This results in a longer data signal propagation time between the two computation stages than between two adjacent computation stages within the sub-combination, potentially limiting the processing speed and throughput of the computation core circuit. However, the asynchronous FIFO module has relaxed timing, which can help shorten the propagation time of a data signal from the last computation stage of a sub-combination (e.g., the first hash engine) to the first computation stage of a downstream sub-combination (e.g., the second hash engine) in the vertical structure, thereby improving the performance of the computation core circuit having a vertical structure. This allows the vertical structure to provide benefits without degrading the processing speed and throughput of the computation core circuit.

[0049] As such, in some embodiments, Figure 1 Taking the operation core circuit shown in FIG. 1 as an example, the first operation stage 121 - a and the operation stages before it may be located in the first row on the semiconductor chip, and the second operation stage 121 - b and the operation stages after it may be located in the second row on the semiconductor chip, and the first row and the second row may be adjacent to each other in the vertical direction. In some embodiments, Figure 3The illustrated operation core circuit is an example, the fifth operation stage 121-e and the operation stages before it can be located in the first row on the semiconductor chip, the operation stages between the sixth operation stage 121-f and the first operation stage 121-a can be located in the second row on the semiconductor chip, the operation stages between the second operation stage 121-b and the third operation stage 121-c can be located in the third row on the semiconductor chip, and the fourth operation stage 121-d and the operation stages after it can be located in the fourth row on the semiconductor chip. The first to fourth rows can be adjacent to each other in the vertical direction in this order.

[0050] Figure 7A and Figure 7B is a schematic diagram of an operation core circuit with a vertical structure according to an embodiment of the present disclosure. As Figure 7A illustrated, the first hash engine 221 and the second hash engine 222 of the operation core circuit 200A are adjacent to each other in the vertical direction, and the data signal passes from the first hash engine 221 to the second hash engine 222 via the asynchronous FIFO module 230, wherein both the first clock signal and the second clock signal are opposite to the direction of the data signal. As Figure 7B illustrated, the first hash engine 221 and the second hash engine 222 of the operation core circuit 200B are adjacent to each other in the vertical direction, and the data signal passes from the first hash engine 221 to the second hash engine 222 via the asynchronous FIFO module 230, wherein both the first clock signal and the second clock signal are the same as the direction of the data signal. The relative position relationship of the first hash engine 221 and the second hash engine 222 in the vertical direction as depicted in the figure is merely exemplary and not limiting, and the relative position relationship of the first hash engine 221 and the second hash engine 222 in the vertical direction can be reversed according to actual needs.

[0051] In addition, the present disclosure also provides an example arrangement for distributing clock signals to each operation core circuit in a computing chip comprising a plurality of operation core circuits.

[0052] As Figure 11A and Figure 11B illustrated, the computing chip 1200, 1200' can comprise a top-level module 1210 and a plurality of operation core circuits 1220. The top-level module 1210 comprises a clock source 1211. The clock source 1211 is configured to provide clock signals for the operation core circuits 1220 of the computing chip 1200. In Figure 11A and Figure 11B , the operation core circuits 1220 are shown as conventional operation core circuits not containing an asynchronous FIFO module, but it can be understood that this is merely exemplary, Figure 11A and Figure 11B The arrangement illustrated can be applicable to the operation core circuits containing an asynchronous FIFO module according to the present disclosure, which will be described in detail later.

[0053] like Figure 11A and Figure 11B As shown, computing chips 1200 and 1200' include multiple computing core circuits 1220 arranged in multiple columns 1220-1, 1220-2, 1220-3, and 1220-4. Each computing core circuit 1220 may include: an input module (not shown) configured to receive a data block; an operation module (not shown) configured to perform a hash operation on the received data block, the operation module including multiple operation stages arranged in a pipeline structure so that data signals based on the data block are sequentially transmitted along the multiple operation stages, each of the multiple operation stages operates on the data signal received from the previous operation stage and provides the data signal operated by the operation stage to the next operation stage; and a clock module (not shown) configured to provide clock signals to the multiple operation stages. Each clock module may receive a clock signal from a clock source 1211 via a clock channel.

[0054] It should be understood that although in the illustrated example, the computing chips 1200, 1200' include four columns and four rows of operation core circuits, this is merely illustrative and not restrictive, and any suitable number of operation core circuits can be arranged into any suitable number of columns according to actual circumstances.

[0055] like Figure 11A As shown, in some embodiments, the clock modules of each column of computing core circuits receive clock signals via a common clock channel. For example, column 1220-1 receives clock signals via a common clock channel 1231, column 1220-2 receives clock signals via a common clock channel 1232, column 1220-3 receives clock signals via a common clock channel 1233, and column 1220-4 receives clock signals via a common clock channel 1234. Figure 11BAs shown, in some embodiments, the plurality of columns 1220-1, 1220-2, 1220-3, 1220-4 includes a first column of compute core circuits and a second column of compute core circuits (e.g., 1220-1 and 1220-2) arranged adjacent to each other and in the stated order, the clock module of the first column of compute core circuits 1220-1 and the clock module of the second column of compute core circuits 1220-2 receiving a clock signal via a common clock channel 1235. The compute chip 1200 can also include multiple pairs of such first and second columns of compute core circuits (e.g., 1220-3 and 1220-4), each pair can take the above arrangement to receive a clock signal via a common clock channel, for example, columns 1220-3 and 1220-4 receive a clock signal via a common clock channel 1236. Two adjacent columns of compute core circuits share a clock channel, which can further save the area on chip for clock channels, thus can further reduce the chip size, or save more area for setting more compute core circuits.

[0056] In Figure 11A and Figure 11B In the depicted embodiment, in some examples, the direction of transfer of the clock signal and the data signal in each compute core circuit can be the same, while in other examples, the direction of transfer of the clock signal and the data signal in each compute core circuit can be opposite.

[0057] The present disclosure also provides a compute chip including one or more compute core circuits as described in any of the above embodiments.

[0058] The compute chip 300 according to some embodiments of the present disclosure is described below in connection with Figure 8 The compute chip 300 can include a top-level module 310 and a plurality of compute core circuits 320 having an asynchronous FIFO module as described above. In Figure 8 In the depicted embodiment, the compute core circuits 320 are shown as not having the vertical structure described above. However, in other embodiments, the compute core circuits 320 can have the vertical structure described above, which will be described later.

[0059] As Figure 8As shown, the top-level module 310 includes a clock source 311. The clock source 311 is configured to provide clock signals to the compute core circuit 320 of the compute chip 300. The compute core circuit 320 is arranged in a plurality of columns 320-1, 320-2, 320-3, 320-4, with a first clock module of each column of compute core circuits receiving a first clock signal via a common clock channel and a second clock module of each column of compute core circuits receiving a second clock signal via a common clock channel. For example, one of the first and second clock modules of each compute core circuit in column 320-1 receives a clock signal via a common clock channel 331 and the other receives a clock signal via a common clock channel 332.

[0060] These compute core circuits 320 can also have other arrangements for distributing clock signals from the clock source 311. Similar to the arrangement of Figure 11B , adjacent columns of compute core circuits can share clock channels. Some example arrangements for distributing clock signals to compute core circuits are described below in connection with Figures 9A-9D . In Figures 9A-9D , F represents an asynchronous FIFO module, H1 represents a first hash engine (or a set of compute stages upstream of the asynchronous FIFO module), and H2 represents a second hash engine (or a set of compute stages downstream of the asynchronous FIFO module). Note that in Figures 9A-9D , the input of a clock signal into the asynchronous FIFO module can not be specifically embodied, but this is merely for clarity in illustrating the point being made. Figures 9A-9D

[0061] The plurality of columns of compute core circuits can include a first column of compute core circuits and a second column of compute core circuits (e.g., 320-1 and 320-2) that are adjacent to each other and arranged in the order stated. In some embodiments, one of the first and second clock modules of the first column of compute core circuits and one of the first and second clock modules of the second column of compute core circuits can receive a clock signal via a common clock channel.

[0062] As shown in Figure 9A , in some examples, the second clock module of the first column of compute core circuits 320-1 and the first clock module of the second column of compute core circuits 320-2 receive a clock signal via a common clock channel as the respective second clock signal and first clock signal, respectively. The first clock module of the first column of compute core circuits 320-1 receives a first clock signal via a common separate clock channel. The second clock module of the second column of compute core circuits 320-2 receives a second clock signal via a common separate clock channel. In such examples, the clock signal is in opposite direction to the data signal in the first column of compute core circuits 320-1 and the clock signal is in the same direction as the data signal in the second column of compute core circuits 320-2. ​

[0063] like Figure 9B As shown, in some examples, the second clock module of the first column computing core circuit 320-1 and the second clock module of the second column computing core circuit 320-2 receive clock signals as their respective second clock signals via a common clock channel. The first clock module of the first column computing core circuit 320-1 receives the first clock signal via a common, separate clock channel. The first clock module of the second column computing core circuit 320-2 receives the first clock signal via a common, separate clock channel. In such examples, the clock signals in the first column computing core circuit 320-1 and the second column computing core circuit 320-2 are inversely proportional to the data signal.

[0064] like Figure 9C As shown, in some examples, the first clock module of the first column computing core circuit 320-1 and the second clock module of the second column computing core circuit 320-2 receive clock signals via a common clock channel as their respective first clock signal and second clock signal. The first clock module of the second column computing core circuit 320-2 receives the first clock signal via a common, separate clock channel. The second clock module of the first column computing core circuit 320-1 receives the second clock signal via a common, separate clock channel. In such an example, the clock signal in the second column computing core circuit 320-2 is in the opposite direction of the data signal, while the clock signal in the first column computing core circuit 320-1 is in the same direction as the data signal.

[0065] like Figure 9D As shown, in some examples, the first clock module of the first column computing core circuit 320-1 and the first clock module of the second column computing core circuit 320-2 receive clock signals as their respective first clock signals via a common clock channel. The second clock module of the first column computing core circuit 320-1 receives a second clock signal via a common separate clock channel. The second clock module of the second column computing core circuit 320-2 receives a second clock signal via a common separate clock channel. In such examples, the clock signals in the first column computing core circuit 320-1 and the second column computing core circuit 320-2 are in the same direction as the data signal.

[0066] The computing chip may further include multiple pairs of such first column operation core circuits and second column operation core circuits (eg, 320 - 3 and 320 - 4 ), and each pair may adopt any of the above arrangements to receive the clock signal.

[0067] exist Figures 9A-9DIn the example shown, by sharing clock channels between adjacent columns of computational core circuits, the number of clock channels required in the computing chip is reduced, thereby further reducing chip size or freeing up more area for arranging more computational core circuits. This beneficial effect is more pronounced when the computing chip contains more columns of computational core circuits.

[0068] Figure 8 The operation core circuit in can also be the operation core circuit 320' having the vertical structure described above. Figures 10A-10D It is shown in Figure 8 When the computing chip 300 includes a plurality of computing core circuits 320', an arrangement for distributing clock signals to the computing core circuits 320' is provided. Figures 10A-10D , F represents an asynchronous FIFO module, H1 represents a first hash engine (or a set of operation stages upstream of the asynchronous FIFO module), and H2 represents a second hash engine (or a set of operation stages downstream of the asynchronous FIFO module). In addition, the relative positional relationship between H1 and H2 in the vertical direction depicted in the accompanying drawings is merely exemplary and non-restrictive, and the relative positional relationship between H1 and H2 in the vertical direction may be reversed according to actual needs. In addition, the description here also applies to an operation core circuit including three or more vertically adjacent rows of sub-combinations, and will not be repeated here.

[0069] Multiple computing core circuits 320′ are arranged in multiple columns 320-1′, 320-2′, 320-3′, and 320-4′. In some embodiments, similar to the case of computing core circuit 320, the first clock module of each column of computing core circuit 320′ can receive a first clock signal via a common clock channel, and the second clock module of each column of computing core circuit 320′ can receive a second clock signal via the common clock channel.

[0070] In other cases, clock channels can also be shared between adjacent columns of compute core circuits 320'. The plurality of columns 320-1', 320-2', 320-3', 320-4' of compute core circuits 320' includes a first column of compute core circuits and a second column of compute core circuits (e.g., 320-1', 320-2') that are adjacent to each other and arranged in the stated order. In some embodiments, one of the first clock module and the second clock module of the first column of compute core circuits and one of the first clock module and the second clock module of the second column of compute core circuits can receive a clock signal via a common clock channel. Additionally, in some embodiments, the plurality of columns includes the first column of compute core circuits, the second column of compute core circuits, and a third column of compute core circuits that are adjacent to each other and arranged in the stated order, another of the first clock module and the second clock module of the second column of compute core circuits and one of the first clock module and the second clock module of the third column of compute core circuits receive a clock signal via a common clock channel.

[0071] In some embodiments, as shown in FIG. 3A, the second clock module of the first column of compute core circuits 320-1' and the second clock module of the second column of compute core circuits 320-2' receive a clock signal via a common clock channel as respective second clock signals. Alternatively, in other embodiments, the first clock module of the first column of compute core circuits 320-1' and the first clock module of the second column of compute core circuits 320-2' can receive a clock signal via a common clock channel as respective first clock signals. Figure 10A In some embodiments, as shown in FIG. 3A, the second clock module of the first column of compute core circuits 320-1' and the second clock module of the second column of compute core circuits 320-2' receive a clock signal via a common clock channel as respective second clock signals. Alternatively, in other embodiments, the first clock module of the first column of compute core circuits 320-1' and the first clock module of the second column of compute core circuits 320-2' can receive a clock signal via a common clock channel as respective first clock signals. Figure 10D In some embodiments, as shown in FIG. 3A, the second clock module of the first column of compute core circuits 320-1' and the second clock module of the second column of compute core circuits 320-2' receive a clock signal via a common clock channel as respective second clock signals. Alternatively, in other embodiments, the first clock module of the first column of compute core circuits 320-1' and the first clock module of the second column of compute core circuits 320-2' can receive a clock signal via a common clock channel as respective first clock signals.

[0072] In some embodiments, as shown in FIG. 3A, the second clock module of the first column of compute core circuits 320-1' and the second clock module of the second column of compute core circuits 320-2' receive a clock signal via a common clock channel as respective second clock signals. Alternatively, in other embodiments, the first clock module of the first column of compute core circuits 320-1' and the first clock module of the second column of compute core circuits 320-2' can receive a clock signal via a common clock channel as respective first clock signals. Figure 10A and Figure 10DIn the example, the other clock modules of the first column of operation core circuits 320-1′ and the second column of operation core circuits 320-2′ that do not share a common clock channel each receive a clock signal via a separate clock channel. Thus, the clock channel can be shared between the operation core circuits in each pair of two adjacent columns. In other embodiments, the other clock module of the second column of operation core circuits 320-2′ can also share a clock channel with the clock module of another column of operation core circuits that is adjacent to the second column of operation core circuits 320-2′ and opposite to the first column of operation core circuits 320-1′. Thus, the clock channel can be shared between two of the operation core circuits in each group of three adjacent columns. By analogy, each column of operation core circuits can share a clock channel with the operation core circuits in the adjacent columns on both sides of it.

[0073] For example, Figure 10B As shown, in addition to the first column of computing core circuits 320-1′ and the second column of computing core circuits 320-2′, computing chip 1200 also includes a third column of computing core circuits 320-3′ adjacent to the second column of computing core circuits 320-2′. The first clock module of the second column of computing core circuits 320-2′ and the first clock module of the third column of computing core circuits 320-3′ receive clock signals via a common clock channel as their respective first clock signals. In an alternative embodiment in which the first clock modules of the first column of computing core circuits 320-1′ and the second column of computing core circuits 320-2′ share a clock channel, the second clock module of the second column of computing core circuits 320-2′ and the second clock module of the third column of computing core circuits 320-3′ receive clock signals via the common clock channel as their respective second clock signals.

[0074] Further Figure 10B As shown, computing chip 1200 further includes a fourth column of computing core circuits 320-4' adjacent to the third column of computing core circuits 320-3'. The second clock module of the third column of computing core circuits 320-3' and the second clock module of the fourth column of computing core circuits 320-4' receive clock signals via a common clock channel as their respective second clock signals.

[0075] In addition, although Figure 10A and Figure 10B In the embodiment, the asynchronous FIFO module is arranged on the left side of the hash engines H1 and H2 and H1 is set on H2, but this is only exemplary and not restrictive. Figure 10C As shown, the asynchronous FIFO modules of the operation core circuits of columns 320-2' and 320-4' are arranged on the right side of the hash engines H1 and H2, as shown in FIG. Figure 10DAs shown, the H2 of column 320-2' is arranged above the H1, which still enables the clock signal distribution arrangement provided by the present disclosure. In fact, the position of the asynchronous FIFO module in each operation core circuit relative to the hash engine and the relative position between the first hash engine and the second hash engine can be reasonably arranged according to actual conditions, and it is not necessarily required that the arrangement of each operation core circuit be the same.

[0076] Correspondingly, it is also not necessarily required that the relative relationship of the transmission direction of the data signal and the clock signal in each operation core circuit be the same. The transmission direction of the data signal and the clock signal in some of the operation core circuits of the computing chip can be the same, while the transmission direction of the data signal and the clock signal in some other operation core circuits can also be opposite, which is not particularly limited, but can be reasonably arranged according to actual conditions.

[0077] Since the aspect ratio of the conventional operation core circuit of the computing chip is usually large (because up to 128 operation stages are arranged), the arrangement of the operation core circuit on the computing chip (which is generally based on a silicon wafer) is very limited. The operation core circuit with a vertical structure provided by the present disclosure can have a significantly reduced aspect ratio and can be more flexibly and freely arranged on the computing chip. The inclusion of the asynchronous FIFO module can also help to improve the performance of the operation core circuit with a vertical structure. In addition, through the sharing of the clock channel between adjacent columns of the operation core circuit, the chip area can be further saved, and a larger number of operation core circuits can also be arranged on a chip of the same size to efficiently undertake complex operation tasks.

[0078] It can be understood that, although the sharing of the clock channel between adjacent columns of the operation core circuit is described in the above embodiments, the sharing of the clock channel between adjacent rows of the operation core circuit in a similar manner is also feasible and is also covered within the scope of the present disclosure.

[0079] The words "left," "right," "front," "back," "top," "bottom," "over," "under," "upper," "lower," and the like in the description and the claims, if any, are used for descriptive purposes and not necessarily for describing permanent relative positions. It is to be understood that the words so used are interchangeable under appropriate circumstances such that the embodiments of the present disclosure described herein are capable of operation in other orientations than described or otherwise shown in the illustrations. For example, if the device is turned over, the feature described as above other features can now be described as below. The device can be oriented in other ways (rotated at 90 degrees or at other orientations) and the relative spatial relationships would be correspondingly interpreted.

[0080] In the description and claims, when an element is referred to as being "on", "attached", "connected", "coupled", or "in contact" with another element, it can be directly on, attached, connected, coupled, or in contact with the other element, or one or more intervening elements can also be present. In contrast, when an element is referred to as being "directly on", "directly attached", "directly connected", "directly coupled", or "directly in contact" with another element, there are no intervening elements present. In the description and claims, when a feature is arranged "adjacent" to another feature, it can mean that the feature has a portion that overlaps the adjacent feature or a portion that is above or below the adjacent feature.

[0081] As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any implementation of the described implementations described herein as exemplary is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, the disclosed implementations are not intended to be limited to any expressed or implied theory of operation.

[0082] As used herein, the word "substantially" means including any minor variations as a result of design, manufacturing, and / or other factors that can be expected to occur in real world implementations. The word "substantially" also allows for differences that are within normal manufacturing tolerances and / or other factors that can be expected to occur in real world implementations.

[0083] In addition, the terms "first", "second", and like terms are also used herein, merely for purposes of reference, and thus do not imply or create (1) an order or sequence etc., unless specifically stated, (2) limitation on the number of components etc.

[0084] It is also to be understood that the phraseology "comprising", "including", "containing", "consisting", "consisting essentially of", and the like, when used in the present specification, specifies the presence of stated features, integers, steps, operations, elements, and / or components but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0085] In the present disclosure, the term "providing" is used in a broad sense and is intended to encompass all means of obtaining an object, and thus "providing an object" includes, but is not limited to, "purchasing", "preparing / manufacturing", "arranging / setting", "installing / fitting", and / or "ordering" the object, etc.

[0086] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0087] Those skilled in the art will realize that the boundaries between the above described operations merely illustrative. The multiple operations can be combined into a single operation, a single operation can be distributed in additional operations and operations can be executed at least partially overlapping in time. Moreover, alternative embodiments can include a number of instances of a particular operation, and the order of the operations can be altered in other various embodiments. However, other modifications, variations, and alternatives are also possible. The aspects and elements of all such embodiments can be combined in any manner and / or combination with other aspects or elements of other embodiments, as would be understood by one of ordinary skill in the art, to provide a number of additio nal embodiments. The present description and drawings should be considered in a descriptive sense only and not for purposes of limitation. Thus, whereas the present disclosure has been described in detail with respect to specific embodiments, it will be apparent that various alterations, modifications, and improvements will occur to others upon reading the above description. Therefore, the format, system, and / or method disclosed herein should not be construed as limiting, but is intended to be broadly defined by the appended claims.

[0088] While particular embodiments of the present disclosure have been described in detail, it is to be understood that the foregoing examples and explanations are intended to be illustrative only and are not intended to limit the scope of the disclosure. Embodiments disclosed herein can be combined in any manner and / or combination with other aspects or elements of other embodiments, as would be understood by one of ordinary skill in the art, without departing from the spirit and scope of the present disclosure. Those skilled in the art will appreciate that modifications, variations, and alternatives to the embodiments described herein can be practiced without departing from the spirit and scope of the present disclosure. It is intended that the following claims define the scope of the disclosure and that modifications and variations such as those discussed above are within the scope of the disclosure.

Claims

1. An operation core circuit comprising: an input module configured to receive a data block; an operation module configured to perform a hash operation on the received data block, the operation module including a plurality of operation stages arranged in a pipeline structure such that a data signal based on the data block is sequentially passed along the plurality of operation stages, each of the plurality of operation stages performing an operation on a data signal received from a previous operation stage and providing a data signal operated by the operation stage to a subsequent operation stage; an asynchronous first-in-first-out (FIFO) module disposed between first and second adjacent operation stages of the plurality of operation stages, the first operation stage preceding the second operation stage, the asynchronous FIFO module configured to receive a data signal output from the first operation stage with a first clock signal to write to a FIFO buffer and to read the received data signal from the FIFO buffer with a second clock signal different from the first clock signal to output to the second operation stage; a first clock module configured to provide the first clock signal to the asynchronous FIFO module and to the first operation stage and to operation stages of the plurality of operation stages preceding the first operation stage; and a second clock module configured to provide the second clock signal to the asynchronous FIFO module and to the second operation stage and to operation stages of the plurality of operation stages succeeding the second operation stage, wherein a frequency of the first clock signal is the same as a frequency of the second clock signal, and wherein a passing time of a data signal between the asynchronous FIFO module and an operation stage is no more than a passing time between operation stages.

2. The operation core circuit of claim 1, wherein: passing directions of both the first and second clock signals are the same as a passing direction of the data signal, or passing directions of both the first and second clock signals are opposite to the passing direction of the data signal. the asynchronous FIFO module is a first asynchronous FIFO module, the operation core circuit further comprising:

3. The arithmetic core circuit according to claim 1, wherein a second asynchronous FIFO module disposed between third and fourth adjacent operation stages of the plurality of operation stages, the third operation stage preceding the fourth operation stage and succeeding the second operation stage, the second asynchronous FIFO module configured to receive a data signal output from the third operation stage with the second clock signal and to output the received data signal to the fourth operation stage with a third clock signal different from the second clock signal; and a third clock module configured to provide the third clock signal to the second asynchronous FIFO module and to the fourth operation stage and to operation stages of the plurality of operation stages succeeding the fourth operation stage, wherein the second clock module is configured to provide the second clock signal to the first and second asynchronous FIFO modules and to the second and third operation stages and operation stages therebetween.

3. The operation core circuit of claim 2, wherein: the first clock module is configured to provide the first clock signal to the first and second asynchronous FIFO modules and to the first and second operation stages and operation stages therebetween, and the second clock module is configured to provide the second clock signal to the first and second asynchronous FIFO modules and to the second and third operation stages and operation stages therebetween.

4. The arithmetic core circuit according to claim 1, wherein The asynchronous FIFO module is a first asynchronous FIFO module, and the operation core circuit further comprises: a third asynchronous FIFO module disposed between a fifth operation stage and a sixth operation stage adjacent in the plurality of operation stages, the sixth operation stage being subsequent to the fifth operation stage and prior to the first operation stage, the third asynchronous FIFO module being configured to receive a data signal output from the fifth operation stage using a fourth clock signal different from the first clock signal and output the received data signal to the sixth operation stage using the first clock signal; and a fourth clock module configured to provide the fourth clock signal to the third asynchronous FIFO module and to the fifth operation stage and operation stages of the plurality of operation stages prior to the fifth operation stage, wherein the first clock module is configured to provide the first clock signal to the first asynchronous FIFO module and the third asynchronous FIFO module and to the sixth operation stage and the first operation stage and operation stages therebetween.

5. The arithmetic core circuit according to any one of claims 1 to 4, wherein, The operation module comprises a first hash engine and a second hash engine, the first hash engine comprising a first plurality of operation stages of the plurality of operation stages, the second hash engine comprising a second plurality of operation stages of the plurality of operation stages subsequent to the first plurality of operation stages, the first hash engine and the second hash engine being configured to sequentially perform a hash algorithm on the data block.

6. The arithmetic core circuit according to claim 5, wherein The first operation stage is a last operation stage of the first plurality of operation stages of the first hash engine, and the second operation stage is a first operation stage of the second plurality of operation stages of the second hash engine.

7. The arithmetic core circuit according to claim 6, wherein The operation core circuit is implemented on a semiconductor chip, and the first hash engine and the second hash engine are arranged adjacent to each other in a vertical direction perpendicular to a direction of transfer of the data signal along a surface of the semiconductor chip.

8. A computing chip comprising one or more operation core circuits according to any one of claims 1-7.

9. A computing chip comprising a plurality of operation core circuits according to any one of claims 1-7, the plurality of operation core circuits being arranged in a plurality of columns, the first clock module of each column of operation core circuits receiving a first clock signal via a common clock channel, and the second clock module of each column of operation core circuits receiving a second clock signal via a common clock channel.

10. The computing chip of claim 9, wherein: the plurality of columns comprises a first column of operation core circuits and a second column of operation core circuits adjacent to each other and arranged in the recited order, one of the first clock module and the second clock module of the first column of operation core circuits and one of the first clock module and the second clock module of the second column of operation core circuits receiving a clock signal via a common clock channel.

11. The compute chip of claim 10, wherein, the operation core circuit is according to claim 7, and wherein: The plurality of columns includes the first column of operation core circuits, the second column of operation core circuits, and a third column of operation core circuits arranged adjacent to each other and in the recited order, another of the first clock module and the second clock module of the second column of operation core circuits and one of the first clock module and the second clock module of the third column of operation core circuits receive a clock signal via a common clock channel.

Citation Information

Patent Citations

  • Block chain mining apparatus

    CN111427891A

  • Clock tree, hash engine, computing chip, computing power board and digital currency mining machine

    CN111930682A

  • Operation core, calculation chip and cryptocurrency mining machine

    CN213399572U