Instruction forwarding circuit and method for data synchronization between computing units
By designing an instruction forwarding circuit for data synchronization between calculation units, the problem of data synchronization time in the prior art is solved, and more efficient data synchronization and improvement of computing unit performance is achieved.
Patent Information
- Application Number
- CN202410587748.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-05-11
AI Technical Summary
In the prior art, data synchronization between calculation units takes a long time, resulting in low utilization of calculation units and degradation of performance.
An instruction forwarding circuit is designed, including an instruction storage module and an instruction scheduling module. By scheduling fence instructions and data synchronization instructions, the time overhead of data synchronization is reduced and the performance of the computing unit is improved.
Through the instruction forwarding circuit, the time overhead of data synchronization between calculation units is reduced, the efficiency of data synchronization is improved, and the performance of computing unit on the data production end is released, so that the computing unit on the data consumer end can load data faster and improve its performance.
Smart Images

Figure CN118535227B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of data processing technology, and more specifically to an instruction forwarding circuit and method for data synchronization between computing units. Background Art
[0002] Generally, data synchronization can be achieved between computing units through store-fence-sync / atomic-load. For example, when data needs to be synchronized from one computing unit A to another computing unit B, computing unit A first writes and stores the data to be synchronized in the target address (such as the first-level cache of computing unit A) through the store instruction; then computing unit A sends a fence instruction and waits for the response corresponding to the fence instruction to ensure that all the data to be synchronized are written to the target address; after receiving the response corresponding to the fence instruction, computing unit A sends a sync instruction or an atomic instruction to notify computing unit B that it can load data from the target address; finally, computing unit B loads data through the load instruction, thereby achieving data synchronization between the two computing units (i.e., computing unit A and computing unit B).
[0003] However, the above solution for data synchronization between computing units is time-consuming, and during the data synchronization process, the computing units are in a state of waiting for instructions for a long time, resulting in low utilization of the computing units and reduced performance. Summary of the invention
[0004] In view of the above problems, the present invention provides an instruction forwarding circuit and method for data synchronization between computing units, which can reduce the time overhead of data synchronization between computing units, improve the efficiency of data synchronization, and improve the performance of computing units.
[0005] According to a first aspect of the present invention, there is provided an instruction forwarding circuit for data synchronization between computing units, comprising: an instruction storage module, configured to store instructions received from a first computing unit, and to send the instructions received from the first computing unit to an instruction scheduling module in sequence; and an instruction scheduling module, configured to schedule instructions to be sent based at least on a current instruction obtained from the instruction storage module, wherein the instructions received from the first computing unit include: fence instructions and data synchronization instructions.
[0006] In some embodiments, the data synchronization instruction is one of a synchronization instruction and an atomic operation instruction.
[0007] In some embodiments, scheduling the instruction to be sent includes: checking the dependency between the current instruction obtained from the instruction storage module and other instructions; and sending the current instruction in response to the current instruction having no dependency with other instructions.
[0008] In some embodiments, scheduling instructions to be sent also includes: in response to the current instruction having a dependency on other instructions, determining whether all instructions having a dependency relationship with the current instruction have been received; and in response to receiving all instructions having a dependency relationship with the current instruction, sending the current instruction.
[0009] In some embodiments, scheduling the instructions to be sent further includes: in response to not receiving all instructions having a dependency relationship with the current instruction, waiting until all instructions having a dependency relationship with the current instruction are received; and sending the current instruction.
[0010] In some embodiments, the instruction scheduling module is further configured to: based on the received response corresponding to the sent instruction, send a deletion request for deleting the corresponding instruction to the instruction storage module.
[0011] In some embodiments, the instruction storage module is further configured to: based on a deletion request received from the instruction scheduling module, delete the instruction corresponding to the deletion request.
[0012] In some embodiments, the current instruction obtained from the instruction storage module includes: a data item indicating the dependency of the current instruction, wherein the data item includes at least: an identifier of the current instruction, and information for indicating the dependency of the current instruction.
[0013] According to a second aspect of the present invention, a method for data synchronization between computing units is provided, comprising: a first computing unit sends a storage instruction to store the data to be synchronized to a target address; the first computing unit sends a fence instruction and a data synchronization instruction to an instruction forwarding circuit according to the first aspect of the present invention; the instruction forwarding circuit sends a fence instruction; in response to the instruction forwarding circuit receiving a response corresponding to the fence instruction, the instruction forwarding circuit sends a data synchronization instruction; and a second computing unit loads the data to be synchronized from the target address.
[0014] According to a third aspect of the present invention, a computing module is provided, comprising: at least two computing units; and a first instruction forwarding circuit, which is an instruction forwarding circuit according to the first aspect of the present invention and is configured for data synchronization between at least two computing units.
[0015] According to a fourth aspect of the present invention, a computing circuit is provided, comprising: at least two computing modules according to the third aspect of the present invention; and a second instruction forwarding circuit, which is an instruction forwarding circuit according to the first aspect of the present invention and is configured for data synchronization between at least two computing modules.
[0016] In some embodiments, the computing circuit is a graphics processing unit and the computing module is a cluster of stream processors.
[0017] According to a fifth aspect of the present invention, a computing system is provided, comprising: at least two computing circuits according to the fourth aspect of the present invention; and a third instruction forwarding circuit, which is an instruction forwarding circuit according to the first aspect of the present invention and is configured for data synchronization between at least two computing circuits.
[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements.
[0020] Figure 1 An exemplary circuit diagram of an instruction forwarding circuit according to an embodiment of the present invention is shown.
[0021] Figure 2 A flow chart of a method in which an instruction scheduling module schedules instructions to be sent according to an embodiment of the present invention is shown.
[0022] Figure 3 A flow chart of a method for data synchronization between computing units according to an embodiment of the present invention is shown.
[0023] Figure 4A A timing diagram showing data synchronization between two computing units.
[0024] Figure 4B A timing diagram of data synchronization between two computing units according to an embodiment of the present invention is shown.
[0025] Figure 5 An exemplary structural diagram of a computing system according to an embodiment of the present invention is shown.
[0026] Figure 6 A timing diagram of data synchronization between two computing units located in different computing modules on the same computing circuit according to an embodiment of the present invention is shown.
[0027] Figure 7 A timing diagram of data synchronization between two computing units in computing modules located on different computing circuits according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0028] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0029] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0030] As mentioned above, data synchronization between computing units is mainly carried out in the manner of store-fence-sync / atomic-load. Specifically, for two computing units that are closer, data synchronization between them can be carried out in the manner of store-fence-sync-load; and for two computing units that are farther away, data synchronization between them can be carried out in the manner of store-fence-atomic-load. However, in the above two data synchronization methods, whether it is a sync instruction or an atomic instruction, it is necessary to wait until the response corresponding to the fence instruction is received before it can be issued. For example, in the example of synchronizing the data of computing unit A (also known as the computing unit at the data production end) to computing unit B (also known as the computing unit at the data consumption end), if the time to send the fence instruction and the time to wait for the response corresponding to the fence instruction are long, computing unit A needs to wait until the response corresponding to the fence instruction is received before sending the sync instruction or the atomic instruction. In this case, computing unit A will be in a state of waiting for instructions for a long time, and will not be able to perform other calculations, resulting in low utilization of computing unit A and reduced performance. Accordingly, the time that computing unit B spends waiting for a sync instruction or an atomic instruction to load data is also long, which hinders the performance of computing unit B.
[0031] In summary, the existing solutions for data synchronization between computing units have the following disadvantages: they are time-consuming, and the long waiting time during the data synchronization process hinders the performance of the computing units.
[0032] In order to at least partially solve one or more of the above problems and other potential problems, an example embodiment of the present invention proposes a scheme for data synchronization between computing units. In the above scheme for data synchronization between computing units, an instruction forwarding circuit for data synchronization between computing units is used, and the instruction forwarding circuit includes: an instruction storage module, which is configured to store instructions received from a first computing unit and send the instructions received from the first computing unit to an instruction scheduling module in sequence; and an instruction scheduling module, which is configured to schedule the instructions to be sent based on at least the current instructions obtained from the instruction storage module, so that the instruction forwarding circuit can replace the computing unit at the data production end to send instructions including fence instructions and data synchronization instructions, thereby releasing the performance of the computing unit at the data production end in advance, and at the same time can reduce the time overhead of data synchronization between computing units, so that the computing unit at the data consumption end can load data faster, thereby improving the performance of the computing unit at the data consumption end.
[0033] The following will be combined Figures 1 to 3 A design scheme of an instruction forwarding circuit for data synchronization between computing units according to an embodiment of the present invention is described in detail.
[0034] Figure 1 An exemplary circuit diagram of an instruction forwarding circuit 100 according to an embodiment of the present invention is shown, wherein the instruction forwarding circuit 100 is configured for data synchronization between a first computing unit 101 and a second computing unit 102 .
[0035] like Figure 1 As shown, the instruction forwarding circuit 100 includes: an instruction storage module 110 and an instruction scheduling module 120. It should be understood that the instruction forwarding circuit 100 may also include additional circuit modules not shown, and the scope of the present invention is not limited in this respect.
[0036] Regarding the instruction storage module 110, it may be configured to store instructions received from the computing unit of the data production end, and send these instructions in sequence to the instruction scheduling module 120. According to an embodiment of the present invention, the instruction storage module 110 may be configured to store instructions received from the first computing unit 101, and send the instructions received from the first computing unit 101 in sequence to the instruction scheduling module 120. In some embodiments, the instruction storage module 110 may store the instructions received from the first computing unit 101 in the instruction list of the instruction storage module 110 in sequence.
[0037] Regarding the instructions received from the first computing unit 101, according to an embodiment of the present invention, the instructions may include barrier instructions and data synchronization instructions.
[0038] The fence instruction can be used to ensure that all the data to be synchronized are written to the target address. Specifically, the fence instruction can be sent to the cache unit used to store the data to be synchronized. When all the data to be synchronized are written to the cache unit, the cache unit will return a response corresponding to the fence instruction to indicate that all the data to be synchronized have been written to the cache unit, so that the operation corresponding to the next instruction can be executed. In other words, the fence instruction can ensure that the operation indicated by the instruction sent before it is fully executed before sending the subsequent instruction.
[0039] Regarding the data synchronization instruction, it can be one of a synchronization instruction (sync instruction) and an atomic operation instruction (atomic instruction), which depends on the distance between the computing units to be synchronized. According to an embodiment of the present invention, when the distance between the computing units to be synchronized is close, the data synchronization instruction can be a sync instruction; when the distance between the computing units to be synchronized is far, the data synchronization instruction can be an atomic instruction.
[0040] Regarding the instruction scheduling module 120, it can be configured to schedule and send the instructions obtained from the instruction storage module 110. According to an embodiment of the present invention, the instruction scheduling module 120 can schedule the instructions to be sent based on at least the current instructions obtained from the instruction storage module 110, so that the second computing module 102 can read the data stored by the first computing unit 101 at the target address based on the instructions sent by the instruction scheduling module 120.
[0041] Regarding scheduling instructions to be sent, according to an embodiment of the present invention, it may include: the instruction scheduling module 120 checks the dependency of the current instruction obtained from the instruction storage module 110 on other instructions; in response to the current instruction not having a dependency on other instructions, sends the current instruction; in response to the current instruction having a dependency on other instructions, determines whether all instructions with a dependency relationship with the current instruction are received; in response to receiving all instructions with a dependency relationship with the current instruction, sends the current instruction; and in response to not receiving all instructions with a dependency relationship with the current instruction, waits until all instructions with a dependency relationship with the current instruction are received; sends the current instruction. The following will be combined with Figure 2 How the instruction scheduling module 120 schedules instructions to be sent is described in detail.
[0042] Figure 2 Shows Figure 1Flowchart of method 200 for scheduling instructions to be sent by instruction scheduling module 120. It should be understood that method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this respect.
[0043] In step 202 , the instruction scheduling module 120 obtains instructions from the instruction storage module 110 .
[0044] According to an embodiment of the present invention, the instruction scheduling module 120 may obtain instructions from the instruction list of the instruction storage module 110. For example, the instruction scheduling module 120 may obtain the instruction at the head of the instruction list from the instruction list of the instruction storage module 110.
[0045] In step 204 , the instruction scheduling module 120 checks whether the obtained current instruction has dependencies with other instructions.
[0046] Regarding the current instruction obtained, it may include information related to the dependency of the current instruction. According to an embodiment of the present invention, the current instruction obtained may include a data item indicating the dependency of the current instruction, and the data item may include: an identifier of the current instruction, information indicating the dependency of the current instruction, etc. For example, in some examples, 0 or -1 may be used as information indicating the dependency of the current instruction, where 0 indicates that the current instruction has no dependency on other instructions, and -1 indicates that the current instruction has a dependency on other instructions.
[0047] Regarding other instructions, they may refer to instructions other than the current instructions obtained and received by the instruction scheduling module 120. For example, in some examples, other instructions may be one or more instructions sent to the instruction scheduling module 120 by other modules other than the instruction storage module 110.
[0048] Dependency refers to the necessary relationship between the reception and transmission of different instructions. For example, in the data synchronization process between computing units as described above, since the sync instruction (or atomic instruction) must be sent only after receiving the response corresponding to the fence instruction, the sync instruction (or atomic instruction) and the response corresponding to the fence instruction can be regarded as having a dependency relationship, that is, the sync instruction (or atomic instruction) and the response corresponding to the fence instruction are dependent.
[0049] According to an embodiment of the present invention, the instruction scheduling module 120 can determine whether the current instruction has dependencies with other instructions by checking the data items indicating the dependencies of the current instruction included in the acquired current instruction. In the above example, if the instruction scheduling module 120 checks that the value of the information indicating the dependencies of the current instruction included in the data items indicating the dependencies of the current instruction is 0, the instruction scheduling module 120 determines that the current instruction does not have dependencies with other instructions, and then proceeds to step 206, and the instruction scheduling module 120 sends the current instruction to its downstream module. If the instruction scheduling module 120 checks that the value of the information indicating the dependencies of the current instruction included in the data items indicating the dependencies of the current instruction is -1, and the instruction scheduling module 120 determines that the current instruction has dependencies with other instructions, then proceeds to step 208, and the instruction scheduling module 120 further determines whether all instructions that have dependencies with the current instruction are received.
[0050] For example, if the current instruction checked by the instruction scheduling module 120 is a fence instruction, since the value of the information indicating the dependency of the instruction in the data item indicating the dependency of the instruction included in the fence instruction is 0, indicating that the fence instruction has no dependency with other instructions, the instruction scheduling module 120 can send the fence instruction to the downstream module. If the current instruction checked by the instruction scheduling module 120 is a sync instruction, since the value of the information indicating the dependency of the instruction in the data item indicating the dependency of the instruction included in the sync instruction is -1, indicating that the sync instruction has a dependency with other instructions, the instruction scheduling module 120 needs to further determine whether all instructions having a dependency relationship with the sync instruction are received. In this example, all instructions having a dependency relationship with the sync instruction refer to the response corresponding to the fence instruction. If the instruction scheduling module 120 determines that the response corresponding to the fence instruction has been received, the sync instruction can be sent to the downstream module; if the instruction scheduling module 120 determines that the response corresponding to the fence instruction has not been received, the instruction scheduling module 120 waits until it is determined that the response corresponding to the fence instruction has been received, and then sends the sync instruction to the downstream module.
[0051] Furthermore, when the instruction scheduling module 120 determines that it has received a response corresponding to the fence instruction, the instruction scheduling module 120 may also set the value of the information indicating the dependency of the instruction in the data item indicating the instruction dependency included in the sync instruction to 0, to indicate that the sync instruction no longer has any dependency, so that the sync instruction can be sent to the downstream module.
[0052] In summary, when it is determined in step 208 that all instructions that have a dependency relationship with the current instruction have been received, proceed to step 206, and the instruction scheduling module 120 sends the current instruction to the downstream module; when it is determined in step 208 that all instructions that have a dependency relationship with the current instruction have not been received, wait at step 208 until it is determined that all instructions that have a dependency relationship with the current instruction have been received, and then proceed to step 206, and the instruction scheduling module 120 sends the current instruction to the downstream module.
[0053] Back to Figure 1 According to some embodiments of the present invention, the instruction scheduling module 120 of the instruction forwarding circuit 100 may also be configured to: based on a received response corresponding to the sent instruction, send a deletion request for deleting the corresponding instruction to the instruction storage module 110; and the instruction storage module 110 may also be configured to: based on the deletion request received from the instruction scheduling module 120, delete the instruction corresponding to the deletion request.
[0054] For example, the instruction scheduling module 120 may obtain the first instruction from the instruction storage module 110 and send the first instruction to its downstream module. In response to the instruction scheduling module 120 receiving a response corresponding to the sent first instruction, the instruction scheduling module 120 may send a deletion request for deleting the first instruction to the instruction storage module 110. Accordingly, the instruction storage module 110 receives the deletion request for deleting the first instruction from the instruction scheduling module 120, and then may delete the first instruction stored in its instruction list based on the deletion request.
[0055] From the above, it can be seen that the instruction forwarding circuit for data synchronization between computing units provided according to the present invention can, through the instruction storage module and instruction scheduling module configured therein, realize that the instruction forwarding circuit replaces the computing unit at the data production end to send instructions related to data synchronization, so as to release the performance of the computing unit at the data production end.
[0056] The following will be combined Figure 3 ,by Figure 1 Taking the instruction forwarding circuit 100 shown as an example, a detailed description is given of how to achieve data synchronization between computing units through the instruction forwarding circuit provided by the present invention.
[0057] Figure 3 A flow chart of a method 300 for data synchronization between computing units according to an embodiment of the present invention is shown. It should be understood that the method 300 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this respect.
[0058] In step 302, the first computing unit 101 sends a storage instruction to store the data to be synchronized to the target address.
[0059] In step 304 , the first computing unit 101 sends a fence instruction and a data synchronization instruction to the instruction forwarding circuit 100 .
[0060] In step 306 , the instruction forwarding circuit 100 sends a fence instruction.
[0061] In step 308 , in response to the instruction forwarding circuit 100 receiving a response corresponding to the fence instruction, the instruction forwarding circuit 100 sends a data synchronization instruction.
[0062] In step 310 , the second computing unit 102 loads the data to be synchronized from the target address.
[0063] For example, in an exemplary embodiment, the first computing unit 101 sends a store instruction to a local L2 cache (L2 cache) to write and store the data to be synchronized in the L2 cache; the first computing unit 101 sends a fence instruction and an atomic instruction to the instruction forwarding circuit 100 respectively, so that the instruction forwarding circuit 100 sends the fence instruction and the atomic instruction instead of the first computing unit 101. Specifically, the instruction forwarding circuit 100 first sends a fence instruction to the L2 cache to ensure that all the data to be synchronized are written into the L2 cache; the instruction forwarding circuit 100 waits for and receives a response corresponding to the fence instruction returned by the L2 cache; and the instruction forwarding circuit 100 sends an atomic instruction to the L2 cache, so that the value of the flag variable of the L2 cache (i.e., the variable flag, which is usually initialized to 0) is set to 1, indicating that all the data to be synchronized have been written into the L2 cache. At the same time, the second computing unit 102 can determine whether data can be loaded from the L2 cache by cyclically loading the flag variable of the L2 cache. If the value of the mark variable of the L2 cache loaded by the second computing unit 102 is 0, it means that the data to be synchronized has not been fully written into the L2 cache, and the second computing unit 102 needs to reload the mark variable of the L2 cache after a predetermined time; if the value of the mark variable of the L2 cache loaded by the second computing unit 102 is 1, it means that the data to be synchronized has been fully written into the L2 cache, and the second computing unit 102 can load the data to be synchronized from the L2 cache.
[0064] Figure 4A and Figure 4B The timing diagrams of data synchronization between two computing units are respectively shown when the two computing units are not configured with the instruction forwarding circuit provided by the present invention and when the two computing units are configured with the instruction forwarding circuit provided by the present invention.
[0065] Figure 4AA timing diagram of data synchronization between two computing units that are not equipped with the instruction forwarding circuit provided by the present invention is shown. Figure 4B A timing diagram of data synchronization between two computing units configured with the instruction forwarding circuit provided by the present invention is shown.
[0066] like Figure 4A As shown, when the instruction forwarding circuit of the present invention is not configured, at time t0, the first computing unit sends the store instruction and the fence instruction to the target address in sequence, and then the first computing unit waits until a response corresponding to the fence instruction is received from the target address at time t1, and then the first computing unit sends an atomic instruction to the target address to instruct the second computing unit to load data from the target address; at time t2, the data to be synchronized is loaded into the second computing unit.
[0067] like Figure 4B As shown, when the instruction forwarding circuit of the present invention is configured, at time t0', the first computing unit first sends the store instruction to the target address, and sends the fence instruction and the atomic instruction to the instruction forwarding circuit. Subsequently, the instruction forwarding circuit sends the fence instruction to the target address, and waits until a response corresponding to the fence instruction is received from the target address at time t1', and then the instruction forwarding circuit sends the atomic instruction to the target address to instruct the second computing unit to load data from the target address; at time t2', the data to be synchronized is loaded into the second computing unit.
[0068] By comparison Figure 4A and Figure 4B It can be seen that Figure 4A In the example, the first computing unit is in a state of waiting for a response from time t0 to time t1 and cannot perform other operations. This is equivalent to the first computing unit being locked. It can only continue to perform operations related to the next instruction after receiving a response and sending an atomic instruction. Figure 4B In the example, after the first computing unit sends out the store instruction, the fence instruction, and the atomic instruction respectively at time t0', it can start to execute operations related to the next instruction, thereby improving the utilization and performance of the first computing unit. In addition, since the instruction forwarding circuit is closer to the target address than the first computing unit, the time for instruction transmission is shortened, so that the second computing unit can know more quickly that the data to be synchronized has been written to the target address, so that these data can be loaded into the second computing unit more quickly, so that the second computing unit can continue to perform other operations. As a result, the utilization and performance of the second computing unit are also improved accordingly.
[0069] In summary, through the instruction forwarding circuit for data synchronization between computing units as described above, the performance of the computing unit at the data production end can be released in advance, and at the same time, the time overhead of data synchronization between computing units can be reduced, so that the computing unit at the data consumption end can load data faster, thereby improving the performance of the computing unit at the data consumption end.
[0070] According to the inventive concept of the present invention, at least one instruction forwarding circuit for data synchronization between computing units as described above can also be configured in a computing system with a hierarchical relationship to further realize data synchronization across boards and across stream processor clusters. Specifically, according to an embodiment of the present invention, a computing module is provided, comprising: at least two computing units, and an instruction forwarding circuit configured for data synchronization between at least two computing units. According to an embodiment of the present invention, a computing circuit is also provided, comprising: at least two of the aforementioned computing modules, and an instruction forwarding circuit configured for data synchronization between at least two computing modules; and a computing system is provided, comprising: at least two of the aforementioned computing circuits, and an instruction forwarding circuit configured for data synchronization between at least two computing circuits. The following will be combined with Figures 5 to 7 The structures of the above-mentioned computing modules, computing circuits and computing systems, as well as the method flow for performing data synchronization thereon are described in detail.
[0071] Figure 5 An exemplary structural diagram of a computing system 500 according to an embodiment of the present invention is shown. It should be understood that the computing system 500 may further include additional circuits, modules and / or units not shown, and the scope of the present invention is not limited in this respect.
[0072] like Figure 5 As shown, the computing system 500 includes: an instruction forwarding circuit 505, and two computing circuits, computing circuit 510-1 and computing circuit 510-2 (collectively referred to as computing circuit 510), wherein the instruction forwarding circuit 505 is configured for data synchronization between computing circuits 510.
[0073] although Figure 5 The computing system 500 is shown to include two computing circuits. In other embodiments of the present invention, the computing system may include more than two computing circuits, which is not limited here.
[0074] Furthermore, each computing circuit may include: at least two computing modules, and an instruction forwarding circuit configured to synchronize data between the at least two computing modules. Figure 5As shown, the computing circuit 510-1 includes: an instruction forwarding circuit 515-1, and twelve computing modules 520-1. It should be understood that the number of computing modules that each computing circuit can include is not limited to the twelve as shown in the figure. In some other examples, the computing circuit can include, for example, ten or sixteen computing modules, and the present invention is not limited to this.
[0075] Furthermore, each computing module may include: at least two computing units, and an instruction forwarding circuit configured to synchronize data between at least two computing units. Figure 5 As shown, the computing module 520-1 includes: an instruction forwarding circuit 525-1, and four computing units 530-1. It should be understood that the number of computing units that each computing module can include is not limited to the four as shown in the figure. In some other examples, the computing circuit can include six, eight or twelve computing units, and the present invention is not limited to this.
[0076] According to an embodiment of the present invention, the computing circuit 510 may be, for example, a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processor (NPU), a deep learning processor (DPU), an acceleration processing unit (APU), a general purpose graphics processing unit (GPGPU), etc. For example, when the computing circuit 510 is a GPU, the computing module (e.g., computing module 520-1, computing module 520-2) may be a stream processor cluster, in which case the computing system 500 may also be referred to as a GPU server.
[0077] Regarding the instruction forwarding circuit 505, the instruction forwarding circuits 515-1 and 515-2, and the instruction forwarding circuit 525-1, they may be Figure 1 The instruction forwarding circuit 100 in the embodiment of the present invention, for the description of these instruction forwarding circuits, please refer to the above text and will not be repeated here.
[0078] Figure 6 and Figure 7 They are shown respectively Figure 5 A timing diagram of data synchronization between different computing units in a computing system 500.
[0079] Figure 6A timing diagram of data synchronization between two computing units located in different computing modules on the same computing circuit according to an embodiment of the present invention is shown. For example, the two computing units to be synchronized may be computing units of two different computing modules on the computing circuit 510-1. For the purpose of brevity and clarity, the two computing units may be respectively recorded as computing unit CU0 and computing unit CU1, wherein computing unit CU0 is located in the first computing module of the computing circuit 510-1, and computing unit CU1 is located in the second computing module of the computing circuit 510-1. In addition, an instruction forwarding circuit RP1 is configured on the first computing module, and an instruction forwarding circuit RP2 is configured on the computing circuit 510-1.
[0080] like Figure 6 As shown, according to an embodiment of the present invention, the computing unit CU0 may send a store instruction to the L2 cache of the second computing module to write and store the data to be synchronized in the L2 cache of the second computing module; the computing unit CU0 may first send the fence instruction and the atomic instruction to the instruction forwarding circuit RP1, respectively, so that the instruction forwarding circuit RP1 may replace the computing unit CU0 in sending the fence instruction and the atomic instruction. Specifically, the instruction forwarding circuit RP1 may first send the fence instruction to the L2 cache of the first computing module, and wait for the response corresponding to the fence instruction returned by the L2 cache of the first computing module. At the same time, the instruction forwarding circuit RP1 may also send the fence instruction to the instruction forwarding circuit RP2, so that the instruction forwarding circuit RP2 may send the fence instruction to the L2 cache of the second computing module. After receiving the response corresponding to the fence instruction returned by the L2 cache of the first computing module, the instruction forwarding circuit RP1 may send the atomic instruction to the instruction forwarding circuit RP2, so that the instruction forwarding circuit RP2 may send the atomic instruction to the L2 cache of the second computing module.
[0081] As can be seen from the above, up to now, the fence instruction and atomic instruction originally sent by the computing unit CU0 have been forwarded by the instruction forwarding circuit RP1 to the instruction forwarding circuit RP2, and will be sent by the instruction forwarding circuit RP2 to the target address storing the data to be synchronized (i.e., the L2 cache of the second computing module). Specifically, the instruction forwarding circuit RP2 sends the fence instruction to the L2 cache of the second computing module, and waits for the response corresponding to the fence instruction returned by the L2 cache of the second computing module to ensure that all the data to be synchronized are written into the L2 cache of the second computing module. In response to receiving the response corresponding to the fence instruction, the instruction forwarding circuit RP2 sends the atomic instruction to the L2 cache of the second computing module, so that the value of the variable flag of the L2 cache of the second computing module is set to 1, to indicate that all the data to be synchronized have been written into the L2 cache of the second computing module. Subsequently, the computing unit CU1 loads the data to be synchronized from the L2 cache of the second computing module.
[0082] In summary, by respectively configuring instruction forwarding circuits according to embodiments of the present invention on different computing modules on the same computing circuit, computing units of different computing modules can achieve data synchronization across computing modules via the instruction forwarding circuits of their respective computing modules.
[0083] Figure 7 A timing diagram of data synchronization between two computing units in computing modules located on different computing circuits according to an embodiment of the present invention is shown. For example, the two computing units to be synchronized may be: computing unit 530-1 of computing module 520-1 on computing circuit 510-1, and a computing unit of computing module 520-2 on computing circuit 510-2. For the purpose of brevity and clarity, these two computing units may be respectively referred to as computing unit CU2 and computing unit CU3, where computing unit CU2 is a computing unit of computing module CM1 on computing circuit CC1, and computing unit CU3 is a computing unit of computing module CM2 on computing circuit CC2. In addition, computing module CM1 (such as Figure 5 The computing module 520-1 is provided with an instruction forwarding circuit RP3, a computing circuit CC1 (such as Figure 5 The computing circuit 510-1 is provided with an instruction forwarding circuit RP4, and a computing system CS (such as Figure 5 The computing system 500 is provided with an instruction forwarding circuit RP5.
[0084] like Figure 7As shown, according to an embodiment of the present invention, the computing unit CU2 may send a store instruction to the L2 cache of the computing module CM2 where the computing unit CU3 is located, so as to write the data to be synchronized into the L2 cache of the computing module CM2. In addition, the computing unit CU2 first sends the fence instruction and the atomic instruction to the instruction forwarding circuit RP3, respectively, so that the instruction forwarding circuit RP3 replaces the computing unit CU2 in sending the fence instruction and the atomic instruction. Specifically, the instruction forwarding circuit RP3 first sends the fence instruction to the L2 cache of the computing module CM1, and waits for the response corresponding to the fence instruction returned by the L2 cache of the computing module CM1. At the same time, the instruction forwarding circuit RP3 also sends the fence instruction to the instruction forwarding circuit RP4, so that the instruction forwarding circuit RP4 sends the fence instruction to the L2 cache of the computing circuit CC1. After receiving the response corresponding to the fence instruction returned by the L2 cache of the computing module CM1, the instruction forwarding circuit RP3 sends the atomic instruction to the instruction forwarding circuit RP4, so that the instruction forwarding circuit RP4 sends the atomic instruction to the L2 cache of the computing circuit CC1. Therefore, the fence instruction and atomic instruction originally sent by the computing unit CU2 have been forwarded by the instruction forwarding circuit RP3 to the instruction forwarding circuit RP4, and will be sent by the instruction forwarding circuit RP4.
[0085] Specifically, the instruction forwarding circuit RP4 sends the fence instruction to the L2 cache of the computing circuit CC1, and waits for the response corresponding to the fence instruction returned by the L2 cache of the computing circuit CC1. At the same time, the instruction forwarding circuit RP4 also sends the fence instruction to the instruction forwarding circuit RP5, so that the instruction forwarding circuit RP5 can send the fence instruction to the L2 cache of the computing module CM2. After receiving the response corresponding to the fence instruction returned by the L2 cache of the computing circuit CC1, the instruction forwarding circuit RP4 sends the atomic instruction to the instruction forwarding circuit RP5, so that the instruction forwarding circuit RP5 can send the atomic instruction to the L2 cache of the computing module CM2. Thus, the instruction forwarding circuit RP4 forwards the fence instruction and the atomic instruction to the instruction forwarding circuit RP5, and then the instruction forwarding circuit RP5 will send these instructions to the L2 cache of the computing module CM2.
[0086] Specifically, the instruction forwarding circuit RP5 sends a fence instruction to the L2 cache of the computing module CM2, and waits for the response corresponding to the fence instruction returned by the L2 cache of the computing module CM2 to ensure that all the data to be synchronized are written into the L2 cache of the computing module CM2. In response to receiving the response corresponding to the fence instruction, the instruction forwarding circuit RP5 sends an atomic instruction to the L2 cache of the computing module CM2, so that the value of the variable flag of the L2 cache of the computing module CM2 is set to 1, indicating that all the data to be synchronized have been written into the L2 cache of the computing module CM2. Subsequently, the computing unit CU3 loads the data to be synchronized from the L2 cache of the computing module CM2.
[0087] In summary, by configuring the instruction forwarding circuit according to the embodiment of the present invention on the computing system, computing circuit and computing module respectively, data on different computing circuits in the computing system can be synchronized via the configured three-level instruction forwarding circuit.
[0088] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0089] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An instruction forwarding circuit for data synchronization between computing units, characterized in that: include: an instruction storage module, configured to store instructions received from the first computing unit, and send the instructions received from the first computing unit to the instruction scheduling module in sequence; as well as The instruction scheduling module is configured to schedule instructions to be sent based on at least the current instructions obtained from the instruction storage module, The instructions received from the first computing unit include: a fence instruction and a data synchronization instruction to be sent to a target address, The instructions to be sent include: the fence instruction and the data synchronization instruction, and Wherein, the instruction forwarding circuit is further configured as follows: Sending the fence instruction to the target address; and Based at least on the received response corresponding to the fence instruction, the data synchronization instruction is sent to the target address to facilitate data synchronization between computing units.
2. The instruction forwarding circuit according to claim 1, characterized in that: The data synchronization instruction is one of a synchronization instruction and an atomic operation instruction.
3. The instruction forwarding circuit according to claim 1, characterized in that: The instructions to be sent by the scheduler include: Checking the dependency of the current instruction obtained from the instruction storage module on other instructions; and In response to the current instruction not having dependencies with other instructions, the current instruction is sent.
4. The instruction forwarding circuit according to claim 3, characterized in that: The instructions to be sent by the scheduler also include: In response to the current instruction having a dependency on other instructions, determining whether all instructions having a dependency relationship with the current instruction have been received; and In response to receiving all instructions having a dependency relationship with the current instruction, the current instruction is sent.
5. The instruction forwarding circuit according to claim 4, characterized in that: The instructions to be sent by the scheduler also include: In response to not receiving all instructions having a dependency relationship with the current instruction, waiting until all instructions having a dependency relationship with the current instruction are received; and Send the current instruction.
6. The instruction forwarding circuit according to claim 1, characterized in that: The instruction scheduling module is further configured to: Based on the received response corresponding to the sent instruction, a deletion request for deleting the corresponding instruction is sent to the instruction storage module.
7. The instruction forwarding circuit according to claim 6, characterized in that: The instruction storage module is further configured to: Based on the deletion request received from the instruction scheduling module, the instruction corresponding to the deletion request is deleted.
8. The instruction forwarding circuit according to claim 1, characterized in that: The current instructions obtained from the instruction storage module include: A data item indicating the dependency of the current instruction, wherein the data item at least includes: an identifier of the current instruction, and information used to indicate the dependency of the current instruction.
9. A method for synchronizing data between computing units, characterized in that: include: The first computing unit sends a storage instruction to store the data to be synchronized to a target address; The first computing unit sends a fence instruction and a data synchronization instruction to the instruction forwarding circuit according to any one of claims 1 to 8; The instruction forwarding circuit sends the fence instruction to the target address; In response to the instruction forwarding circuit receiving a response corresponding to the fence instruction, the instruction forwarding circuit sends the data synchronization instruction to the target address; as well as The data to be synchronized is loaded from the target address by the second computing unit.
10. A computing module, characterized in that: include: at least two computing units; as well as A first instruction forwarding circuit, wherein the first instruction forwarding circuit is the instruction forwarding circuit according to any one of claims 1 to 8, and is configured to be used for data synchronization between the at least two computing units.
11. A computing circuit, characterized in that: include: At least two computing modules according to claim 10; as well as A second instruction forwarding circuit, wherein the second instruction forwarding circuit is the instruction forwarding circuit according to any one of claims 1-8, and is configured to be used for data synchronization between the at least two computing modules.
12. The calculation circuit according to claim 11, characterized in that: The computing circuit is a graphics processing unit, and the computing module is a stream processor cluster.
13. A computing system, characterized in that: include: At least two computing circuits according to claim 11; as well as A third instruction forwarding circuit, wherein the third instruction forwarding circuit is the instruction forwarding circuit according to any one of claims 1 to 8, and is configured to be used for data synchronization between the at least two computing circuits.
Citation Information
Patent Citations
High-parallelism computing system and instruction scheduling method thereof
CN110659070A
Instruction set processing system and method and electronic equipment
CN117608667A