Processing system with an integrated accelerator for a specific domain
By introducing the main processor and accelerator interface unit into the processing system, and using interface instructions to communicate with the accelerator interface unit, simple integration of domain-specific accelerators is achieved, solving the problem of major changes to the toolchain in the prior art, simplifying the integration process and reducing costs.
Patent Information
- Application Number
- CN202080106331.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-12-22
AI Technical Summary
Prior art requires significant changes to the toolchain when integrating domain-specific accelerators (DSAs) and processing systems on the same chip, resulting in complexity of the process.
A processing system is designed, including a main processor and an accelerator interface unit. The main processor communicates with the accelerator interface unit through interface instructions, generates and executes multiple commands, and is coupled to multiple domain-specific accelerators through multiple interface registers to achieve simple integration of the accelerator.
This solution requires only fewer changes to the toolchain, simplifying the integration process of DSA and processing systems, reducing integration complexity and cost.
Smart Images

Figure CN116438512B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of processing systems, and more particularly to a processing system having an integrated domain - specific accelerator. Background Art
[0002] An accelerator is a device designed to handle specific computationally intensive tasks. The main processor of a processing system typically offloads these computational tasks to the accelerator, thus allowing the main processor to continue executing other tasks. Graphics accelerators are perhaps the most well - known accelerators as they are found in almost all current - generation personal computers. However, there are many other different types of accelerators.
[0003] Traditionally, accelerators are coupled to and communicate with the main processor via an external bus such as a Peripheral Component Interconnect Express (PCIe) bus. However, accelerators known as domain - specific accelerators (DSAs) and processing systems have recently been integrated on the same chip.
[0004] However, integrating an accelerator and a processing system is a very important task. Partly because any changes to the instruction set architecture (ISA) to accommodate the instructions needed to operate the DSA using the processing system require significant changes to the toolchain, which is a complex tool used to verify the correct operation of the processing system. Therefore, a simple scheme for integrating a DSA and a processing system on the same chip is needed. Summary of the Invention
[0005] The present invention provides a simplified scheme for integrating a domain - specific accelerator (DSA) and a processing system on the same chip that requires only minor changes to the toolchain. The present invention provides a processing system including a main processor that decodes and fetches instructions and outputs interface instructions in response to the decoded and fetched instructions. The processing system further includes an accelerator interface unit coupled to the main processor. The accelerator interface unit includes a plurality of interface registers and a receiver coupled to the main processor and the plurality of interface registers. The receiver receives the interface instructions from the main processor, generates commands among the plurality of commands according to the interface instructions, determines the identified interface register among the plurality of interface registers according to the interface instructions, and outputs the commands to the identified interface register. The identified interface register executes the commands output by the receiver. The processing system further includes a plurality of domain - specific accelerators coupled to the plurality of interface registers. The domain - specific accelerator among the plurality of domain - specific accelerators receives information from the identified interface register and provides information to the identified interface register.
[0006] The present invention also includes a method of operating an accelerator interface unit. The method includes: receiving an interface instruction from a main processor; generating one of a plurality of commands according to the interface instruction; determining an identified interface register among a plurality of interface registers coupled to a plurality of domain-specific accelerators according to the interface instruction; and outputting the command to the identified interface register. The identified interface register executes the command output by the receiver.
[0007] The present invention also includes a method of operating a processing system. The method includes: using a main processor to decode and fetch an instruction; outputting an interface instruction in response to the decoding of the fetched instruction. The method further includes: receiving the interface instruction from the main processor; generating one of a plurality of commands according to the interface instruction; determining an identified interface register among a plurality of interface registers coupled to a plurality of domain-specific accelerators according to the interface instruction; and outputting the command to the identified interface register. The identified interface register executes the command output by the receiver.
[0008] The features and advantages of the present invention will be better understood by reference to the following detailed description and the accompanying drawings, which illustrate illustrative embodiments of the principles of the present invention. In order to better illustrate the technical means of the present application, so as to implement the present application according to the content of the specification, and to make the above and other objects, features and advantages of the present application more easily understood, specific embodiments of the present application are given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for illustrating the preferred embodiments and do not constitute a limitation to the present application. In addition, in each of the drawings, the same reference numerals are used to indicate the same parts. In the drawings:
[0010] Figure 1 is a block diagram illustrating an example of a processing system 100 according to the present invention.
[0011] Figure 2 is a flowchart illustrating an example of a method 200 of operating a main processor 110 according to the present invention.
[0012] Figures 3A - 3C is a flowchart illustrating an example of a method 300 of operating an accelerator interface unit 130 according to the present invention. DETAILED DESCRIPTION
[0013] Exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0014] As Figure 1 shown, a block diagram illustrating an example of a processing system 100 according to the present invention is presented. As Figure 1 shown, the processing system 100 includes a main processor 110, the main processor 110 includes a main decoder 112, a multi-word GPR 114 coupled to the main decoder 112, and an input stage 116 coupled to the main decoder 112 and the GPR 114. In addition, the main processor 110 further includes an execution stage 120 coupled to the input stage 116 and a switch 122 coupled to the main decoder 112, the execution stage 120, and the GPR 114.
[0015] As Figure 1 further shown, the processing system 100 further includes an accelerator interface unit 130 coupled to the input stage 116 and the switch 122 of the main processor 110. The accelerator interface unit 130 includes a receiver 132 coupled to the input stage 116, and a plurality of interface registers RG1-RGn each coupled to the receiver 132.
[0016] In operation, the receiver 132 receives interface instructions from the main processor 110, the main processor 110 decodes and fetches the instructions, and outputs the interface instructions to the receiver 132 in response to the decoding of the fetched instructions. The receiver 132 does not fetch instructions in the same manner as the decoder 112 of the main processor 110, but only receives interface instructions when the fetched instructions indicate that the main processor 100 provides interface instructions.
[0017] In addition, the receiver 132 generates one of a plurality of commands according to the interface instructions, determines the identified interface register among the plurality of interface registers according to the interface instructions, and outputs the command to the identified interface register that responds to the command.
[0018] In this example, the receiver 132 includes a front end 134 coupled to the input stage 116, an interface decoder 136 coupled to the front end 134, and a timeout counter 138 coupled to the front end 134. In addition, the interface registers RG1-RGn are respectively coupled to the front end 134 and the interface decoder 136.
[0019] In operation, the front end 134 receives interface instructions from the main processor 110, generates commands according to the interface instructions, broadcasts the commands to the interface registers RG, determines identification information according to the interface instructions, and outputs the identification information. The interface decoder 136 then determines the identified interface register according to the identification information, generates an enable signal, and outputs the enable signal to the identified interface register that responds by executing the commands broadcast by the front end 134.
[0020] Each interface register RG has a command register 140 and a response register 142. The command register 140 has a number of 32-bit command storage locations C1-Cx, and the response register 142 has a number of 32-bit response storage locations R1-Ry. Although this example shows each command register 140 as having the same number of command storage locations Cx, alternatively, the command register 140 may have a different number of command storage locations C. Similarly, although this example shows each response register 142 as having the same number of response storage locations Ry, alternatively, the response register 142 may have a different number of response storage locations R.
[0021] In addition, each interface register RG has a first-in first-out (FIFO) output queue 144 coupled to the command register 140 and a FIFO input queue 146 coupled to the response register 142. Each row of the FIFO output queue 144 has the same number of storage locations as the number of storage locations in the command register 140. Similarly, each row of the FIFO input queue 146 has the same number of storage locations as the number of storage locations in the response register 142.
[0022] In addition, the accelerator interface unit 130 includes an output multiplexer 150 coupled to the interface decoder 136 and each interface register RG. Optionally, the accelerator interface unit 130 may include an index out-of-bounds detector 152 coupled to the interface decoder 136. In addition, the accelerator interface unit 130 further includes a switch 154 coupled to the front end 134, and the switch 154 selectively couples the timeout counter 138, the multiplexer 150, or the index out-of-bounds detector 152 (when in use) to the switch 122.
[0023] In this example, the main decoder 112, GPR 114, input stage 116, and execution stage 120 are basically conventional elements commonly found in a main processor such as a RISC-V processor, with the main difference being the provision of an output from the input stage 116 to the accelerator interface unit 130. For example, in a typical RISC-V processor, the GPR has 32 storage locations, where the length of each storage location is 32 bits. In addition, the execution stage typically includes an arithmetic logic unit (ALU), a multiplier, and a load store unit (LSU).
[0024] AsFigure 1 As further shown, the processing system 100 also includes a plurality of domain-specific accelerators DSA1-DSAn coupled to output queue 144 and input queue 146 of interface registers RG1-RGn. The domain-specific accelerators DSA1-DSAn can be implemented together with various conventional accelerators such as video, vision, artificial intelligence, vector, and general matrix multiplication. In addition, the domain-specific accelerators DSA1-DSAn can operate at any desired clock frequency.
[0025] In operation, the domain-specific accelerators DSA1-DSAn receive respective numerical values from the output queue 144 of the corresponding interface registers RG1-RGn, interpret the respective numerical values as operation codes and operands, perform operations based on the operation codes and operands, and provide operation results back to the input queue 146 of the corresponding interface registers RG1-RGn.
[0026] As described in more detail below, many new instructions including DSA command write, push ready, push, read ready, pop, and read instruction are added to the traditional instruction set architecture (ISA). For example, RISC-V ISA has four basic instruction sets (RV32I, RV32E, RV64I, RV128I) and some extended instruction sets that can be added to the basic instruction sets to achieve specific goals (e.g., M, A, F, D, G, Q, C, L, B, J, T, P, V, N, H). In this example, RISC-V ISA is modified to include the new instructions in a custom extension set.
[0027] In addition, each new instruction uses the same instruction format as other instructions in the ISA. For example, RISC-V ISA has six instruction formats. One of the six formats is the I-type format, which has a seven-bit opcode field, a five-bit destination field identifying the destination location in the general-purpose register (GPR), a three-bit function field identifying the operation to be performed, a five-bit operand field identifying the location of the numerical value in the GPR, and a 12-bit immediate field.
[0028] Figure 2 A flowchart showing an example of a method 200 for operating the main processor 110 according to the present invention is shown. As Figure 2 shown, method 200 begins at 208, where the main processor 110 decodes the fetched instruction and outputs an interface instruction in response to the decoding of the fetched instruction.
[0029] In this example, the fetched instruction executed by the main processor 110 is an instruction in the instruction set architecture including the new instructions of the present invention. The interface instruction can then be the same as the fetched instruction, including only the selected fields in the fetched instruction, or include the information of the fetched instruction in a different format. In this example, the interface instruction is the same as the fetched one.
[0030] When the main decoder 112 decodes a DSA command write instruction for a new instruction, method 200 moves to 210. The DSA command write instruction includes an operand field that defines a storage location in GPR 114 that holds the DSA value, a function field that indicates the accelerator interface unit 130 to perform a write operation, and an immediate field that identifies interface register RG and command storage location C within the command register 140 of the identified interface register RG. (Alternatively, interface register RG and command storage location C may be in two separate fields.)
[0031] In addition, in this example, the DSA command write instruction further includes an opcode field that indicates the main decoder 112 of the main processor 110 to move the DSA command write instruction and the DSA value held in the storage location in GPR 114 to the accelerator interface unit 130 via the input stage 116.
[0032] In addition, when using the optional index-out-of-bounds detector 152, the DSA command write instruction includes a destination field that identifies the index-out-of-bounds storage location in GPR 114, and the opcode field further indicates the main decoder 112 to couple the switch 122 to the switch 154 and the index-out-of-bounds storage location in GPR 114.
[0033] For example, in the I-type format of a RISC-V instruction, a five-bit operand field may identify the location of the DSA value in GPR 114, a three-bit function field may identify the write operation to be performed by the accelerator interface unit 130, and a 12-bit immediate field may hold the identification of interface register RG and the identification of command storage location C. The destination register field may in turn identify the index-out-of-bounds storage location.
[0034] In addition, a seven-bit opcode field of a RISC-V instruction may indicate the main decoder 112 to move the DSA command write instruction and the DSA value held in the storage location of GPR 114 to the accelerator interface unit 130 via the input stage 116, and when using the optional index-out-of-bounds detector 152, the seven-bit opcode field couples the switch 122 to the switch 154 and the index-out-of-bounds storage location in GPR 114.
[0035] The index-out-of-bounds storage location may hold the index-out-of-bounds status of the identified interface register. When not using the index-out-of-bounds detector 152, method 200 returns to 208. When using the index-out-of-bounds detector 152, method 200 moves to 212 to check the index-out-of-bounds storage location, returns to 208 when there is no index-out-of-bounds status condition, and generates an error when there is an index-out-of-bounds status condition.
[0036] Figures 3A - 3C A flowchart showing an example of a method 300 for operating an accelerator interface unit 130 in accordance with the present invention is shown. AsFigure 3A As shown, method 300 begins at 308, where the front end 134 of the accelerator interface unit 130 detects and identifies DSA command instructions received from the input stage 116.
[0037] When the DSA command write instruction of the new instruction is identified, method 300 moves to 310, where the front end 134 extracts the function field and the immediate field from the DSA command write instruction. In addition, the front end 134 receives the DSA value held in the storage location of GPR114 from the input stage 116.
[0038] In addition, the front end 134 forwards the immediate field to the interface decoder 136, generates a write command from the function field, and broadcasts the write command and the DSA value to all interface registers RG. Also, when using the index-out-of-bounds detector 152, the front end 134 couples the index-out-of-bounds detector 152 to the switch 154.
[0039] Next, method 300 moves to 312, where the interface decoder 136 identifies the interface register and the command storage location C of the command register 140 of the identified interface register RG from the immediate field of the DSA command write instruction, and outputs an encoded enable signal representing the identified interface register RG to all the identified interface registers. (Instead of the encoded enable signal, it is possible to choose to send individual enable signals to each interface register. The encoded enable signal slightly increases the complexity of the interface register RG, but reduces the number of traces.) After this, method 300 moves to 314, where the identified interface register RG writes the DSA value to the identified command storage location C of the command register 140 of the identified interface register RG in response to the identification of the enable signal.
[0040] When using the index-out-of-bounds detector 152, method 300 moves from 312 to 316 to determine whether the interface register and / or the command storage location is out of index. For example, if there are three interface registers RG and the immediate field of the DSA command write instruction identifies the fifth interface register, the index-out-of-bounds detector 152 detects an index-out-of-bounds condition. Similarly, if there are four command storage locations C1 - C4 and the immediate field identifies a fifth command storage location, the index-out-of-bounds detector 152 detects an index-out-of-bounds condition.
[0041] When one or both exceed the index, method 300 moves to 318 to output the value to GPR114 beyond the index storage location through switch 154 and switch 122. Then, the index-out-of-storage location can be checked to determine if there is an error. When both are within the index, the method moves from 316 to 314, and the identified interface register RG writes the DSA value to the identified command storage location C in command register 140 of the identified interface register RG in response to the enable signal. Method 300 returns from 314 to 308 to wait for another instruction.
[0042] Referring again to Figure 2 , method 200 restarts at 208, and main decoder 112 decodes another fetch instruction such as another DSA command write instruction. In the first embodiment, the write operation includes more than two DSA command write instructions. The DSA value identified by the operand field in one DSA command write instruction in GPR114 represents the DSA opcode (the operation performed by the DSA), and the DSA value identified by the operand field in another DSA command write instruction in GPR114 represents the DSA operand (the value to be operated on).
[0043] In the first embodiment, main decoder 112 and front end 134 process the DSA opcode and DSA operand in the same way, and it is not possible or necessary to separate the two. The DSA command write instruction basically moves the word from GPR114 to command register 140 of the identified interface register RG.
[0044] Several DSA command write instructions are used to fill all the command storage locations C in command register 140. It is determined by the domain-specific accelerator DSA coupled to the identified interface register RG whether the DSA value is a DSA opcode or a DSA operand, and it is the programmer's responsibility to ensure that command register 140 is correctly assembled.
[0045] Alternatively, in the second embodiment, the DSA opcode and DSA operand can be combined and stored together in a storage location in GPR114. For example, several bits in a 32-bit storage location in GPR114 can be allocated to represent the DSA opcode (the operation performed by the DSA), and the remaining bits can represent the DSA operand (the value operated on by the DSA).
[0046] Referring again to Figure 2 , when main decoder 112 decodes another DSA command instruction of a new instruction, method 200 moves to 220, and the DSA command push-ready instruction is decoded. The DSA command push-ready instruction includes a function field indicating the accelerator interface unit 130 to perform the push-ready operation, an immediate field identifying the interface register RG, and a destination field identifying the push-ready storage location in GPR114.
[0047] The DSA command push - ready instruction further includes an opcode field that instructs the main decoder 112 to move the DSA command push - ready instruction via the input stage 116 to the accelerator interface unit 130 and couple the switch 122 to the push - ready storage location in the switch 154 and the GPR 114. The push - ready storage location holds the push - ready state of the identified interface register.
[0048] For example, in the I - type format of the RISC - V instruction, a three - bit function field can identify the push - ready operation to be performed by the accelerator interface unit 130, and a 12 - bit immediate field can hold the identification of the interface register RG. The destination field can further hold the identification of the push - ready storage location in the GPR 114. Additionally, a seven - bit opcode field can instruct the main decoder 112 to move the DSA command push - ready instruction via the input stage 116 to the accelerator interface unit 130 and couple the switch 122 to the push - ready storage location in the switch 154 and the GPR 114.
[0049] Refer again to Figure 3A , the method 300 restarts at 308, and the front - end 134 of the accelerator interface unit 130 detects and identifies another interface instruction received from the input stage 116. When identifying the DSA command push - ready instruction of the new instruction, the method 300 moves to 320, and the front - end 134 extracts the function field and the immediate field from the DSA command push - ready instruction.
[0050] In addition, the front - end 134 forwards the immediate field of the DSA command push - ready instruction to the interface decoder 136, generates a push - ready command according to the function field, broadcasts the push - ready command to all interface registers RG, and couples the output multiplexer 150 to the switch 154.
[0051] Next, the method 300 moves to 322, and the interface decoder 136 identifies the interface register from the immediate field of the DSA command push - ready instruction. The interface decoder 136 also outputs a select signal to the multiplexer 150 and outputs an encoded enable signal indicating the identified interface register to all interface registers RG. After this, the method 300 moves to 324, and the identified interface register RG determines whether the output queue 144 of the identified interface register RG can accept the value held in the command register 140 in response to the identification of the encoded enable signal.
[0052] When the output queue 144 of the identified interface register RG can accept the value held in the command register 140, method 300 moves to 326, the identified interface register RG outputs a ready value to the output multiplexer 150, and the output multiplexer 150 passes the ready value to the push-ready position in GPR114 via the switch 154 and the switch 122 in response to the select signal.
[0053] When the output queue 144 of the identified interface register RG is not ready to accept these values, method 300 moves to 328, the identified interface register RG outputs a not-ready value to the multiplexer 150, and the multiplexer 150 passes the not-ready value to the push-ready position in GPR114 via the switch 122 and the switch 154 in response to the select signal. Then, a loop process is executed until the output ready signal. Alternatively, the loop process may also include other steps. After the ready value has been output and waiting for the next instruction, method 300 returns to 308.
[0054] Refer again to Figure 2 , method 200 moves from 220 to 222 to check the push-ready storage location in GPR114 to determine the push-ready state of the identified interface register. Method 200 loops until the push-ready state indicates that the identified interface register is ready to accept the push command. Alternatively, the loop process may also include other steps. When the push-ready state indicates ready, method 200 returns to 208, and the main decoder 112 decodes another fetch instruction.
[0055] When the DSA command push command of the new instruction is decoded, method 200 moves to 230. The DSA command push command includes a timeout field that locates the first timeout storage location in GPR114 that holds the first timeout value, a function field that indicates the accelerator interface unit 130 to perform the push operation, an immediate number field that identifies the interface register RG and the command storage location C in the command register 140 of the identified interface register RG, and a destination field that identifies the push timeout storage location in GPR114.
[0056] In addition, the DSA command push command includes an opcode field that indicates the main decoder 112 to move the DSA command push command and the first timeout value held in the first timeout storage location of GPR114 to the accelerator interface unit 130 via the input stage 116, and to couple the switch 122 to the switch 154 and the push timeout storage location in GPR114. The push timeout storage location holds the first timeout state.
[0057] For example, in the I-type format of RISC-V instructions, a five-bit operand field can identify the first timeout storage location of the first timeout value in GPR114, a three-bit function field can identify that the push operation is performed by the accelerator interface unit 130, and a 12-bit immediate field can hold the identification of the interface register RG and the command storage location C. The destination register field can then identify the push timeout storage location. Additionally, a seven-bit opcode field can instruct the main decoder 112 to move the DSA command push command and the first timeout value held in the first timeout storage location to the accelerator interface unit 130 via the input stage 116, and couple the switch 122 to the switch 154 and the push timeout storage location in GPR114.
[0058] Reference Figure 3A and Figure 3B , method 300 resumes at 308, and the front end 134 of the accelerator interface unit 130 detects and identifies another interface instruction received from the input stage 116. When the DSA command push command of the new instruction is identified, method 300 moves to 330, and the front end 134 extracts the function field and the immediate field from the DSA command push command.
[0059] In addition, the front end 134 forwards the immediate field of the DSA command push command to the interface decoder 136, generates a push command according to the function field, and broadcasts the push command to all interface registers RG. Additionally, the front end 134 receives the first timeout value held in the first timeout storage unit in GPR114 from the input stage 116, couples the timeout circuit 138 to the switch 154, and forwards the first timeout value to the timeout counter 138, and the timeout counter 138 starts counting.
[0060] Next, method 300 moves to 332, the interface decoder 136 identifies the interface register RG and the command storage location C from the middle field of the DSA command push command, and outputs an encoded enable signal indicating the identified interface register to all interface registers RG.
[0061] After that, method 300 moves to 334, and in response to the identification of the encoded enable signal, the identified interface register RG stacks one or more numerical values from the identified command storage location C in the command register 140 of the identified interface register RG into the output queue 144 of the identified interface register RG.
[0062] In addition, the identified interface register RG outputs a transfer signal to the corresponding domain-specific accelerator DSA, indicating that one or more numerical values are in the output queue 144 and ready for transfer. The transfer signal can be a notification signal to the corresponding domain-specific accelerator DSA, or an acknowledgement of a query from the corresponding domain-specific accelerator DSA.
[0063] After that, the identified interface register RG transfers the value to the corresponding domain specific accelerator DSA using any conventional handshaking protocol. Once the relevant DSA has received all the required operation codes and required operands, the DSA will perform the required task and return the response value to the input queue 146 of the identified interface register RG in a manner similar to how individual values are received from the output queue 144.
[0064] In addition, method 300 moves to 336 when the timeout counter 138 reaches, the timeout counter 138 outputs a timeout value to the switch 154, and the switch 154 passes the timeout value to the push timeout storage location in the GPR114 via the switch 154 and the switch 122.
[0065] Referring again to Figure 2 , method 200 moves from 230 to 232 to check the push timeout storage location in the GPR114 to determine the first timeout status of the identified interface register. When the first timeout status is set, this status indicates that an error has occurred. When the first timeout status is not set, method 200 returns to 208 to decode the next fetch instruction.
[0066] When the DSA command read ready instruction of the new instruction is decoded, method 200 moves from 208 to 240. The DSA command read ready instruction includes a function field indicating the read ready operation to be performed by the accelerator interface unit 130, an immediate digit field identifying the interface register, and a destination field identifying the read ready storage location in the GPR114.
[0067] The DSA command read ready instruction also includes an operation code field that indicates the main decoder 112 to move the DSA command read ready instruction to the accelerator interface unit 130 via the input stage 116 and couple the switch 122 to the read ready storage location in the GPR114. The read ready storage location holds the read ready status of the identified interface register.
[0068] For example, in the I - type format of a RISC - V instruction, a three - bit function field can identify the read ready operation to be performed by the accelerator interface unit 130, and a 12 - bit immediate digit field can hold the register identification. The destination register field can then identify the read ready storage location. In addition, a seven - bit operation code field can indicate the main decoder 112 to move the DSA command read ready instruction to the accelerator interface unit 130 via the input stage 116 and couple the switch 122 to the switch 154 and the read ready location in the GPR114.
[0069] Referring again to Figure 3A and Figure 3B, Method 300 restarts at 308, and the front end 134 of the accelerator interface unit 130 detects and identifies another instruction received from the input stage 116. When identifying the DSA command read-ready instruction of the new instruction, Method 300 moves to 340, and the front end 134 extracts the function field and the immediate field from the DSA command read-ready instruction. In addition, the front end 134 forwards the immediate field of the DSA command read-ready instruction to the interface decoder 136, generates a read-ready command according to the function field, broadcasts the read-ready command to all interface registers RG, and couples the output multiplexer 150 to the switch 154.
[0070] Next, Method 300 moves to 342, and the interface decoder 136 identifies the interface register RG from the immediate field of the DSA command read-ready instruction. The interface decoder 136 also outputs a selection signal to the multiplexer 150 and outputs an encoded enable signal indicating the identified interface register to all interface registers RG. After that, Method 300 moves to 344. In response to the identification of the enable signal, the identified interface register RG determines whether the input queue 146 of the identified interface register RG holds the response value to be read received from the corresponding specific domain accelerator DSA.
[0071] When the input queue 146 of the identified interface register RG holds the value to be read, Method 300 moves to 346, and the identified interface register RG outputs a read-ready value to the output multiplexer 150. In response to the selection signal, the output multiplexer 150 passes the read-ready value to the read-ready storage location in the GPR114 via the switch 154 and the switch 122.
[0072] When the input queue 146 of the identified interface register RG is empty, Method 300 moves to 348, and the identified interface register RG outputs an unready value to the multiplexer 150. In response to the selection signal, the multiplexer 150 passes the unready value to the read-ready storage location in the GPR114 via the switch 154 and the switch 122. Then, a loop process is executed until a read-ready value has been output. Alternatively, the loop process may also include other steps. After a read-ready value has been output to wait for the next instruction, Method 300 returns to 308.
[0073] Referring again to Figure 2 , Method 200 moves from 240 to 242 to check the read-ready storage location in the GPR114 to determine the read-ready state of the identified interface register. Method 200 loops until the read-ready state indicates that the input queue 146 of the identified interface register RG holds the value to be read. Alternatively, the loop process may also include other steps.
[0074] After that, method 200 returns to 208 to decode the next fetched instruction. When the DSA command pop instruction of the new instruction is decoded, method 200 moves to 250. The DSA command pop instruction includes a timeout field that defines a second timeout storage location holding a second timeout value in GPR114, a function field that indicates the accelerator interface unit 130 to perform a pop operation, an immediate number field that identifies the interface register RG and the response storage location R, and a destination field that identifies the pop timeout storage location in GPR114.
[0075] In addition, the DSA command pop instruction includes an opcode field that indicates the main decoder 112 to move the DSA command pop instruction and the second timeout value held in the second timeout storage location in GPR114 to the accelerator interface unit 130 via the input stage 116, and couple the switch 122 to the switch 154 and the pop timeout storage location in GPR114. The pop timeout storage location holds the second timeout state.
[0076] For example, in the I-type format of a RISC-V instruction, a five-bit operand field can identify the second timeout storage location of the second timeout value in GPR114, a three-bit function field can identify the pop operation performed by the accelerator interface unit 130, and a 12-bit immediate number field can identify the interface register RG and the response storage location R in the response register 142 of the identified interface register RG. The destination register field can further identify the pop timeout storage location. In addition, a seven-bit opcode field can indicate the main decoder 112 to move the DSA command pop instruction and the second timeout value held in the second timeout storage location in GPR114 to the accelerator interface unit 130 via the input stage 116.
[0077] Reference Figures 3A - 3C Referring, method 300 restarts at 308, and the front end 134 of the accelerator interface unit 130 detects and identifies another interface instruction received from the input stage 116. When identifying the DSA command pop instruction of the new instruction, method 300 moves to 350, and the front end 134 extracts the function field and the immediate number field from the DSA command pop instruction.
[0078] In addition, the front end 134 forwards the immediate number field of the DSA command pop instruction to the interface decoder 136, generates a pop command according to the function field, and broadcasts the pop command to all interface registers RG. In addition, the front end 134 receives the second timeout value held in the second timeout storage unit in GPR114 from the input stage 116, couples the timeout circuit 138 to the switch 154, and forwards the second timeout value to the timeout counter 138, and the timeout counter 138 starts counting.
[0079] Next, method 300 moves to 352, where interface decoder 136 identifies the interface register and the response storage location R from the immediate digit field of the DSA command pop instruction, and outputs an encoded enable signal indicating the identified interface register to all interface registers RG. After that, method 300 moves to 354, and in response to receiving the encoded enable signal, the identified interface register RG pops one or more response words from the input queue 146 of the identified interface register RG into one or more response storage locations R in the response register 142 of the identified interface register RG.
[0080] In addition, when the timeout counter 138 reaches, method 300 moves to 356, where the timeout counter 138 outputs a second timeout value to switch 154, and switch 154 passes the timeout value to the pop timeout storage location in GPR 114 via switch 122.
[0081] Referring again to Figure 2 , method 200 moves from 250 to 252 to check the pop timeout storage location to determine the second timeout status of the identified interface register. When the second timeout status is set, this status indicates that an error has occurred. When the second timeout status is not set, method 200 returns to 208 to decode the next fetch instruction.
[0082] When decoding the DSA command read instruction of a new instruction, method 200 moves from 208 to 260. The DSA command read instruction includes a function field indicating the accelerator interface unit 130 to perform a read operation, an immediate digit field identifying the interface register RG and the response storage location R in the response register 142 of the identified interface register RG, and a destination field identifying the read storage location in GPR 114.
[0083] In addition, the DSA command read instruction includes an opcode field, which indicates that the main decoder 112 moves the DSA command read instruction to the accelerator interface unit 130 via the input stage 116 and couples switch 122 to the read storage location in switch 154 and GPR 114. For example, in the I-type format of a RISC-V instruction, a three-bit function field can identify the read operation to be performed by the accelerator interface unit 130, and a 12-bit immediate digit field can identify the interface register RG and the response storage location R in the response register 142 of the identified interface register RG.
[0084] The destination register field can further identify the read storage location. In addition, a seven-bit opcode field can indicate that the main decoder 112 moves the DSA command read instruction to the accelerator interface unit 130 via the input stage 116 and couples switch 122 to the read storage location in switch 154 and GPR 114. Reading the storage location in GPR 114 holds the value returned from the DSA.
[0085] Referring again to Figures 3A - 3C , method 300 restarts at 308, and the front end 134 of the accelerator interface unit 130 detects and identifies another interface instruction received from the input stage 116. When identifying the DSA command read instruction of the new instruction, method 300 moves to 360 to extract the function field and the immediate field from the DSA command read instruction. In addition, the front end 134 forwards the immediate field of the DSA command read instruction to the interface decoder 136, generates a read command according to the function field, and broadcasts the read command to all interface registers RG. In addition, the front end 134 couples the output multiplexer 150 to the switch 154.
[0086] Next, method 300 moves to 362, and the interface decoder 136 identifies the interface register and the response storage location R from the immediate field of the DSA command read instruction. In addition, the interface decoder 136 outputs a selection signal to the output multiplexer 150 and outputs an encoded enable signal indicating the identified interface register to all interface registers RG.
[0087] After that, method 300 moves to 364. In response to the identification of the enable signal, the identified interface register RG transfers the response word from the response storage location R to the output multiplexer 150, and the output multiplexer 150 transfers the response word R to the switch 122 in response to the selection signal. Then, the response word is transferred to the read storage location in GPR114 via the switch 122.
[0088] The present invention provides many advantages. One of the greatest advantages is that the new instruction is general, so compared with other methods (such as multi-input multi-output (MIMO) methods or ISA extensions using specific instructions), only a small modification to the existing tool chain is required. In addition, the interaction delay, computational scalability, and multi-accelerator cooperation are all good. In addition, the programmable granularity is also good.
[0089] Now, various embodiments of the present disclosure have been described in detail, and examples thereof are shown in the accompanying drawings. Although described in conjunction with various embodiments, it should be understood that these various embodiments are not intended to limit the present disclosure. On the contrary, the present disclosure is intended to cover alternatives, modifications, and equivalents that may be included within the scope of the present disclosure as interpreted according to the claims. In addition, in the foregoing detailed description of the various embodiments of the present disclosure, many specific details are set forth to provide a thorough understanding of the present disclosure. However, those of ordinary skill in the art will recognize that the present disclosure may be practiced without these specific details or their equivalents. In other instances, well-known methods, processes, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the various embodiments of the present disclosure.
[0090] It should be noted that although for clarity, the method may be described herein as a series of numbered operations, the numbers do not necessarily prescribe the order of the operations. It should be understood that some operations may be skipped, performed in parallel, or performed without maintaining a strict order. The drawings showing various embodiments in accordance with the present disclosure are semi-diagrammatic and not drawn to scale, and in particular, some dimensions are exaggerated for clarity of presentation and shown in the drawings. Similarly, although for ease of description, the views in the drawings generally represent similar orientations, such depictions in the drawings are largely arbitrary. Generally, the various embodiments in accordance with the present disclosure may be operated in any direction.
[0091] Certain portions of the detailed description are presented in terms of processes, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. These descriptions and representations are used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. In the present disclosure, a process, logic block, process, etc. is considered to be a self-consistent sequence of operations or instructions leading to a desired result. These operations are those that manipulate physical quantities. Usually, but not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transmitted, combined, compared, and otherwise manipulated in a computing system. Sometimes, mainly for reasons of common usage, it has proven convenient to refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, etc.
[0092] However, it should be borne in mind that all such and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless clearly stated otherwise explicitly from the following discussion, it should be understood that throughout this disclosure, discussions using terms such as "generate", "determine", "allocate", "aggregate", "utilize", "virtualize", "process", "access", "execute", "store", etc. refer to the actions and processes of a computer system or similar electronic computing device or processor. A computer system or similar electronic computing device or processor manipulates data represented as physical (electronic) quantities in a computer system memory, registers, other such information storage, and / or other computer-readable media, and converts it into other data similarly represented as physical quantities in a computer system memory or register or other such information storage, transmission, or display device.
[0093] The technical solutions of the embodiments of the present application have been clearly and completely described above in conjunction with the accompanying drawings of the embodiments of the present application. It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that these numbers may be interchanged at appropriate times so that the embodiments of the present invention described herein may be implemented in an order different from that illustrated or described herein.
[0094] If the functions described in the method of this embodiment are implemented in the form of software function units and sold or used as independent products, they can be stored in a storage medium readable by a computing device. Based on this understanding, the part or parts of the technical solutions that contribute to the prior art in the embodiments of this application can be embodied in the form of a software product stored in a storage medium, including a plurality of instructions for causing a computing device (which can be a personal computer, a server, a mobile computing device, a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application. The above storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, optical discs, etc., which can store program codes.
[0095] The various embodiments in the specification of this application are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. The described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without departing from the inventive scope of this application belong to the scope protected by this application.
[0096] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments are obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not limited to the embodiments shown herein, but to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A processing system, comprising: A main processor that decodes and fetches instructions and outputs interface instructions in response to the decoded and fetched instructions; An accelerator interface unit coupled to the main processor, the accelerator interface unit comprising: A plurality of interface registers; and A receiver coupled to the main processor and the plurality of interface registers, the receiver receiving the interface instructions from the main processor, generating a command among a plurality of commands according to the interface instructions, determining an identified interface register among the plurality of interface registers according to the interface instructions, and outputting the command to the identified interface register, and the identified interface register executing the command output by the receiver; and A plurality of domain-specific accelerators coupled to the plurality of interface registers, wherein a domain-specific accelerator among the plurality of domain-specific accelerators receives information from the identified interface register and provides information to the identified interface register, wherein each interface register comprises: A command register having a plurality of command storage locations; An output queue coupled to the command register and a domain-specific accelerator among the plurality of domain-specific accelerators; A response register having a plurality of response storage locations; and An input queue coupled to the response register and the domain-specific accelerator.
2. The processing system according to claim 1, wherein, The main processor comprises: A main decoder that decodes and fetches instructions; General-purpose registers coupled to the main decoder; An input stage coupled to the main decoder, the general-purpose registers, and a front end; and An execution stage coupled to the input stage.
3. The processing system according to claim 1, wherein, The receiver comprises: A front end coupled to the main processor, the front end receiving the interface instructions from the main processor, generating a command according to the interface instructions, broadcasting the command to the plurality of interface registers, determining identification information according to the interface instructions, and outputting the identification information; and An interface decoder coupled to the front end, the interface decoder determining the identified interface register according to the identification information, generating an enable signal, and outputting the enable signal to the identified interface register.
4. The processing system according to claim 3, wherein, When the interface instruction is a write instruction, the front end generates a write command among the plurality of commands according to the interface instruction, receives a value from the main processor in addition to the interface instruction, and broadcasts the write command and the value to the plurality of interface registers; and The identified interface register writes the value into the command register of the identified interface register in response to the enable signal.
5. The processing system according to claim 4, wherein, The accelerator interface unit further comprises a multiplexer coupled to the interface decoder and the plurality of interface registers.
6. The processing system according to claim 5, wherein, When the interface instruction is a push-ready instruction, the front end generates a push-ready instruction among the plurality of instructions according to the interface instruction and broadcasts the push-ready instruction to the plurality of interface registers; The interface decoder outputs a selection signal in addition to the enable signal in response to the determination of the identified interface register; The identified interface register determines whether the output queue of the identified interface register can accept the value stored in the command register in response to the enable signal. When the output queue of the identified interface register can accept the value in the command register, it outputs a ready value to the multiplexer, and when the output queue of the identified interface register cannot accept the value stored in the command register, it outputs a not-ready value to the multiplexer; and The multiplexer transmits a ready signal or a not-ready signal in response to the selection signal.
7. The processing system according to claim 6, wherein, When the interface instruction is a push instruction, the front end generates a push command among the multiple commands according to the interface instruction and broadcasts the push command to the multiple interface registers; and The identified interface register pushes the value stored in the command register onto the output queue in response to the enable signal.
8. The processing system according to claim 5, wherein, When the interface instruction is a read-ready instruction, the front end generates a read-ready command among the multiple commands according to the interface instruction and broadcasts the read-ready command to the multiple interface registers; The interface decoder outputs a selection signal in addition to the enable signal in response to the determination of the identified interface register; The identified interface register determines whether the input queue of the identified interface register holds a response value from the above-mentioned domain-specific accelerator. When the input queue of the identified interface register holds the response value, it outputs a ready value to the multiplexer, and when the input queue of the identified interface register does not hold the response value, it outputs a not-ready value to the multiplexer; and The multiplexer transmits a ready signal or a not-ready signal in response to the selection signal.
9. The processing system according to claim 8, wherein, When the interface instruction is a pop instruction, the front end generates a pop command among the multiple commands according to the interface instruction and broadcasts the pop command to the multiple interface registers; and The identified interface register pops the response value in the input queue from the domain-specific accelerator into the response register of the identified interface register in response to the enable signal.
10. The processing system according to claim 9, wherein, When the interface instruction is a read instruction, the front end generates a read command among the multiple commands according to the interface instruction and broadcasts the read command to the multiple interface registers; The interface decoder outputs the selection signal in addition to the enable signal in response to the determination of the identified interface register; The identified interface register outputs the response value held in the response register to the multiplexer in response to the enable signal; and The multiplexer transmits the response value in response to the selection signal.
11. A method of operating an accelerator interface unit, the method comprising: Receiving an interface instruction from a main processor; Generating a command among multiple commands according to the interface instruction; Determine the identified interface register among multiple interface registers coupled to a plurality of domain - specific accelerators according to the interface instruction, wherein each interface register includes: a command register having a plurality of command storage locations; an output queue coupled to the command register and a domain - specific accelerator among the plurality of domain - specific accelerators; a response register having a plurality of response storage locations; and an input queue coupled to the response register and the domain - specific accelerator; and Output the command to the identified interface register, and the identified interface register executes the command output by the receiver.
12. The method according to claim 11, wherein: Determining the identified interface register includes: Determine identification information from the interface instruction; Determine the identified interface register according to the identification information; Generate an enable signal and output the enable signal to the identified interface register; and Outputting the command to the identified interface register includes: Broadcast the command to the plurality of interface registers.
13. The method according to claim 11, further comprising: When the interface instruction is a write instruction, generate a write command among the plurality of commands from the interface instruction; In addition to the interface instruction, receive a value from the main processor; Broadcast the write command and the value to the plurality of interface registers; and In response to the enable signal, write the value into the command register.
14. The method according to claim 13, further comprising: When the interface instruction is a push - ready instruction, generate a push - ready command among the plurality of commands from the interface instruction and broadcast the push - ready command to the plurality of interface registers; In response to the determination of the identified interface register, output a selection signal in addition to the enable signal; In response to the enable signal, determine whether the output queue of the identified interface register can accept the value stored in the command register. Output a ready value when the output queue of the identified interface register can accept the value stored in the command register, and output a not - ready value when the output queue of the identified interface register cannot accept the value stored in the command register; and Transmit a ready signal or a not - ready signal in response to the selection signal.
15. The method according to claim 13, wherein When the interface instruction is a push instruction, further include: Generate a push command among the plurality of commands according to the interface instruction and broadcast the push instruction to the plurality of interface registers; In response to the determination of the identified interface register, output a selection signal in addition to the enable signal; and In response to the enable signal, push the value stored in the command register into the output queue.
16. The method according to claim 13, wherein When the interface instruction is a read - ready instruction, generate a read - ready command among the plurality of commands from the interface instruction and broadcast the read - ready command to the plurality of interface registers in response to the read - ready instruction; In response to the determination of the identified interface register, output a selection signal in addition to the enable signal; Determine whether the input queue of the interface register holds a response value from the domain - specific accelerator. Output a ready value when the input queue of the identified interface register holds a response value, and output a not - ready value when the input queue of the identified interface register does not hold a response value; and Transmit a ready signal or a not-ready signal in response to the selection signal.
17. The method according to claim 16, wherein When the interface instruction is a pop instruction, generate a pop instruction among the multiple instructions according to the interface instruction, and broadcast the pop instruction to the multiple interface registers in response to the pop instruction; and In response to the enable signal, pop the response value from the domain-specific accelerator into the response register of the identified interface register.
18. The method according to claim 17, wherein, When the interface instruction is a read instruction, generate a read command among the multiple commands according to the interface instruction, and broadcast the read command to the multiple interface registers in response to the read instruction; In response to the determination of the identified interface register, output a selection signal in addition to the enable signal; Output the response value held in the response register in response to the enable signal; and Transmit the response value in response to the selection signal.
19. A method for operating a processing system, the method comprising: Use a main processor to decode and fetch an instruction; Output an interface instruction in response to the decoding of the fetched instruction; Receive the interface instruction from the main processor; Generate a command among the multiple commands according to the interface instruction; According to the interface instruction, determine the identified interface register among the multiple interface registers coupled to the multiple domain-specific accelerators, wherein each interface register includes: a command register having a plurality of command storage locations; an output queue coupled to the command register and the domain-specific accelerator among the multiple domain-specific accelerators; a response register having a plurality of response storage locations; and an input queue coupled to the response register and the domain-specific accelerator; and Output the command to the identified interface register, and the identified interface register executes the command output by the receiver.
Citation Information
Patent Citations
Data processing system and method for performing enhanced pipelined operations on instructions for normal and specific functions
US6532530B1