Instruction processing device, instruction processing method and board

By introducing a locking circuit into the processor to lock instructions with a longer processing time, the problem of ROB blockage in traditional processors when processing efficient artificial intelligence instructions is solved, which improves processing efficiency and saves ROB space.

CN113835759BActive Publication Date: 2025-05-02SHANGHAI CAMBRICON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010591063.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-24
Publication Date
2025-05-02
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

When traditional processors process efficient artificial intelligence instructions, due to the processing principle of reorder buffer (ROB), instructions with a longer processing time will cause ROB to be blocked, which will block the instruction finger fetching and decoding, affecting the processing efficiency.

Method used

The locking circuit is used to lock the instructions with a longer processing time and separate them from the instructions with a shorter processing time to avoid blocking the reorder buffer (ROB) and subsequent decoded new instructions.

Benefits of technology

By processing long and short instructions separately, the processor's processing efficiency is improved, the ROB and instruction finger fetching is blocked, and in some cases the ROB space is saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113835759B_ABST
    Figure CN113835759B_ABST
Patent Text Reader

Abstract

The present disclosure discloses an instruction processing device, an instruction processing method and a board. The instruction processing device may be included in a combined processing device, and the combined processing device may also include a universal interconnection interface and other processing devices. The instruction processing device interacts with other processing devices to jointly complete the calculation operation specified by the user. The combined processing device may also include a storage device, which is connected to the instruction processing device and the other processing devices respectively, and is used to store data of the instruction processing device and the other processing devices. The instruction processing scheme disclosed in the present disclosure can improve processing efficiency by processing instructions with longer processing time separately from instructions with shorter processing time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of processors, and in particular to an instruction processing device, an instruction processing method and a board. Background Art

[0002] In traditional processor design, the design method of out-of-order instruction processing is generally used, that is, instructions that have no dependencies on each other can be processed out of order. In order to achieve out-of-order processing, traditional processors generally use a reorder buffer (ROB) to record the true order of all instructions, and then release the instruction order after the instructions are executed.

[0003] For efficient artificial intelligence (AI) processors, since the instructions in the instruction set meet the needs of artificial intelligence algorithms more, the amount of data that an instruction may calculate, the functions implemented and / or the required computing time are much larger than the instructions of traditional processors. If the traditional ROB is still used to implement out-of-order instruction processing, problems may arise. Since the processing principle of ROB is that when the instructions stored in ROB are executed, the instructions are released sequentially. Therefore, the next instruction cannot be released when the previous instruction is not executed. When the processing time required for the current instruction is particularly long, the instruction cannot be released for a long time, resulting in the inability to release subsequent instructions, thereby blocking the ROB, and further blocking instruction fetching and decoding. Summary of the invention

[0004] In order to at least solve the technical problems mentioned above, the present disclosure proposes a solution in multiple aspects to use a locking circuit to lock instructions with a long processing time. Through the instruction processing solution disclosed in the present disclosure, instructions with a long processing time can be processed separately from instructions with a short processing time, so that instructions with a long processing time will not block the ROB, nor will they block new instructions decoded later, thereby improving processing efficiency.

[0005] In a first aspect, the present disclosure provides an instruction processing device, comprising: an instruction sending circuit, a reordering buffer circuit and a locking circuit, wherein the instruction sending circuit is configured to: send the instruction to the reordering buffer circuit and / or the locking circuit based on the type of instruction to be processed; and the locking circuit is configured to: receive instructions from the instruction sending circuit; and lock the general register associated with the received instruction.

[0006] In a second aspect, the present disclosure provides an instruction processing method, which includes: sending the instruction to a reorder buffer circuit and / or a locking circuit based on the type of the instruction to be processed; and at the locking circuit: receiving the instruction; and performing locking processing on a general register associated with the received instruction.

[0007] In a third aspect, the present disclosure provides a processor, comprising the instruction processing device of the first aspect.

[0008] In a fourth aspect, the present disclosure provides a board comprising the processor of the third aspect or the instruction processing device of the first aspect.

[0009] In a fifth aspect, the present disclosure provides a computing device, comprising the board of the fourth aspect.

[0010] Through the instruction processing device, instruction processing method, processor, board and computing device provided above, the solution disclosed herein utilizes a locking circuit to lock instructions with a long processing time, so that in artificial intelligence application scenarios including, for example, neural network operations, or other general scenarios, instructions with a long processing time can be processed separately from instructions with a short processing time, so that instructions with a long processing time will not block the ROB, nor will they block new instructions decoded later, thereby improving processing efficiency. Further, in some embodiments of the disclosure, for certain instructions, such as read-only general purpose register (GPR) instructions, the ROB may not be entered at all, thereby further saving ROB space. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] By reading the detailed description below with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0012] Figure 1 is a simplified block diagram showing a conventional instruction processing system including a ROB;

[0013] Figure 2 is a simplified block diagram showing an instruction processing apparatus according to an embodiment of the present disclosure;

[0014] Figure 3 is a flow chart showing an instruction processing method according to an embodiment of the present disclosure;

[0015] Figure 4 An example process of a method for sending instructions according to an embodiment of the present disclosure is shown;

[0016] Figure 5 An example process of a method for sending instructions according to another embodiment of the present disclosure is shown;

[0017] Figure 6 is a structural diagram showing a combined processing device according to an embodiment of the present disclosure; and

[0018] Figure 7 It is a schematic diagram showing the structure of a board card according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0020] It should be understood that the terms "first", "second", "third", and "fourth" in the claims, specifications, and drawings of the present disclosure are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprise" used in the specifications and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their collections.

[0021] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.

[0022] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0023] As mentioned in the background technology section, for efficient artificial intelligence (AI) processors, if the traditional ROB is still used to implement out-of-order instruction processing, problems may arise. This is because the processing principle of the traditional ROB is that when the instructions stored in the ROB are executed, the instructions are released in sequence. Therefore, the next instruction cannot be released when the previous instruction is not executed.

[0024] In order to better understand the technology involved in this disclosure, Figure 1 FIG. 1 shows a simplified block diagram of a conventional instruction processing system including a ROB. Figure 1 As shown, the instruction processing system 100 includes an instruction queue 110 , a re-order buffer (ROB) 120 , a register 130 , a reservation station 140 , an operation unit 150 , a memory access unit 160 , and a result bus 170 .

[0025] The instruction queue 110 receives decoded instructions from the instruction unit (not shown). In the instruction issuance stage, if the reservation station 140 is empty and there is an empty slot in the ROB 120, the instruction queue 110 will issue an instruction to the reservation station 140 and designate an item in the ROB 120 for temporarily storing the result of the instruction. Based on different operation types, there may be corresponding reservation stations, such as an addition reservation station and a multiplication reservation station. The instruction queue 110 may send instructions to the corresponding reservation station. For simplicity, in Figure 1 Not shown in detail.

[0026] During the instruction issuance process, the value of register 130 and the result status field are read. The result status field indicates whether the register is available. For example, empty indicates that the value of the register is available, otherwise it indicates the ROB number that generates the register result. If the result status field indicates that the register is available, the value of the register is read. If the ROB number that writes the register is recorded in the result status field, it indicates that the register has been renamed to ROB 120. At this time, ROB 120 is read according to the ROB number.

[0027] During the execution phase, each functional unit checks the instructions in the reservation station 140. If the operands required by an instruction in the reservation station 140 are ready, the instruction is executed. For example, the relevant operands are sent to the operation unit 150 to execute the instruction. If the operands of an instruction in the reservation station 140 are not ready, the instruction listens to the result bus 170 according to the result ROB number and receives the value of the result bus 170.

[0028] In the write-back phase, when the instruction is executed, for example, the operation unit 150 obtains the operation result, and when writing back, not only the result of the instruction operation but also the ROB number of the instruction should be written back. The operation unit 150 sends the relevant information (for example, the operation result, the ROB number) to the result bus and releases the reservation station 140. The ROB 120 and the register 130 listen to the result bus 170 and modify their own states according to the content of the result bus 170. For example, the register 130 modifies the result status field of the corresponding item according to the result bus 170, and the ROB 120 modifies the state and content of the corresponding item according to the result bus 170.

[0029] In the commit phase, if the result of the first instruction in the instruction queue has been written back and no exception has occurred, the result of the instruction is written back from ROB 120 to register 130 or memory, and the corresponding item of ROB is released. If an exception occurs in the first instruction in the queue, the operation queue and ROB are cleared.

[0030] References Figure 1 The following briefly describes how a traditional instruction processing system including ROB implements an instruction pipeline. As can be seen from the above process, in the write-back phase, the operation result of the instruction is first written to ROB 120, and in the commit phase, the content of ROB 120 is written back to the register or memory. The commit of instructions is in order, and can only be carried out after all previous instructions have been committed. This kind of instruction pipeline is out-of-order execution and orderly termination. The instruction execution is out-of-order, but the commit is in order.

[0031] From the above processing principle of ROB, it can be seen that if the processing time required for the previous instruction is particularly long, the instruction cannot be released for a long time, resulting in the failure to release subsequent instructions, thereby blocking ROB and further blocking instruction fetching and decoding.

[0032] In view of this, the present disclosure proposes a solution in multiple aspects to use a locking circuit to lock instructions with a long processing time. Through the instruction processing solution disclosed in the present disclosure, instructions with a long processing time can be processed separately from instructions with a short processing time, so that instructions with a long processing time will not block the ROB, nor will they block new instructions decoded later, thereby improving processing efficiency.

[0033] The specific implementation of the present disclosure is described in detail below with reference to the accompanying drawings.

[0034] Figure 2 A simplified block diagram of an instruction processing device 200 according to an embodiment of the present disclosure is shown. In one or more embodiments, the instruction processing device 200 can be used for any type of processor, including but not limited to a general-purpose processor (such as a CPU, etc.), an artificial intelligence processor (for example, as a coprocessor or accelerator) and / or a heterogeneous processor including the foregoing two, and the present disclosure is not limited in this respect.

[0035] like Figure 2 As shown, the instruction processing device 200 includes an instruction issuing circuit 210 , a reordering buffer circuit 220 and a locking circuit 230 .

[0036] In some embodiments, the instruction sending circuit 210 may be configured to send the instruction to the reorder buffer circuit 220 and / or the lock circuit 230 based on the type of the instruction to be processed. As mentioned above, since some instructions (e.g., instructions in an artificial intelligence processor, such as convolution operation instructions) may calculate a much larger amount of data and / or require a much longer calculation time than instructions of a traditional processor, if they are still processed in a traditional manner, the ROB will be blocked, further blocking instruction fetching and decoding. Therefore, in the embodiments of the present disclosure, different types of instructions to be processed, such as instructions with a longer processing time and instructions with a shorter processing time, may be processed separately, thereby avoiding blocking the ROB.

[0037] Specifically, in some embodiments, the instruction sending circuit 210 can be further configured to: determine the type of instruction to be processed; in response to determining that the instruction to be processed is of a first type, send the instruction of the first type to the reordering buffer circuit 220; and in response to determining that the instruction to be processed is of a second type, send the instruction of the second type to the locking circuit 230.

[0038] In some examples, the type of instruction can be predefined or identified and marked. For example, some rules can be pre-formulated, wherein according to the characteristics of the instruction, such as the type of operation to be performed, such as addition, multiplication, convolution and other operations, or the function to be implemented by the instruction or the amount of data calculated, these instructions are divided into a first type and a second type. These rules can be represented as an instruction classification table or other data structure containing the instruction classification results, for example. Then, in the instruction decoding stage, the instruction decoder identifies and marks the type of instruction based on the previously pre-configured instruction type table while decoding the instruction. Thus, the instruction sending circuit 210 can directly determine the type of instruction based on the tag of the instruction.

[0039] In other examples, the type of instruction may be determined in real time by the instruction sending circuit 210. For example, the instruction sending circuit 210 may determine whether the instruction is of the first type or the second type based on the aforementioned characteristics of the instruction or a pre-generated instruction type table when sending the instruction.

[0040] In some embodiments, the type of instruction is divided at least in part based on the execution time of the instruction. For example, the first type of instruction is an instruction whose execution time does not exceed a predetermined threshold, that is, an instruction with a shorter processing time, such as an instruction of a general-purpose processor, such as scalar addition, subtraction, multiplication, division and other operation instructions; and the second type of instruction is an instruction whose execution time exceeds a predetermined threshold, that is, an instruction with a longer processing time, such as an instruction of an artificial intelligence processor, such as convolution instructions, vector addition and multiplication instructions, pooling instructions, etc. It should be understood by those skilled in the art that there may also be instructions with a longer processing time in a general-purpose processor, such as image processing instructions, matrix operation instructions, etc., and there may also be instructions with a shorter processing time in an artificial intelligence processor. Therefore, the instruction sending circuit 210 in the disclosed embodiment may not need to distinguish the processor to which the instruction belongs, but directly determine the instruction type based on the pre-configured instruction division rules (for example, a pre-generated instruction type table or other data structure), and then process them separately.

[0041] In some embodiments, the reorder buffer circuit 220 can operate in its existing configuration. Figure 1 As described, the reservation station 140 turns the ordered instructions into out-of-order instructions, and the ROB 120 re-arranges the out-of-order instructions back to the ordered instructions, ensuring that the out-of-order instructions are submitted in order. Figure 2 The reorder buffer circuit (or ROB) 220 in the embodiment reorders the out-of-order executed instructions by sequentially inputting and sequentially releasing instructions, thereby achieving the purpose of sequential release. The reorder buffer circuit 220 can be used to store the results of instructions that have been executed but not yet submitted, thereby ensuring the orderly submission of instruction results.

[0042] Each entry (instruction) in ROB 220 may contain the following fields: instruction operation type, destination field (storage address or register number), value field, and ready field. The instruction operation type field may, for example, specify whether the instruction is a branch (without destination result), a storage instruction (with a storage address destination), or a register operation (arithmetic logic unit ALU operation or load instruction, which has a register destination). The destination field provides the register number (for load instructions and ALU operations) or the memory address (for storage instructions) to which the instruction result should be written. The value field is used to save the instruction result value before the instruction is submitted. The ready field indicates that the instruction has completed execution and the result value is ready.

[0043] The result register of an instruction can be renamed to the ROB number of its result, so that subsequent instructions may read operands from the ROB. Register renaming can eliminate pipeline conflicts that may be caused by dependencies between instructions (for example, data dependencies such as write-after-write (WAW) and read-after-write (WAR)).

[0044] After the instruction is executed, the result is not written directly back to the register, but written to the ROB. As long as an instruction is not committed, it will not modify the contents of the register or memory, and it is easy to cancel the instruction.

[0045] Therefore, in some embodiments, the reorder buffer circuit 220 may be configured to receive instructions from the instruction issuing circuit 210 in order; and release the instructions in the same order based on the execution status of the instructions.

[0046] In some examples, the reorder buffer circuit 220 can be implemented using a first-in-first-out (FIFO) queue. When an instruction is sent to the reservation station, the instruction needs to be inserted into the tail of the ROB queue, and the executed instruction will be removed from the head of the ROB queue. Only when the execution status of the first instruction in the queue is ready, that is, the ready field of the item in the ROB indicates that the result value is ready, the instruction result of the instruction is submitted and the corresponding item in the ROB is released.

[0047] Therefore, for the first type of instructions received from the instruction issuing circuit 210 , such as instructions with a shorter processing time, the reorder buffer circuit 220 may process them according to the existing configuration.

[0048] In some embodiments, the locking circuit 230 may be configured to: receive instructions from the instruction sending circuit 210; and lock the general register associated with the received instruction. As described above, the instruction sending circuit 210 may send a second type of instruction to the locking circuit 230. The second type of instruction is, for example, an instruction that takes a long time to process, such as an instruction involving a neural network operation. By using the locking circuit 230 to process the second type of instruction separately, it is possible to avoid the second type of instruction from blocking the reorder buffer circuit 220 due to a long execution time.

[0049] The locking circuit 230 locks the general purpose registers (GPRs) used in the second type of instructions so that the GPRs used by the instructions will not be released in advance and the reorder buffer circuit 220 will not be blocked.

[0050] Specifically, the locking circuit 230 will lock the GPR used by the received instruction. That is, after locking, the locked GPR cannot be released and used until the instruction is executed. Figure 1 After the computing unit 150 in the instruction sends a signal that the instruction has been executed back to the locking circuit 230, the locking circuit 230 will release the lock of the corresponding GPR, and then subsequent instructions using the same GPR can continue to be calculated.

[0051] Therefore, in some embodiments, the locking circuit 230 may be further configured to: in response to receiving a signal indicating that the execution of the instruction is complete, unlock the GPR associated with the instruction.

[0052] In some embodiments, the locking and unlocking process of the locking circuit 230 can be implemented by setting the status bit of the GPR. For example, when the status bit of the GRP is set to "1", it can indicate that the GPR is locked and cannot be used at this time; on the contrary, when the status of the GPR is set to "0", it can indicate that the GPR is not locked, that is, the GPR is released and can be used again.

[0053] By processing the second type of instructions separately and using the locking circuit 230 to lock the associated GPR, even if the second type of instructions (for example, neural network instructions) take a particularly long time to execute, it will not block new instructions decoded later. As long as the GPR used by the new instruction is different from the locked GPR, the instruction will not be blocked by the locking circuit 230 and can continue to be decoded and executed.

[0054] In the above embodiment, since the instruction status of the second type of instruction is not stored in the reorder buffer circuit 220, the reorder buffer circuit 220 is not required to release the second type of instruction. In this way, not only the space of the ROB is released, but also the subsequent instructions will not be unable to be released because the previous locked second type of instruction has not been executed, so the instruction fetching and decoding can be accelerated.

[0055] References from previous article Figure 1 As can be seen from the description, ROB can also implement the register renaming function, thereby eliminating pipeline conflicts that may be caused by dependencies between instructions (for example, data dependencies such as write-after-write (WAW) and read-after-write (WAR)).

[0056] Furthermore, in some embodiments, the instruction sending circuit 210 can be further configured to: in response to determining that the instruction to be processed is of the second type, also send this second type of instruction to the reorder buffer circuit 220; and indicate to the reorder buffer circuit 220 that the execution status of this second type of instruction is completed.

[0057] As described above, the reorder buffer circuit 220 can store information such as the instruction operation type, the destination field (storage address or register number), the value field and the ready field for each item (each instruction). The ready field indicates that the instruction has been executed and the result value is ready. Since the GPR used by the second type of instruction has been locked by the locking circuit 230 and can only be released by the locking circuit 230 in response to receiving a signal indicating that the instruction has been executed, there is no need to preserve the order of the second type of instruction in the reorder buffer circuit 220. In view of this, the instruction sending circuit 210 can directly indicate to the reorder buffer circuit 220 that the execution status of the second type of instruction is completed. At this time, the reorder buffer circuit 220 can set the ready field of the item (instruction) to ready. In this way, the reorder buffer circuit 220 does not need to wait for the instruction to be executed, but assumes that the instruction has been executed, and only needs to be released according to the normal operation sequence. In this way, even if the second type of instruction enters the reorder buffer circuit 220, since it is assumed to be executed, it will not block the reorder buffer circuit 220 even if the execution time is long.

[0058] As mentioned above, the register renaming function is mainly to eliminate pipeline conflicts that may be caused by dependencies between instructions (for example, data dependencies such as write-after-write (WAW) and read-after-write (WAR)). In view of this, for the second type of instruction, the instruction can be sent to the reorder buffer circuit 220 only when the instruction involves a write operation to the GPR.

[0059] In these embodiments, the instruction sending circuit 210 can be further configured to: determine whether the second type of instruction involves a write operation to a general register; and in response to determining that the second type of instruction involves a write operation to a general register, also send the second type of instruction to the reorder buffer circuit 220, and indicate to the reorder buffer circuit 220 that the execution status of the second type of instruction is completed.

[0060] Optionally, in some other embodiments, the instruction sending circuit 210 is further configured to: in response to determining that the second type of instruction does not involve a write operation to a general register, not send the second type of instruction to the reorder buffer circuit 220. Since register renaming is not required when the second type of instruction does not involve a write operation to a GPR, the instruction does not need to be sent to the reorder buffer circuit 220, thereby saving space in the reorder buffer circuit 220.

[0061] Reference above Figure 2 The instruction processing device 200 of the embodiment of the present disclosure is described. Those skilled in the art can understand that the instruction processing device 200 is for example Figure 1Specifically, in some implementations, the instruction sending circuit 210 in the instruction processing device 200 may be implemented in Figure 1 The instruction queue 110 shown in the figure can be in or outside the queue, so as to select different order-preserving circuits according to the type of instruction, such as the reordering buffer circuit 220 and / or the locking circuit 230. The reordering buffer circuit 220 can be generally similar to Figure 1 The configuration of the ROB 120 is such that only when receiving the second type of instruction, it is assumed that the instruction has been executed. The locking circuit 230 can be a buffer similar to the ROB. The locking circuit 230 is used to lock the GPR used by the second type of instruction.

[0062] Through the instruction processing device provided above, the solution disclosed herein utilizes a locking circuit to lock instructions with a long processing time, so that in artificial intelligence application scenarios including, for example, neural network operations, or other general scenarios, instructions with a long processing time can be processed separately from instructions with a short processing time, so that instructions with a long processing time will not block the ROB, nor will they block new instructions decoded later, thereby improving processing efficiency. Further, in some embodiments of the disclosure, for certain instructions, such as read-only GPR instructions, the ROB may not be entered at all, thereby further saving ROB space.

[0063] Figure 3 A flowchart of an instruction processing method 300 according to an embodiment of the present disclosure is shown. As mentioned above, in one or more embodiments, the instruction processing method 300 can be applied to any type of processor, including but not limited to a general-purpose processor (such as a CPU, etc.), an artificial intelligence processor (for example, as a coprocessor or accelerator) and / or a heterogeneous processor including the foregoing two, and the present disclosure is not limited in this respect.

[0064] like Figure 3 As shown, in step S310, based on the type of the instruction to be processed, the instruction is sent to the reorder buffer circuit and / or the locking circuit. Step S310 may be, for example, Figure 2 The instruction is sent to the circuit 210 for execution.

[0065] As described above with reference to the instruction processing apparatus 200 , in the embodiments of the present disclosure, different types of instructions to be processed, such as instructions with longer processing time and instructions with shorter processing time, may be processed separately, thereby avoiding clogging of the ROB.

[0066] Figure 4 FIG. 4 shows an example process of a method 400 for sending instructions according to an embodiment of the present disclosure. Figure 4As shown, at step S411, it is determined whether the instruction to be processed is of the first type or the second type. At step S412, in response to determining that the instruction to be processed is of the first type, the instruction of the first type is sent to the reordering buffer circuit. And at step S413, in response to determining that the instruction to be processed is of the second type, the instruction of the second type is sent to the locking circuit. In some embodiments, the type of instruction is divided at least in part based on the execution time of the instruction. For example, the first type of instruction is an instruction whose execution time does not exceed a predetermined threshold; and the second type of instruction is an instruction whose execution time exceeds a predetermined threshold. The determination of the instruction type can refer to the previous description and will not be repeated here.

[0067] As mentioned above, considering utilizing the register renaming function of the reorder buffer circuit to eliminate pipeline conflicts that may be caused by dependencies between instructions, in some embodiments, Figure 4 The method may further include step S414, in response to determining that the instruction to be processed is of the second type, sending the instruction of the second type to the reorder buffer circuit; and indicating to the reorder buffer circuit that the execution status of the instruction of the second type is completed. Thus, by defaulting the instruction of the second type as completed in the reorder buffer circuit, the register renaming function of the reorder buffer circuit can be utilized without blocking the reorder buffer circuit.

[0068] Further, in some embodiments, whether to send the second type of instruction to the reorder buffer circuit is determined according to whether the second type of instruction involves a write operation to a GPR.

[0069] Figure 5 An example flow of an instruction sending method 500 according to another embodiment of the present disclosure is shown.

[0070] like Figure 5 As shown, at step S511, it is determined whether the type of instruction to be processed is the first type or the second type. If the instruction is the first type, the method proceeds to step S512, and the first type of instruction is sent to the reorder buffer circuit. If the instruction is the second type, the method proceeds to step S513.

[0071] At step S513, it is further determined whether the second type of instruction involves a write operation to the GPR. If the instruction involves a write operation to the GPR, the method proceeds to step S514, the second type of instruction is sent to both the locking circuit and the reordering buffer circuit, and the execution status of the second type of instruction is indicated to the reordering buffer circuit as completed. If the instruction does not involve a write operation to the GPR, for example, only involves a read operation to the GPR, the method proceeds to step S515, at which time the second type of instruction only needs to be sent to the locking circuit, and does not need to be sent to the reordering buffer circuit.

[0072] Back to Figure 3 If the instruction is sent to the locking circuit, then, in step S320, the locking circuit: receives the instruction; and performs locking processing on the general register associated with the received instruction.

[0073] Specifically, the locking circuit will lock the GPR used by the received instruction. That is, after locking, the locked GPR cannot be released and used until the instruction is executed. Figure 1 After the computing unit 150 in the processor sends a signal that the instruction has been executed back to the locking circuit, the locking circuit will release the lock of the corresponding GPR, and then subsequent instructions using the same GPR can continue to be calculated.

[0074] Therefore, in some embodiments, the locking circuit may further include the step of: in response to receiving a signal indicating that the execution of the instruction is complete, unlocking the GPR associated with the instruction.

[0075] In some embodiments, the reorder buffer circuit may further include the steps of: receiving instructions in sequence; and releasing the instructions in sequence based on the execution status of the instructions.

[0076] The instruction processing method executed by the instruction processing device of the embodiment of the present disclosure has been described above with reference to the flowchart. Those skilled in the art can understand that, since only the order preservation processing is required for the type of instruction, it is possible to prevent instructions with long processing time from blocking the ROB, and it will not block new instructions decoded later, thereby improving processing efficiency. Further, in some embodiments of the present disclosure, for certain instructions, such as read-only GPR instructions, it is possible not to enter the ROB at all, thereby further saving the space of the ROB.

[0077] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0078] It should be further explained that although Figure 3-Figure 5 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3-5At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0079] Figure 6 is a structural diagram showing a combined processing device 600 according to an embodiment of the present disclosure. Figure 6 As shown, the combined processing device 600 includes the aforementioned instruction processing device 602, which can be configured to execute the instruction processing method described in conjunction with the above drawings. In addition, the combined processing device also includes a universal interconnection interface 604 and other processing devices 606. According to the present disclosure, the instruction processing device 602 can interact with other processing devices 606 through the universal interconnection interface 604 to jointly complete the operation specified by the user.

[0080] According to the scheme disclosed herein, the other processing device may include one or more types of processors in general and / or special processors such as a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence processor, etc., and the number thereof may not be limited but determined according to actual needs. In one or more embodiments, the other processing device may serve as an interface between the instruction processing device disclosed herein (which may be embodied as an artificial intelligence such as a related computing device for neural network computing) and external data and control, and perform basic control including but not limited to data handling and starting and stopping the instruction processing device; other processing devices may also cooperate with the instruction processing device to jointly complete computing tasks.

[0081] According to the scheme disclosed herein, the universal interconnect interface can be used to transmit data and control instructions between the instruction processing device and other processing devices. For example, the instruction processing device can obtain the data to be calculated from other processing devices via the universal interconnect interface and write it into the storage device (or storage circuit or memory) on the computing device chip. Further, the instruction processing device can obtain control instructions from other processing devices via the universal interconnect interface and write them into the control cache on the instruction processing device chip. Alternatively or optionally, the universal interconnect interface can also read the data in the storage circuit of the instruction processing device and transmit it to other processing devices.

[0082] Optionally, the combined processing device may further include a storage device 608, which may be connected to the instruction processing device and other processing devices, respectively. In one or more embodiments, the storage device may be used to store data of the instruction processing device and other processing devices, especially data that cannot be fully stored in the internal or on-chip storage device of the instruction processing device or other processing devices.

[0083] According to different application scenarios, the combined processing device disclosed in the present invention can be used as a SOC chip system for mobile phones, robots, drones, video surveillance equipment and other devices, effectively reducing the core area of ​​the control part, improving the processing speed, and reducing the overall power consumption. In this case, the universal interconnection interface of the combined processing device is connected to certain components of the device. Certain components such as cameras, displays, mice, keyboards, network cards or wifi interfaces.

[0084] In some embodiments, the present disclosure further discloses a processor, which includes the above instruction processing device or combined processing device. In other embodiments, the present disclosure further discloses a chip packaging structure, which includes the above processor.

[0085] In some embodiments, the present disclosure also discloses an integrated circuit board, which includes the above chip packaging structure. Figure 7 , which provides the aforementioned exemplary board card, which, in addition to the aforementioned chip 702 , may also include other supporting components, including but not limited to: a storage device 704 , an interface device 706 and a control device 708 .

[0086] The memory device is connected to the chip in the chip package structure via a bus for storing data. The memory device may include multiple groups of memory cells 710. Each group of memory cells is connected to the chip via a bus. It is understood that each group of memory cells may be DDR SDRAM ("Double Data Rate SDRAM, double rate synchronous dynamic random access memory").

[0087] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read out on the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of the above-mentioned storage units. Each group of storage units may include multiple DDR4 particles (chips). In one embodiment, the chip may include 4 72-bit DDR4 controllers, of which 64 bits are used for data transmission and 8 bits are used for ECC verification.

[0088] In one embodiment, each group of storage units includes a plurality of double rate synchronous dynamic random access memories arranged in parallel. DDR can transmit data twice in one clock cycle. A controller for controlling DDR is arranged in the chip to control the data transmission and data storage of each storage unit.

[0089] The interface device in the figure is electrically connected to the chip in the chip packaging structure. The interface device is used to realize data transmission between the chip and an external device 712 (such as a server or a computer). For example, in one embodiment, the interface device can be a standard PCIE interface. For example, the data to be processed is transmitted to the chip by the server through the standard PCIE interface to realize data transfer. In another embodiment, the interface device can also be other interfaces. This disclosure does not limit the specific manifestations of the above-mentioned other interfaces. The interface device can realize the switching function. In addition, the calculation results of the chip are still transmitted back to the external device (such as a server) by the interface device.

[0090] The control device in the figure is electrically connected to the chip. The control device is used to monitor the state of the chip. Specifically, the chip and the control device can be electrically connected through an SPI interface. The control device may include a single-chip microcomputer (Micro Controller Unit, MCU). In one or more embodiments, the chip may include multiple processing chips, multiple processing cores or multiple processing circuits, which can drive multiple loads. Therefore, the chip can be in different working states such as multi-load and light load. The control device can realize the regulation of the working state of multiple processing chips, multiple processing and / or multiple processing circuits in the chip.

[0091] In some embodiments, the present disclosure also discloses a computing device or equipment, which includes the above-mentioned integrated circuit board. According to different application scenarios, the computing device or equipment may include a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage, a wearable device, a means of transportation, a household appliance, and / or a medical device. Among them, the means of transportation may include an airplane, a ship and / or a vehicle; the household appliance may include a television, an air conditioner, a microwave oven, a refrigerator, an electric rice cooker, a humidifier, a washing machine, an electric lamp, a gas stove, and a range hood; the medical device may include an MRI, an ultrasound machine and / or an electrocardiograph.

[0092] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0093] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of each unit, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, optical, acoustic, magnetic or other forms.

[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software program module.

[0096] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, when the technical solution of the present disclosure can be embodied in the form of a software product (such as a computer-readable storage medium), the computer software product is stored in a memory, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0097] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0098] The foregoing content can be better understood in accordance with the following terms:

[0099] Clause 1. An instruction processing device, comprising: an instruction issuing circuit, a reordering buffer circuit and a locking circuit, wherein:

[0100] The instruction sending circuit is configured to: send the instruction to the reorder buffer circuit and / or the locking circuit based on the type of the instruction to be processed; and

[0101] The locking circuit is configured as follows:

[0102] receiving an instruction from the instruction sending circuit; and

[0103] The general register associated with the received instruction is locked.

[0104] Clause 2. The instruction processing apparatus according to clause 1, wherein the instruction issuing circuit is further configured to:

[0105] Determining the type of the instruction to be processed;

[0106] In response to determining that the instruction to be processed is of a first type, sending the instruction of the first type to the reorder buffer circuit; and

[0107] In response to determining that the pending instruction is of a second type, the instruction of the second type is sent to the locking circuit.

[0108] Clause 3. The instruction processing apparatus according to clause 2, wherein the instruction issuing circuit is further configured to:

[0109] determining whether the second type of instruction involves a write operation to a general register; and

[0110] In response to determining that the second type of instruction involves a write operation to a general register, the second type of instruction is further sent to the reorder buffer circuit, and the execution status of the second type of instruction is indicated to the reorder buffer circuit as execution completion.

[0111] Clause 4. The instruction processing apparatus according to clause 3, wherein the instruction issuing circuit is further configured to:

[0112] In response to determining that the second type of instruction does not involve a write operation to a general register, the second type of instruction is not sent to the reorder buffer circuit.

[0113] Clause 5. The instruction processing apparatus according to clause 2, wherein the instruction issuing circuit is further configured to:

[0114] In response to determining that the instruction to be processed is of a second type, further sending the instruction of the second type to the reorder buffer circuit; and

[0115] Indicating to the reorder buffer circuit that the execution status of the second type of instruction is completed.

[0116] Clause 6. An instruction processing device according to any one of Clauses 2-5, wherein the first type of instruction is an instruction whose execution time does not exceed a predetermined threshold, and the second type of instruction is an instruction whose execution time exceeds a predetermined threshold.

[0117] Clause 7. An instruction processing apparatus according to clause 6, wherein the second type of instructions comprises neural network instructions.

[0118] Clause 8. An instruction processing apparatus according to any one of clauses 1 to 7, wherein the locking circuit is further configured to:

[0119] In response to receiving a signal indicating that the execution of the instruction is complete, a general register associated with the instruction is unlocked.

[0120] Clause 9. An instruction processing apparatus according to any one of clauses 1 to 8, wherein the reorder buffer circuit is configured to:

[0121] receiving instructions in sequence from the instruction sending circuit; and

[0122] The instructions are released in the order based on the execution status of the instructions.

[0123] Clause 10. A method of processing an instruction, comprising:

[0124] Based on the type of instruction to be processed, sending the instruction to a reorder buffer circuit and / or a lock circuit; and

[0125] At the locking circuit:

[0126] receiving the instruction; and

[0127] The general register associated with the received instruction is locked.

[0128] Clause 11. The instruction processing method according to clause 10, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the lock circuit further comprises:

[0129] Determining the type of the instruction to be processed;

[0130] In response to determining that the to-be-processed instruction is of a first type, sending the instruction of the first type to the reorder buffer circuit; and

[0131] In response to determining that the pending instruction is of a second type, the instruction of the second type is sent to the locking circuit.

[0132] Clause 12. The instruction processing method according to clause 11, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the lock circuit further comprises:

[0133] determining whether the second type of instruction involves a write operation to a general register; and

[0134] In response to determining that the second type of instruction involves a write operation to a general register, the second type of instruction is further sent to the reorder buffer circuit, and the execution status of the second type of instruction is indicated to the reorder buffer circuit as execution completion.

[0135] Clause 13. The instruction processing method according to clause 12, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the lock circuit further comprises:

[0136] In response to determining that the second type of instruction does not involve a write operation to a general register, the second type of instruction is not sent to the reorder buffer circuit.

[0137] Clause 14. The instruction processing method according to clause 11, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the lock circuit further comprises:

[0138] In response to determining that the instruction to be processed is of a second type, further sending the instruction of the second type to the reorder buffer circuit; and

[0139] Indicating to the reorder buffer circuit that the execution status of the second type of instruction is completed.

[0140] Clause 15. An instruction processing method according to any one of Clauses 11-14, wherein the first type of instruction is an instruction whose execution time does not exceed a predetermined threshold, and the second type of instruction is an instruction whose execution time exceeds a predetermined threshold.

[0141] Clause 16. The instruction processing method of clause 15, wherein the second type of instructions comprises neural network instructions.

[0142] Clause 17. The instruction processing method according to any one of clauses 10 to 16 further comprises:

[0143] At the locking circuit:

[0144] In response to receiving a signal indicating that the execution of the instruction is complete, a general register associated with the instruction is unlocked.

[0145] Clause 18. The instruction processing method according to any one of clauses 10 to 17 further comprises:

[0146] At the reorder buffer circuit:

[0147] receiving the instructions in order; and

[0148] The instructions are released in the order based on the execution status of the instructions.

[0149] Clause 19. A processor comprising an instruction processing device according to any one of clauses 1-9.

[0150] Clause 20. A board comprising the processor according to Clause 19 or the instruction processing device according to any one of Clauses 1-9.

[0151] Clause 21. A computing device comprising the board according to clause 20.

[0152] The embodiments of the present disclosure are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of ​​the present disclosure. At the same time, changes or deformations made by those skilled in the art based on the ideas of the present disclosure, the specific implementation methods and the scope of application of the present disclosure, all belong to the scope of protection of the present disclosure. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. An instruction processing device, comprising: Instruction sending circuit, reordering buffer circuit and locking circuit, where: The instruction sending circuit is configured to: send the instruction to the reorder buffer circuit and / or the locking circuit based on the type of the instruction to be processed; and The locking circuit is configured as follows: receiving an instruction from the instruction sending circuit; and Locking the general register associated with the received instruction; Wherein, the instruction sending circuit is further configured to: Determining the type of the instruction to be processed; In response to determining that the instruction to be processed is of a first type, sending the instruction of the first type to the reorder buffer circuit; and In response to determining that the pending instruction is of a second type, the instruction of the second type is sent to the locking circuit.

2. The instruction processing device according to claim 1, wherein the instruction issuing circuit is further configured to: determining whether the second type of instruction involves a write operation to a general register; and In response to determining that the second type of instruction involves a write operation to a general register, the second type of instruction is further sent to the reorder buffer circuit, and the execution status of the second type of instruction is indicated to the reorder buffer circuit as execution completion.

3. The instruction processing device according to claim 2, wherein the instruction issuing circuit is further configured to: In response to determining that the second type of instruction does not involve a write operation to a general register, the second type of instruction is not sent to the reorder buffer circuit.

4. The instruction processing device according to claim 1, wherein the instruction issuing circuit is further configured to: In response to determining that the instruction to be processed is of a second type, further sending the instruction of the second type to the reorder buffer circuit; and Indicating to the reorder buffer circuit that the execution status of the second type of instruction is completed.

5. The instruction processing device according to any one of claims 1 to 4, wherein the first type of instructions are instructions whose execution time does not exceed a predetermined threshold, and the second type of instructions are instructions whose execution time exceeds a predetermined threshold.

6. The instruction processing device according to any one of claims 1 to 4, wherein the locking circuit is further configured to: In response to receiving a signal indicating that the execution of the instruction is complete, a general register associated with the instruction is unlocked.

7. The instruction processing device according to any one of claims 1 to 4, wherein the reorder buffer circuit is configured to: receiving instructions in sequence from the instruction sending circuit; and The instructions are released in the order based on the execution status of the instructions.

8. A method for processing an instruction, comprising: Based on the type of the instruction to be processed, sending the instruction to a reorder buffer circuit and / or a locking circuit; as well as At the locking circuit: receiving the instruction; and Locking the general register associated with the received instruction; Wherein, based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the lock circuit further includes: Determining the type of the instruction to be processed; In response to determining that the instruction to be processed is of a first type, sending the instruction of the first type to the reorder buffer circuit; and In response to determining that the pending instruction is of a second type, the instruction of the second type is sent to the locking circuit.

9. The instruction processing method according to claim 8, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the locking circuit further comprises: determining whether the second type of instruction involves a write operation to a general register; as well as In response to determining that the second type of instruction involves a write operation to a general register, the second type of instruction is further sent to the reorder buffer circuit, and the execution status of the second type of instruction is indicated to the reorder buffer circuit as execution completion.

10. The instruction processing method according to claim 9, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the locking circuit further comprises: In response to determining that the second type of instruction does not involve a write operation to a general register, the second type of instruction is not sent to the reorder buffer circuit.

11. The instruction processing method according to claim 8, wherein based on the type of the instruction to be processed, sending the instruction to the reorder buffer circuit and / or the locking circuit further comprises: In response to determining that the to-be-processed instruction is of a second type, further sending the second type of instruction to the reorder buffer circuit; as well as Indicating to the reorder buffer circuit that the execution status of the second type of instruction is completed.

12. The instruction processing method according to any one of claims 8 to 11, wherein the first type of instructions are instructions whose execution time does not exceed a predetermined threshold, and the second type of instructions are instructions whose execution time exceeds a predetermined threshold.

13. The instruction processing method according to any one of claims 8 to 11, further comprising: At the locking circuit: In response to receiving a signal indicating that the execution of the instruction is complete, a general register associated with the instruction is unlocked.

14. The instruction processing method according to any one of claims 8 to 11, further comprising: At the reorder buffer circuit: receiving said instructions in order; as well as The instructions are released in the order based on the execution status of the instructions.

15. A board comprising the instruction processing device according to any one of claims 1-7.

Citation Information

Patent Citations

  • Prioritising instructions according to category of instruction

    CN104346223A

  • System and method for parking re-issued instruction of microprocessor

    CN104657145A