Processors, graphics cards, computer equipment, and methods to eliminate dependencies
By setting up multiple paths between the instruction processing unit and the dependency processing unit, the dependency relationship of different types of registers used by the same instruction is solved, and more efficient parallel execution and pipeline processing are achieved.
Patent Information
- Application Number
- CN202410977461.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-19
AI Technical Summary
In the prior art, due to the different usage timing and duration of different registers by the same instruction, the usage efficiency of registers is reduced, and the dependencies between different instructions cannot be effectively terminated, which affects the efficiency of parallel execution.
Multiple paths are set up between the instruction processing unit and the dependency processing unit, and the dependency relationship of different types of registers used by the same instruction is released through different types of channels. The scoreboard mechanism and the sleep wake-up mechanism are used to cancel the first-used register dependency relationship in advance.
It improves the efficiency of register usage, increases the possibility of parallel execution between different instructions, reduces register usage time, and optimizes the processing efficiency of multi-stage pipelines.
Smart Images

Figure CN119003002B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of chips, and in particular to a processor, a graphics card, a computer device, and a dependency removal method. Background Art
[0002] When executing instructions, the processor often uses a multi-stage pipeline mechanism. Taking a relatively simple five-stage pipeline as an example, the five-stage pipeline includes: instruction fetch stage, decoding stage, execution stage, memory access stage, and write-back stage. The five-stage pipeline can increase the efficiency when multiple instructions are executed in parallel.
[0003] The existence of a dependency between two instructions means that the two instructions are assigned to use the same register, and the latter instruction must wait until the former instruction has finished using the register before it can start using the register. Dependencies are divided into three types: read first, write later, write first, read later, and write first, write later. For example, if the first instruction and the second instruction have a read-first, write-later dependency, the processor needs to wait until the write-back phase corresponding to the first instruction is completed before the second instruction can release the dependency on the first instruction in order to execute the second instruction's write to the register.
[0004] Since the second instruction needs to wait until the write-back phase corresponding to the first instruction is executed before the second instruction can release its dependency on the first instruction, the first instruction occupies the register for a long time, resulting in reduced register utilization efficiency. Summary of the invention
[0005] The embodiment of the present application provides a processor, a graphics card, a computer device, and a method for removing dependencies. The technical solution is as follows:
[0006] In one aspect, an embodiment of the present application provides a processor, the processor comprising: an instruction processing unit and a dependency processing unit, at least two paths provided between the instruction processing unit and the dependency processing unit, the at least two paths comprising a first type path and a second type path;
[0007] The instruction processing unit is used to send a first dependency release signal to the dependency processing unit through the first type path; the dependency processing unit is used to release the first dependency relationship corresponding to the first type register based on the first dependency release signal;
[0008] The instruction processing unit is used to send a second dependency release signal to the dependency processing unit through the second type path; the dependency processing unit is used to release the second dependency relationship corresponding to the second type register based on the second dependency release signal;
[0009] The first type register and the second type register are both registers used by the same instruction.
[0010] On the other hand, an embodiment of the present application provides a graphics card, which includes the processor as described above. Optionally, the processor is a graphics processing unit.
[0011] On the other hand, an embodiment of the present application provides a computer device, which includes the processor as described above. Optionally, the processor is a graphics processing unit.
[0012] On the other hand, an embodiment of the present application provides a method for dependency resolution. The method is executed by a processor, which includes an instruction processing unit and a dependency processing unit. At least two paths are provided between the instruction processing unit and the dependency processing unit, and the at least two paths include a first type of path and a second type of path. The method includes:
[0013] The instruction processing unit sends a first dependency resolution signal to the dependency processing unit through the first type of path; the dependency processing unit resolves a first dependency relationship corresponding to a first type of register based on the first dependency resolution signal;
[0014] The instruction processing unit sends a second dependency resolution signal to the dependency processing unit through the second type of path; the dependency processing unit resolves a second dependency relationship corresponding to a second type of register based on the second dependency resolution signal;
[0015] Wherein, both the first type of register and the second type of register are registers used by the same instruction.
[0016] In the embodiment of the present application, by providing a first type of path and a second type of path between the instruction processing unit and the dependency processing unit. For two types of registers used by the same instruction, the first dependency relationship is resolved using the first type of path and the second dependency relationship is resolved using the second type of path respectively. Since the usage timing and usage duration of different types of registers by the same instruction are different, the dependency relationship corresponding to the register that is used up first can be resolved earlier, thereby reducing the occupation duration of the register that is used up first by this instruction and improving the usage efficiency of this part of the registers. At the same time, since this part of the registers can be released earlier, the possibility of parallel instructions between different instructions is increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1It is a schematic structural diagram of a processor provided by an exemplary embodiment of the present application;
[0019] Figure 2 It is a schematic diagram of a five-stage pipeline provided by an exemplary embodiment of the present application;
[0020] Figure 3 It is a flowchart of a multi-stage pipeline provided by an exemplary embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of a device related to a load instruction provided by an exemplary embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of a device related to a store instruction provided by an exemplary embodiment of the present application;
[0023] Figure 6 It is a schematic structural diagram of a processor provided by an exemplary embodiment of the present application;
[0024] Figure 7 It is a schematic structural diagram of three candidate paths provided by an exemplary embodiment of the present application;
[0025] Figure 8 It is a schematic diagram of the usage timing of three candidate paths in a multi-stage pipeline provided by an exemplary embodiment of the present application;
[0026] Figure 9 It is a schematic diagram of the dependence release of a load instruction provided by an exemplary embodiment of the present application;
[0027] Figure 10 It is a schematic diagram of the dependence release of a store instruction provided by an exemplary embodiment of the present application;
[0028] Figure 11 It is a schematic diagram of the dependence release of a load instruction provided by an exemplary embodiment of the present application;
[0029] Figure 12 It is a flowchart of a dependence release method provided by an exemplary embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0031] First, the terms involved in the embodiments of the present application are introduced.
[0032] Instruction: It is a command that instructs a computer to perform a certain operation and is the smallest functional unit for a computer to run. An instruction is a statement in machine language and is a group of meaningful binary codes. The set of all instructions of a computer constitutes the instruction system of the computer, also known as the instruction set.
[0033] Instruction format: The readable representation of an instruction. An instruction usually includes an opcode (Operation Code) and an operand. The opcode is used to describe the type of operation to be performed by the instruction, and the operand provides the data or the address of the data required when executing the instruction. In an instruction, the opcode is indispensable, but the operand can be absent, or there can be one operand or two operands.
[0034] Operation type: At least one of the data transfer type, arithmetic logic type, shift type, transfer type, input / output type. For example, the data transfer type is used for data transfer between memory and registers; another example is the arithmetic logic operation type, which is used to perform arithmetic operations or logic operations. Arithmetic operations include addition, subtraction, multiplication, division, etc., and logic operations include AND, NOT, OR, etc.
[0035] Instruction type: Classified according to the number of operands, instruction types include zero-operand instructions, one-operand instructions, two-operand instructions, etc. A zero-operand instruction has no operand or the operand is implicitly determined in the instruction format; a one-operand instruction has one operand in the instruction format, or there is also an implicit operand (actually a two-operand instruction); a two-operand instruction has two operands in the instruction format, one of which is the destination operand and the other is the source operand. The source operand is the operand that can only be read, and the destination operand is the operand that can be both read and written (store the operation result).
[0036] Data type operand: It is the data that needs to be loaded, stored or calculated.
[0037] Control type operand: It is an operand different from the data type operand. The control type operand is used to indicate a memory address or an execution condition. For example, the control type operand includes at least one of the following: memory address, offset value of the memory address, predicate value in the predicate register, bypass information.
[0038] Predicate register: A register used to implement conditional execution. For example, 1 represents true, and the current instruction or a certain branch of the current instruction needs to be executed; 0 represents false, and the current instruction or a certain branch of the current instruction does not need to be executed.
[0039] Bypassing Information: It is a technique used to improve performance in the pipeline design of the central processing unit (CPU) and the graphics processing unit (GPU). In pipeline processing, the operation of each stage depends on the output of the previous stage. However, if the result of the operation of a stage needs to be used immediately in the next clock cycle and cannot wait for the completion of the entire pipeline, bypass technology is needed to directly pass data from one stage to the next stage that needs it.
[0040] Load instruction: used to read data from memory into registers. When the processor is a GPU, the memory is video memory.
[0041] Store instruction: used to write data in register to memory.
[0042] Pipeline hazard: In some cases, the next instruction in the next clock cycle of the pipeline cannot be executed. This situation is called a hazard, which is divided into three types: data hazard, structural hazard and control hazard. The embodiment of the present application involves data hazard.
[0043] Data hazard: When at least two instructions are executed in parallel in the pipeline, if the second instruction needs to use the execution result of the first instruction, and the first instruction has not yet written back, a conflict will occur. Data hazard is also called data dependency.
[0044] Data hazards are divided into three categories: Read After Write (RAW), Write After Write (WAW), and Write After Read (WAR).
[0045] RAW means that the second instruction reads the register, which needs to be later than the first instruction writes to the register; WAW means that the second instruction writes to the register, which needs to be later than the first instruction writes to the register; WAR means that the second instruction writes to the register, which needs to be later than the first instruction reads the register.
[0046] Sleep-wake mechanism: a data hazard dependency release mechanism. If the second instruction depends on the execution result of the first instruction, the second instruction is put into sleep mode before the first instruction is completed. When the first instruction is completed, the second instruction is woken up to continue execution.
[0047] Scoreboard mechanism: Another dependency resolution mechanism for data hazards. The scoreboard is used to record the dependency relationships between instructions. When the dependency between the second instruction and the first instruction is released, the second instruction is immediately executed.
[0048] An instruction is a command that instructs a computer to perform a certain operation, also known as a machine instruction, which is the smallest functional unit for a computer to run. An instruction is a statement in machine language and a set of meaningful binary codes. The set of all instructions of a computer constitutes the instruction system of the computer, also known as the instruction set. Instructions are executed by a processor, such as a CPU and a GPU.
[0049] Figure 1 FIG. 7 shows a schematic structural diagram of a processor 100 provided by an exemplary embodiment of the present application. The processor 100 includes: an instruction fetch unit 10, a decoding unit 20, a dependency processing unit 30, a load / store unit 40, a computing unit 50, a register file 60, a memory 70, and a write-back unit 80.
[0050] The instruction fetch unit 10 is connected to the decoding unit 20, the decoding unit 20 is connected to the dependency processing unit 30, and the dependency processing unit 30 is respectively connected to the load / store unit 40 and the computing unit 50.
[0051] The instruction fetch unit 10 is connected to the memory 70.
[0052] The computing unit 50 is connected to one or more registers in the register file 60. The register file 60 includes registers with different functions and different types. For example, general-purpose registers, instruction registers, etc.
[0053] The load / store unit 40 is respectively connected to the register file 60 and the memory 70, and the load / store unit 40 is also connected to the write-back unit 80. When the processor 100 is a GPU, the memory 70 is also called the video memory. When the processor 100 is a CPU, the memory 70 can also be considered as a device outside the processor 100.
[0054] The write-back unit 80 is respectively connected to the register file 60 and the memory 70, and the write-back unit 80 is also connected to the dependency processing unit 30.
[0055] When the above-mentioned processor 100 executes instructions, it adopts a multi-stage pipeline mechanism to execute. Taking a simple five-stage pipeline as an example, the execution process of the instruction is divided into five basic stages to achieve parallel processing of instructions and improve the efficiency of the processor.
[0056] Such as Figure 2As shown, the five-stage pipeline includes five stages: the Instruction Fetch (IF) stage, the Instruction Decode (ID) stage, the Execution (EX) stage, the Memory Access (MEM) stage, and the Write-Back (WB) stage. The design of the five-stage pipeline allows multiple instructions to be processed simultaneously in different stages, thereby reducing the total execution time of each instruction. This parallelism can significantly improve the performance of the processor, especially in high-throughput scenarios. For example, in Figure 2 during the execution of Instruction 1, its five stages sequentially occupy cycles 1 to 5; during the execution of Instruction 2, its five stages sequentially occupy cycles 2 to 6; during the execution of Instruction 3, its five stages sequentially occupy cycles 3 to 7. Specifically:
[0057] In the instruction fetch stage, the instruction fetch unit 10 fetches the instruction from the memory address pointed to by the Program Counter (PC) and loads it into the instruction cache. The program counter is then updated to the address of the next instruction. Optionally, the program counter is one of the counters between the instruction fetch unit 10 and the instruction decode unit 20, and is used to store the address of the next instruction that the processor 100 needs to process. Figure 1 The instruction cache and the program counter are not shown in
[0058] In the instruction decode stage, the instruction decode unit 20 parses the opcode and operands of the instruction in the instruction cache, determines the operation type of the instruction and the required operations. This stage may also involve converting the memory address or immediate value in the instruction into a form that the processor can understand. After the instruction decode unit 20 completes the decoding, it will issue the instruction to the dependency processing unit 30.
[0059] The operation types of common instructions include at least one of: data transfer type, arithmetic logic type, shift type, transfer type, input / output type. For example, the data transfer type is used to transfer data between memory and registers; another example is the arithmetic logic operation type, which is used to perform arithmetic operations or logical operations. Arithmetic operations include addition, subtraction, multiplication, division, etc., and logical operations include AND, NOT, OR, etc.
[0060] In the execution phase, the dependency processing unit 30 is responsible for processing the dependencies between different instructions. According to the dependencies between the instructions and the operation type of each instruction, each instruction is scheduled to the load storage unit 40 or the computing unit 50 for execution. For example, if the operation type of the instruction is a data transfer type, the dependency processing unit 30 will schedule the instruction to the load storage unit 40 for execution; if the operation type of the instruction is a calculation type, the dependency processing unit 30 will schedule the instruction to the computing unit 50 for execution. For the current instruction that depends on other instructions, the dependency processing unit 30 needs to release the dependency of the current instruction on other instructions before scheduling the current instruction to the load storage unit 40 or the computing unit 50.
[0061] In the memory access phase, the load storage unit 40 or the computing unit 50 will send a memory access request to the memory according to the requirements of different instructions. The memory access request can be a memory read request or a memory write request. The memory read request is used to request to read data in the memory into a register, and the memory write request is used to request to write data in the register into the memory.
[0062] In the write-back stage, the write-back unit 80 writes the data in the memory or the execution result of the instruction into the register corresponding to the target operand.
[0063] In the embodiment of the present application, the load storage unit 40, the calculation unit 50, and the write back unit 80 can all be understood as instruction processing units, that is, units for processing different instructions.
[0064] The present application embodiment relates to an improvement in the execution of load instructions and store instructions by a load storage unit. For more details on the multi-stage pipeline of load instructions and store instructions, please refer to Figure 3 . Figure 3 The execution phase and memory access phase shown each include at least one sub-phase. The execution phase includes at least one sub-phase of dependency release, read control class operands, memory address calculation, and read data class operands. The memory access phase includes at least one sub-phase of issuing a memory request and receiving return information. In different embodiments, the memory request sent to the memory may be a memory read request or a memory write request; the return information may be the data in the memory, or a write success response or a write failure response after writing data to the memory.
[0065] Loading instructions:
[0066] The load instruction is used to read data from memory 70 into a certain register. In different embodiments, the load instruction has different forms. For example, load instruction 1 is "load src dst", which means to load the data at memory address src into register dst. Another example, load instruction 2 is "load R0, [R1,#4]", which means to load the data at the offset memory address obtained by offsetting the base memory address stored in register R1 by 4 bytes into register R0. Still another example, load instruction 3 is "if p, load src dst", which means that if the predicate register p is true, load data from memory address src into the general-purpose register dst.
[0067] The load instruction includes at least one source operand and at least one destination operand. The source operand is used to indicate the execution condition and / or the memory address, and the destination operand is used to indicate the register to be written. In the above examples, src, R1, #4, and p are source operands, and R0 and dst are destination operands. Moreover, these source operands are all control-type source operands.
[0068] Taking the load instruction "if p, load src dst" as an example, in combination with reference Figure 3 and Figure 4 , in the read control-type operand sub-phase, the load storage unit 40 first reads the control-type operand p. p is the value in the predicate register, 1 means true, and 0 means false. Based on p, it is determined whether to load. If loading is required, in the memory address sub-phase, the memory address src is obtained or calculated. Since the load instruction does not need to read data-type operands from registers, there is no read data-type operand sub-phase. In the memory access phase, the load storage unit 40 sends a memory read request to the memory based on the memory address src, and the write-back unit 80 receives the return information, which is the data returned by the memory. In the write-back phase, the write-back unit 80 writes the data into the general-purpose register dst. After the write-back phase ends, the write-back unit 80 sends a dependency release signal to the dependency processing unit 30 to update the dependency information configured for this load instruction. For example, when there are other instructions that depend on the use of the predicate register p and the general-purpose register dst by the load instruction, the write-back unit 80 notifies the dependency processing unit 30 to release the dependencies of other instructions on the current instruction.
[0069] Store instruction:
[0070] The store instruction is used to write the data in a certain register into the memory 70. In different embodiments, the store instruction has different forms. For example, the store instruction 1 is "store src dst", which means to load the data in the register src into the memory address dst. Another example is that the store instruction 2 is "store R0,[R1,#8]", which means to store the data in the register R0 into the offset memory address that is 8 bytes offset from the base memory address stored in the register R1. Still another example is that the store instruction 3 is "ifp, load src dst", which means that if the predicate register p is true, write the data in the register src into the register dst.
[0071] The store instruction includes at least one source operand and at least one destination operand. The source operand is used to indicate the execution condition and / or the register to be read, and the destination operand is used to indicate the memory address. In the above examples, src and p are source operands, and R1, #8, and dst are destination operands. The operand src is a data type operand, and p, R1, #8, and dst are control type operands.
[0072] Taking the store instruction "if p, store src dst" as an example, in combination with reference Figure 3 and Figure 5 , in the sub-phase of reading the control type operand, the load store unit 40 reads the value in the predicate register p to determine whether to store; if storage is required, in the sub-phase of calculating the memory address, obtain or calculate the memory address dst; in the sub-phase of reading the data type operand, the load store unit 40 reads the data in the general register src; in the memory access phase, the load store unit 40 sends a memory write request to the memory based on the data in the general register src and the memory address dst to write the data into the memory. Then, the load store unit 40 sends a dependency release signal to the dependency processing unit 30 to update the dependency information configured for this store instruction. For example, when there are other instructions that depend on the use of the predicate register p and the general register src by the store instruction, the load store unit 40 notifies the dependency processing unit 30 to release the dependency of other instructions on the current instruction.
[0073] Since the same instruction may involve multiple operands, and multiple operands may correspond to different registers. In the related art, the same instruction only supports one type of dependency relationship. If there are multiple other instructions that depend on the current instruction at the same time, this one type of dependency relationship is uniformly used for maintenance and release. However, the usage timing and usage duration of different registers for the same instruction are different. Using the same dependency relationship to maintain different registers will cause a part of the registers that are used up first to be occupied for a long time until the last register is used up, and then the dependency of other instructions on the current instruction can be uniformly released.
[0074] An embodiment of the present application provides a solution for releasing instruction dependencies, hoping to divide multiple registers involved in the same instruction into at least two categories, wherein the dependencies corresponding to the registers that are used first are released first, and the dependencies corresponding to the registers that are used later are released later, so that the registers that are used first can be used by other instructions earlier.
[0075] Figure 6 The schematic diagram of the structure of a processor 200 provided by an exemplary embodiment of the present application is shown. The processor 200 at least includes: an instruction processing unit 90 and a dependency processing unit 30.
[0076] The instruction processing unit 90 is a unit for executing different instructions. Optionally, the instruction processing unit 90 is a unit for processing instructions after decoding. The instruction processing unit 90 includes: at least one of a load storage unit and a write back unit. It is not excluded that in other embodiments, the instruction processing unit 90 includes a calculation unit.
[0077] The dependency processing unit 30 is a unit for processing the dependency relationship between different instructions. Optionally, the dependency processing unit 30 is used to configure the dependency relationship between different instructions and to release the dependency relationship between different instructions.
[0078] In this embodiment, at least two paths are provided between the instruction processing unit 90 and the dependency processing unit 30 , and the at least two paths include: a first type path 32 and a second type path 34 .
[0079] The instruction processing unit 90 is used to send a first dependency release signal to the dependency processing unit 30 through the first type path 32; the dependency processing unit 30 is used to release the dependency relationship corresponding to the first type register based on the first dependency release signal.
[0080] The instruction processing unit 90 is further configured to send a second dependency release signal to the dependency processing unit 30 through the second type path 34; the dependency processing unit 30 is configured to release the dependency relationship corresponding to the second type register based on the second dependency release signal.
[0081] The first type register and the second type register are both registers used by the same instruction, and the same instruction ends using the first type register and the second type register at different times.
[0082] Optionally, the dependency release signal is a signal sent by the instruction processing unit 90 to the dependency processing unit 30 for releasing the dependency relationship associated with a certain instruction. In the case where an instruction supports at least two types of dependency relationships, the dependency release signal is used to release at least one of the dependency relationships associated with a certain instruction. Optionally, the dependency release signal carries at least one of the information such as instruction identifier, dependency relationship identifier, register identifier, release reason, status information, timestamp, resource release information.
[0083] Optionally, the first type of register is the register corresponding to the first type of operand, and the second type of register is the register corresponding to the second type of operand. Both the first type of operand and the second type of operand are operands of the same instruction. Optionally, the sending times of the first dependency release signal and the second dependency release signal are different. For example, the sending time of the first dependency release signal is earlier than the sending time of the second dependency release signal; or, the sending time of the second dependency release signal is earlier than the sending time of the first dependency release signal.
[0084] The path is a signal transmission path implemented based on hardware, and this signal transmission path includes at least one of the circuits, caches, and other possible electronic devices in the chip. That is, the path is a signal transmission path composed of at least one of the metal lines, caches, and other possible electronic devices on the chip. The path in this embodiment is a logical concept, which is used to represent the signal transmission path for the instruction processing unit 90 to send the dependency release signal to the dependency processing unit 30.
[0085] In some embodiments, the dependency processing unit 30 is integrated in the decoding unit 20; in some embodiments, the dependency processing unit 30 is independent outside the decoding unit 20; in some embodiments, the dependency processing unit 30 includes the decoding unit 20 or a part of the functions of the decoding unit 20; in some embodiments, the dependency processing unit 30 cooperates with the decoding unit 20 to complete dependency processing. The embodiments of the present application do not limit the implementation form of the dependency processing unit 30.
[0086] In some embodiments, the above-mentioned processor 200 further includes at least one unit such as an instruction fetching unit, a decoding unit, a cache, a memory, and a register. The embodiments of the present application do not limit the other units further included in the processor 200 and the connection forms between the other units.
[0087] In some embodiments, there may be three or more signal paths between the instruction processing unit 90 and the dependency processing unit 30, and the first type of path and the second type of path are any two of the three or more signal paths. At this time, the registers used by the same instruction can be three or more, and are released according to three or more release times, and the embodiments of the present application do not limit this.
[0088] In summary, the processor provided in this embodiment sets a first type of path and a second type of path between the instruction processing unit and the dependency processing unit. For two types of registers used by the same instruction, the first type of path is used to release the first dependency and the second type of path is used to release the second dependency. Since the same instruction uses different types of registers at different times and for different durations, the dependency corresponding to the register that is used first can be released earlier, thereby reducing the duration of the instruction occupying the register that is used first and improving the efficiency of using this part of the registers. At the same time, since this part of the registers can be released earlier, the possibility of parallel instructions between different instructions is increased.
[0089] In different embodiments, there are many possible design methods for the first type of path 32 and the second type of path 34. Exemplarily, the first type of path 32 and the second type of path 34 are any two paths among the three candidate paths. Figure 7 As shown, the three candidate pathways include: a first pathway, a second pathway, and a third pathway.
[0090] The first path is a signal transmission path between the load storage unit 40 and the dependency processing unit 30 .
[0091] The second path is a signal transmission path between the load storage unit 40 and the dependency processing unit 30 .
[0092] The third path is a signal transmission path between the write-back unit 80 and the dependency processing unit 30 .
[0093] The three candidate paths send transmission dependency release signals at different times. Figure 8 The timing at which three candidate paths transmit dependency release signals in a multi-stage pipeline is shown.
[0094] The load storage unit 40 is used to send a first dependency release signal to the dependency processing unit 30 through a first path based on a first time T1, where the first time T1 is the end time of reading the control type source operand.
[0095] The load storage unit 40 is used to send a first dependency release signal or a second dependency release signal to the dependency processing unit 30 through the second path based on the second time T2, and the second time T2 is the end time of issuing a memory read request or a memory write request. Exemplarily, for a load instruction, the second time T2 is the end time of issuing a memory read request, and the load storage unit 40 sends the first dependency release signal to the dependency processing unit 30 through the second path at the second time T2; for a store instruction, the second time T2 is the end time of issuing a memory write request, and the load storage unit 40 sends the second dependency release signal to the dependency processing unit 30 through the second path at the second time T2.
[0096] A write-back unit 80, configured to send a second dependency release signal to a dependency processing unit 30 through a third path based on a third time T3, where the third time T3 is an end time of a write-back stage, and the write-back stage is a stage of writing data back to a register.
[0097] It should be noted that "based on xx time" can be understood as: at xx time, or at a subsequent time slightly later than xx time, or at a time indicated by the sum of xx time and a delay time. The delay time includes the time required for the load / store unit 40 or the write-back unit 80 to generate a dependency release signal.
[0098] In different embodiments, the processor includes all or part of the candidate paths among the above three candidate paths. The embodiments of the present application at least provide the following embodiments:
[0099] · The first type of path is the first path ①, and the second type of path is the second path ②;
[0100] · The first type of path is the first path ①, and the second type of path is the third path ③;
[0101] · The first type of path is the second path ②, and the second type of path is the third path ③.
[0102] The following introduces the above several embodiments respectively.
[0103] Figure 9 FIG. shows a schematic diagram of dependency release of a load instruction provided by an exemplary embodiment of the present application. Assume that the first type of path includes the first path ①, and the second type of path includes the third path ③.
[0104] For a load instruction, the execution stage includes three sub-stages: dependency processing, reading control-class source operands, and memory address calculation. After the load / store unit 40 reads the control-class source operands, the subsequent stages of the current instruction no longer need to use the first type of register corresponding to the control-class source operands. At the end time T1 of reading the control-class operands, the load / store unit 40 uses the first path ① to send a first dependency release signal to the dependency processing unit 30, and the dependency processing unit 30 releases the dependency relationship corresponding to the first type of register based on the first dependency release signal.
[0105] Taking the load instruction as: if p, load src dst as an example, in combination with reference to Figure 3 and Figure 9, in the read control operand sub-phase, the load / store unit 40 reads the control source operand p from the predicate register to determine whether to load based on the control source operand p. The read control source operand p is cached in the load / store unit 40. Since the predicate register is no longer needed in the subsequent stages of the load instruction, the load / store unit 40 sends a first dependency release signal to the dependency processing unit 30 through the first path ① at the end time T1 of reading the control operand, so as to trigger the dependency processing unit 30 to release the first dependency corresponding to the predicate register. For example, if there is a read-after-write dependency between the load instruction and other instructions, after sending the first dependency release signal, other instructions can start or continue to execute without waiting for the write-back stage of the load instruction to end.
[0106] At the end time T3 of the write-back phase, the write-back unit 80 writes the data in the memory address src into the general register corresponding to the destination operand dst. At this time, the write-back unit 80 sends a second dependency release signal to the dependency processing unit 30 through the third path ③, and the dependency processing unit 30 releases the second dependency corresponding to the general register based on the second dependency release signal.
[0107] In the related art, the dependency corresponding to the load instruction needs to be uniformly released at the third time T3, and only then will the predicate register and the general register used by the load instruction be released simultaneously. Compared with the related art, in the embodiment of the present application, the predicate register can be released at the first time T1, which can greatly advance the release time of the predicate register, thereby improving the usage efficiency of the predicate register.
[0108] Figure 10 Shows a schematic diagram of dependency release of a store instruction provided by an exemplary embodiment of the present application. Assume that the first type of path includes the first path ① and the second type of path includes the second path ②.
[0109] For the store instruction, the execution phase includes four sub-phases: dependency processing, reading the control source operand, memory address calculation, and reading the data source operand sub-phase. After reading the control source operand, the first type of register corresponding to the control source operand is no longer needed in the subsequent stages of the store instruction. The load / store unit 40 sends a first dependency release signal to the dependency processing unit 30 through the first path ① at the first time T1, and the dependency processing unit 30 releases the dependency corresponding to the first type of register based on the first dependency release signal. In the memory address calculation sub-phase, the load / store unit 40 obtains or calculates the memory address based on the destination operand; in the read data source operand sub-phase, it reads the data in the general register.
[0110] In the memory access stage, the load / store unit 40 issues a memory write request, which carries the data type source operand and the memory address. The data type source operand and the memory address will be stored in the cache first and then written into the memory. Therefore, the second type register corresponding to the data type source operand and / or the memory address will no longer be needed. At the second moment T2 after the load / store unit 40 issues the memory write request, the load / store unit 40 uses the second path ② to send a second dependency release signal to the dependency processing unit 30, and the dependency processing unit 30 releases the dependency relationship corresponding to the second type register based on this second dependency release signal.
[0111] Taking the store instruction as: if p, store src dst as an example, in combination with reference Figure 3 and Figure 9 , in the read control type operand sub-stage, the load / store unit 40 reads the control type source operand p from the predicate register to determine whether to load based on the control type source operand p. The read control type source operand p will be cached in the load / store unit 40. Since the predicate register is no longer needed in the subsequent stages of the store instruction, the load / store unit 40 sends a first dependency release signal to the dependency processing unit 30 through the first path ① at the end moment T1 of the read control type source operand sub-stage, so as to trigger the dependency processing unit 30 to release the first dependency relationship corresponding to the predicate register based on the first dependency release signal. For example, if there is a read-after-write dependency between this store instruction and other instructions, after sending the first dependency release signal, other instructions can start execution or continue execution without waiting for the recall stage of this load instruction to end. In the read data type source operand sub-stage, the load / store unit 40 reads the data from the general register corresponding to the data type source operand src.
[0112] In the memory access stage, the load / store unit 40 issues a memory write request, which carries the data type source operand src and the memory address dst. The data type source operand src and the memory address dst will be stored in the cache first and then written into the memory. Therefore, after issuing the memory write request, the load / store unit 40 uses the second path ② to send a second dependency release signal to the dependency processing unit 30, and the dependency processing unit 30 releases the instruction dependency relationship corresponding to the second type register based on the second dependency release signal. The second type register is the general register corresponding to the data type source operand src.
[0113] In the related art, the dependency relationship corresponding to the store instruction needs to be uniformly released at the second moment T2, and only then will the predicate register and the general register used by the store instruction be released simultaneously. Compared with the related art, in the embodiment of the present application, the predicate register can be released at the first moment T1, which can greatly advance the release timing of the predicate register, thereby improving the usage efficiency of the predicate register.
[0114] Figure 11 A schematic diagram of dependency release of a load instruction provided by an exemplary embodiment of the present application is shown. Assume that the first type of path includes the second path ②, and the second type of path includes the third path ③.
[0115] For the load instruction, the execution phase includes three sub-phases: dependency processing, reading control source operands, and memory address calculation. After the load storage unit 40 reads the control source operand, the subsequent phases of the current instruction do not need to use the first type register corresponding to the control source operand. However, due to the absence of the first path ①, the load storage unit 40 waits for the second moment T2 after issuing the memory read request, and uses the second path ② to send a first dependency release signal to the dependency processing unit 30. The dependency processing unit 30 releases the first dependency relationship corresponding to the first type register based on the first dependency release signal.
[0116] Take the load instruction: if p, load src dst as an example, combined with reference Figure 3 and Figure 9 In the sub-phase of reading control class operands, the load storage unit 40 reads the control class source operand p from the predicate register to determine whether loading is required based on the control class source operand p. The control class source operand p after reading will be cached in the load storage unit 40. However, due to the absence of the first path ①, the load storage unit 40 waits for the second moment T2 after the memory read request is issued, and uses the second path ② to send the first dependency release signal to the dependency processing unit 30 to trigger the dependency processing unit 30 to release the first dependency relationship corresponding to the predicate register. For example, if there is a read-before-write dependency between the load instruction and other instructions, after sending the first dependency release signal, other instructions can start or continue to execute without waiting for the write-back phase of the load instruction to end.
[0117] At the end time T3 of the write-back phase, the write-back unit 80 writes the data in the memory address src to the general register corresponding to the target operand dst. At this time, the write-back unit 80 sends a second dependency release signal to the dependency processing unit 30 through the third path ③, and the dependency processing unit 30 releases the second dependency relationship corresponding to the general register based on the second dependency release signal.
[0118] In the related art, the dependencies corresponding to the load instruction need to be released uniformly at the third time T3, at which time the predicate register and the general register used by the load instruction will be released at the same time. Compared with the related art, the predicate register in the embodiment of the present application can be released at the second time T2, which can greatly advance the release time of the predicate register, thereby improving the use efficiency of the predicate register.
[0119] Since there is no sub - stage for reading data - type operands in the execution stage of the load instruction, compared with Figure 9 the embodiment, the performance difference of the predicate register released at the second moment T2 compared with that released at the first moment T1 is not significant.
[0120] From the above - mentioned various embodiments, it can be seen that the embodiments of the present application can release registers in advance, reduce the occupancy time overhead of actual registers, and increase the usage efficiency of registers. Especially in the scenario where there is a read - before - write dependency between two instructions, after the first instruction responsible for reading finishes reading the first - type register corresponding to the control - type operand, the first - type register can be immediately released, and the second instruction responsible for writing later can use the first - type register earlier. Therefore, in the scenario of read - before - write dependency, the embodiments of the present application can greatly improve the usage efficiency of registers.
[0121] At this time, under the same compilation algorithm tendency, the embodiments of the present application can also reduce register allocation to a greater extent. The embodiments of the present application cooperate with the compilation algorithm to work, and can increase the parallel probability of load instructions under the condition of the same register usage. For example, a stream of instructions is given as an example:
[0122] / / Start of the example code segment
[0123] Load instruction 1 destination operand, source operand 1, source operand 2;
[0124] (Update the register where source operand 1 is located)
[0125] Load instruction 2 destination operand, source operand 3, source operand 4; / / Source operand 4 re - uses the register where source operand 1 is located
[0126] (Update the register where source operand 4 is located)
[0127] Load instruction 3 destination operand, source operand 5, source operand 6; / / Source operand 6 re - uses the register where source operand 4 is located
[0128] Compute instruction;
[0129] / / End of the example code segment
[0130] In the above - mentioned example code segment, if the two dependency relationships used by load instruction 1 are met, the register where source operand 1 is located can be updated in advance before load instruction 1 writes to the register corresponding to the destination operand after reading source operand 1. The same applies to the next load instruction 2. Assuming that the update operation cost is small enough, the memory read requests of load instruction 1, load instruction 2, and load instruction 3 can be executed in parallel on the storage system.
[0131] Based on the above respective embodiments, the dependency processing unit 30 is configured to, based on the first dependency release signal, adopt a first dependency release mechanism to process and release the first dependency relationship corresponding to the first type of register; and based on the second dependency release signal, adopt a second dependency release mechanism to process and release the second dependency relationship corresponding to the second type of register.
[0132] Optionally, the first dependency release mechanism is a scoreboard mechanism, and the second dependency release mechanism is a sleep-wake mechanism.
[0133] The scoreboard mechanism records the dependency relationships between different instructions through a scoreboard. When the scoreboard monitors that the dependency between the second instruction and the first instruction is released, the second instruction is immediately executed.
[0134] The sleep-wake mechanism is that if the second instruction depends on the execution result of the first instruction, when the first instruction has not been executed completely, the second instruction is put into a sleep state; after the first instruction is executed completely, the second instruction is woken up to continue execution.
[0135] Compared with the sleep-wake mechanism, the scoreboard mechanism has a smaller latency. Therefore, using the scoreboard mechanism to release the first dependency relationship can release the first type of register earlier. This is because under the scoreboard mechanism, the instructions whose dependencies have not been released after decoding will wait in the instruction queue and be directly issued to the subsequent units for execution after the dependencies are released; while the sleep-wake mechanism needs to be rescheduled after the dependencies are released and can be issued to the subsequent units for execution only after passing through the fetch stage and the decoding stage in sequence. Therefore, using the scoreboard mechanism for the first type of register that is used up first and the sleep-wake mechanism for the second type of register that is used up later can make the first type of register released earlier.
[0136] Based on the above respective embodiments, the above processor further includes a decoding unit 20, configured to set the first dependency relationship and the second dependency relationship for the above instructions respectively in the decoding stage. Optionally, the decoding unit 20 analyzes the dependency relationships between different instructions, such as at least one of read-after-write, write-after-write, and write-after-read introduced above. Assume that the current instruction is the first instruction. If both the first instruction and the second instruction use the first type of register and the second instruction depends on the first instruction, a first dependency relationship is configured for the first instruction; if both the first instruction and the third instruction use the second type of register and the third instruction depends on the first instruction, a second dependency relationship is configured for the first instruction.
[0137] When configuring the first dependency relationship, the decoding unit 20 allocates a first resource for the first instruction, and the first resource is the resource required for the first dependency relationship, such as the counter (Counter) number used to independently count the first dependency relationship; when configuring the second dependency relationship, the decoding unit 20 allocates a second resource for the first instruction, and the second resource is the resource required for the second dependency relationship, such as the counter used to independently count the second dependency relationship.
[0138] In some embodiments, the decoding unit 20 is configured to send a first dependency configuration signal to the dependency processing unit 30, and the dependency processing unit 30 is configured to configure the first dependency relationship for the above-mentioned instruction based on the first dependency configuration signal; the decoding unit 20 is further configured to send a second dependency configuration signal to the dependency processing unit 30, and the dependency processing unit 30 is configured to configure the second dependency relationship for the above-mentioned instruction based on the second dependency configuration signal.
[0139] The decoding unit 20 is further configured to, in the case where there is insufficient resource to set the first and second dependency relationships, wait for the retransmission condition to be satisfied and then retransmit the address of the (first) instruction to the program counter, and the program counter is used to store the address of the next instruction to be executed by the processor. Taking the first resource as the resource required for the first dependency relationship and the second resource as the resource required for the second dependency relationship as an example, the retransmission condition is that the first resource required by the instruction is sufficient and the second resource is sufficient.
[0140] There is only one case of insufficient resource:
[0141] In some embodiments, if the first resource required by the instruction is insufficient and the second resource is sufficient, the instruction can be set to the sleep state, and the wake-up condition is set to that the idle first resource is greater than or equal to the first resource required by the instruction; after the instruction is awakened again when the first resource is sufficient, it is detected again whether the above retransmission condition is satisfied, and if the retransmission condition is satisfied, the address of the instruction is retransmitted to the program counter. If the first resource required by the instruction is sufficient and the second resource is insufficient, the instruction can be set to the sleep state, and the wake-up condition is set to that the idle second resource is greater than or equal to the second resource required by the instruction; after the instruction is awakened again when the second resource is sufficient, it is detected again whether the above retransmission condition is satisfied, and if the retransmission condition is satisfied, the address of the instruction is retransmitted to the program counter.
[0142] In other embodiments, if the first resource required by the instruction is insufficient and the second resource is sufficient, only the second dependency relationship can be configured for the instruction and the first dependency relationship is not configured; if the first resource required by the instruction is sufficient and the second resource is insufficient, only the first dependency relationship can be configured for the instruction and the second dependency relationship is not configured. This can enable the current instruction to run earlier.
[0143] The case where both resources are insufficient:
[0144] In some embodiments, if the first resource required by the instruction is insufficient and the second resource is also insufficient, the instruction can be set to the sleep state. However, due to software and hardware limitations, the wake-up condition cannot be set to the first resource being sufficient and the second resource being sufficient (that is, it is impossible to set 2 conditions simultaneously, only 1 condition can be set). At the same time, since the release duration of the first dependency of the historical instructions running previously is usually greater than that of the second dependency, the wake-up condition set for this instruction is that the idle first resource is greater than or equal to the first resource required by the instruction. After the instruction is woken up again when the first resource is sufficient, it is detected again whether the above retransmission condition is satisfied. If the retransmission condition is satisfied, the address of the instruction is retransmitted to the program counter. If there is still one kind of resource insufficient, it can be processed according to the situation where only one kind of resource is insufficient as described above.
[0145] In this embodiment, the decoding unit configures two types of dependencies for the instruction, enabling the same instruction to support two different dependencies, and then supporting the release of different types of registers at different times. At the same time, when the resources for configuring the two types of dependencies are insufficient, the sleep-wake mechanism is used to temporarily put the first instruction in the sleep state, and then retransmit it to the program counter for execution when the resources are sufficient, ensuring the smooth execution of the instruction.
[0146] Based on the above various embodiments, the load / store unit 40 is further configured to end the execution of the (first) instruction, and release the first dependency and the second dependency in the case of an internal exception. The load / store unit 40 may encounter an internal exception during the execution phase of the first instruction. When an internal exception occurs, the load / store unit 40 sends an exception signal to the decoding unit 20 or other exception handling units to end the execution of the instruction. If the instruction is configured with both the first dependency and the second dependency at the same time, the load / store unit 40 also needs to notify the dependency processing unit 30 to release the first dependency and the second dependency.
[0147] In this embodiment, by the load / store unit releasing the first dependency and the second dependency simultaneously in the case of an internal exception, an appropriate exception handling mechanism can be implemented in the scenario where the same instruction supports two dependencies, ensuring the normal operation of the multi-stage pipeline.
[0148] Figure 12 The flowchart of a dependency release method provided by an exemplary embodiment of the present application is shown. The method is executed by the processor provided by the above various embodiments. The processor includes: an instruction processing unit, a dependency processing unit, and at least two paths provided between the instruction processing unit and the dependency processing unit. The at least two paths include a first type of path and a second type of path. The method includes:
[0149] Step 222: The instruction processing unit sends a first dependency release signal to the dependency processing unit through the first type of path;
[0150] Step 224: The dependency processing unit releases the first dependency relationship corresponding to the first type of register based on the first dependency release signal;
[0151] Step 242: The instruction processing unit sends a second dependency release signal to the dependency processing unit through the second type of path;
[0152] Step 244: The dependency processing unit releases the second dependency relationship corresponding to the second type of register based on the second dependency release signal;
[0153] Wherein, both the first type of register and the second type of register are registers used by the same instruction. The end times of using the first type of register and the second type of register by the same instruction are different.
[0154] Optionally, the instruction processing unit is a unit for executing different instructions. Optionally, the instruction processing unit is a unit for processing instructions after decoding. The instruction processing unit includes at least one of a load / store unit and a write-back unit. It does not exclude the possibility that the instruction processing unit includes a computing unit in other embodiments.
[0155] Optionally, the dependency processing unit is a unit for processing dependency relationships between different instructions. Optionally, the dependency processing unit 30 is used to configure dependency relationships between different instructions and release dependency relationships between different instructions.
[0156] Optionally, the dependency release signal is a signal sent by the instruction processing unit to the dependency processing unit for releasing the dependency relationship associated with a certain instruction. In the case where an instruction supports at least two dependency relationships, the dependency release signal is used to release at least one of the dependency relationships associated with a certain instruction. Optionally, the dependency release signal carries at least one of information such as an instruction identifier, a dependency relationship identifier, a register identifier, a release reason, status information, a timestamp, and resource release information.
[0157] Optionally, the first type of register is a register corresponding to the first type of operand, and the second type of register is a register corresponding to the second type of operand. Both the first type of operand and the second type of operand are operands of the same instruction. Optionally, the sending times of the first dependency release signal and the second dependency release signal are different. For example, the sending time of the first dependency release signal is earlier than the sending time of the second dependency release signal; or, the sending time of the second dependency release signal is earlier than the sending time of the first dependency release signal.
[0158] A path is a signal transmission path implemented based on hardware, and this signal transmission path includes at least one of circuits, caches, and other possible electronic devices in a chip. That is to say, a path is a signal transmission path composed of at least one of metal lines, caches, and other possible electronic devices on a chip. The path in this embodiment is a logical concept, used to represent the signal transmission path for an instruction processing unit to send a dependency release signal to a dependency processing unit.
[0159] In some embodiments, the dependency processing unit is integrated in the decoding unit; in some embodiments, the dependency processing unit is independent outside the decoding unit; in some embodiments, the dependency processing unit includes the decoding unit or a part of the functions of the decoding unit; in some embodiments, the dependency processing unit cooperates with the decoding unit to complete dependency processing. The implementation form of the dependency processing unit in the embodiments of this application is not limited.
[0160] In some embodiments, the above-mentioned processor further includes at least one unit among an instruction fetching unit, a decoding unit, a cache, a memory, and a register. The embodiments of this application are not limited to other units that the processor further includes and the connection forms between other units.
[0161] In some embodiments, there can be three or more signal paths between the instruction processing unit and the dependency processing unit. The first type of path and the second type of path are any two of the three or more signal paths. At this time, the registers used by the same instruction can be three or more, and are released according to three or more release times. The embodiments of this application do not limit this.
[0162] In summary, the method provided in this embodiment, by setting the first type of path and the second type of path between the instruction processing unit and the dependency processing unit. For two types of registers used by the same instruction, the first type of path is used to release the first dependency relationship and the second type of path is used to release the second dependency relationship. Since the usage times and usage durations of different types of registers by the same instruction are different, it can enable the dependency relationship corresponding to the register that is used up first to be released earlier, thereby reducing the occupation duration of the register that is used up first by this instruction and improving the usage efficiency of this part of the registers. At the same time, since this part of the registers can be released earlier, the possibility of parallel instructions between different instructions is increased.
[0163] In different embodiments, there are various possible design methods for the first type of path and the second type of path. Exemplarily, the first type of path and the second type of path are any two paths among three candidate paths. The three candidate paths include: the first path, the second path, and the third path.
[0164] The first path is the signal transmission path between the load / store unit and the dependency processing unit.
[0165] The second path is another signal transmission path between the load / store unit and the dependency processing unit.
[0166] The third path is the signal transmission path between the write-back unit and the dependency processing unit.
[0167] These three candidate paths send the transmission dependency release signal at different times.
[0168] The load / store unit is configured to send a first dependency release signal to the dependency processing unit through the first path based on a first time T1, where the first time T1 is the end time of reading the control type source operand.
[0169] The load / store unit is configured to send a first dependency release signal or a second dependency release signal to the dependency processing unit through the second path based on a second time T2, where the second time T2 is the end time of issuing a memory read request or a memory write request. Exemplarily, for a load instruction, the second time T2 is the end time of issuing a memory read request, and the load / store unit sends a first dependency release signal to the dependency processing unit through the second path at the second time T2; for a store instruction, the second time T2 is the end time of issuing a memory write request, and the load / store unit sends a second dependency release signal to the dependency processing unit through the second path at the second time T2.
[0170] The write-back unit 80 is configured to send a second dependency release signal to the dependency processing unit through the third path based on a third time T3, where the third time T3 is the end time of the write-back stage, and the write-back stage is the stage of writing data back to the register.
[0171] It should be noted that "based on xx time" can be understood as: at xx time, or at a subsequent time slightly later than xx time, or at the time indicated by the sum of xx time and the delay time. The delay time includes the time required for the load / store unit to generate the dependency release signal.
[0172] In different embodiments, the processor includes all or part of the three candidate paths described above. The embodiments of the present application at least provide the following embodiments:
[0173] · The first type of path is the first path, and the second type of path is the second path;
[0174] · The first type of path is the first path, and the second type of path is the third path;
[0175] · The first type of path is the second path, and the second type of path is the third path.
[0176] The following is an introduction to the above several embodiments respectively.
[0177] In some embodiments, with reference Figure 9and Figure 9 and related descriptions, the instruction processing unit includes a load / store unit and a write-back unit; the first type of path includes: a first path between the load / store unit and the dependency processing unit; the second type of path includes: a third path between the write-back unit and the dependency processing unit; the instructions include load instructions;
[0178] Among them, the instruction processing unit sends a first dependency release signal through the first type of path, including:
[0179] The load / store unit sends a first dependency release signal through the first path based on a first moment, and the first moment is the end moment of reading the first type of operand;
[0180] Among them, the instruction processing unit sends a second dependency release signal through the second type of path, including:
[0181] The write-back unit sends a second dependency release signal through the third path based on a third moment, and the third moment is the end moment of the write-back stage;
[0182] Among them, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and a write-back stage, and the first moment is a moment in the execution stage.
[0183] In the related art, the dependency relationship corresponding to the load instruction needs to be uniformly released at the third moment T3. At this time, the predicate register and the general-purpose register used by the load instruction will be released simultaneously. Compared with the related art, in the embodiment of the present application, the register corresponding to the first type of operand can be released at the first moment T1, which can greatly advance the release timing of the register, thereby improving the usage efficiency of the register.
[0184] In some embodiments, in combination with reference to Figure 10 and Figure 10 related descriptions, the instruction processing unit includes a load / store unit; the first type of path includes: a first path between the load / store unit and the dependency processing unit; the second type of path includes: a second path between the load / store unit and the dependency processing unit; the above instructions include store instructions;
[0185] Among them, the instruction processing unit sends a first dependency release signal through the first type of path, including:
[0186] The load / store unit sends a first dependency release signal through the first path based on a first moment, and the first moment is the end moment of reading the first type of operand;
[0187] Among them, the instruction processing unit sends a second dependency release signal through the second type of path, including:
[0188] The load / store unit sends a second dependency release signal through a second path based on a second moment, which is the end moment of issuing a memory write request;
[0189] Among them, the store instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and a memory access stage, the first moment is a moment in the execution stage, and the second moment is a moment in the memory access stage.
[0190] In the related art, the dependency relationships corresponding to store instructions need to be uniformly released at the second moment T2. Only at this time will the predicate register and the general-purpose register used by the store instruction be released simultaneously. Compared with the related art, in the embodiments of the present application, the register corresponding to the first type of operand can be released at the first moment T1, which can greatly advance the release timing of this register, thereby improving the usage efficiency of this register.
[0191] In some embodiments, with reference to Figure 11 , the instruction processing unit includes a load / store unit and a write-back unit; the first type of path includes: a second path between the load / store unit and the dependency processing unit; the second type of path includes: a third path between the write-back unit and the dependency processing unit; the instruction is a load instruction;
[0192] Among them, the instruction processing unit sends a first dependency release signal through the first type of path, including:
[0193] The load / store unit sends a first dependency release signal through a second path based on a second moment, which is the end moment of issuing a memory read request;
[0194] Among them, the instruction processing unit sends a second dependency release signal through the second type of path, including:
[0195] The write-back unit sends a second dependency release signal through a third path based on a third moment, which is the end moment of the write-back stage;
[0196] Among them, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes a memory access stage and a write-back stage, and the second moment is a moment in the memory access stage.
[0197] In the related art, the dependency relationships corresponding to load instructions need to be uniformly released at the third moment T3. Only at this time will the predicate register and the general-purpose register used by the load instruction be released simultaneously. Compared with the related art, in the embodiments of the present application, a part of the registers can be released at the second moment T2, which can greatly advance the release timing of this part of the registers, thereby improving the usage efficiency of this register.
[0198] Since there is no sub-stage for reading data operands in the execution stage of the load instruction, therefore, compared with Figure 9Compared with the embodiments, the performance difference of releasing this part of registers at the second moment T2 and at the first moment T1 is not significant.
[0199] Based on the above various embodiments, it can be seen that the embodiments of the present application can release registers in advance, reduce the occupancy time overhead of actual registers, and increase the usage efficiency of registers. Especially in the scenario where there is a read-after-write dependency between two instructions, after the first instruction responsible for reading reads the first-type register corresponding to the control operand, the first-type register can be immediately released, and the second instruction responsible for writing later can use the first-type register earlier. Therefore, in the scenario of read-after-write dependency, the embodiments of the present application can greatly improve the usage efficiency of registers.
[0200] At this time, under the same compilation algorithm tendency, the embodiments of the present application can also reduce register allocation to a greater extent. The embodiments of the present application cooperate with the compilation algorithm to work, and can increase the parallel probability of load instructions under the condition of the same register usage.
[0201] In some embodiments, the dependency processing unit releases the first dependency relationship corresponding to the first-type register based on the first dependency release signal, including: the dependency processing unit processes and releases the first dependency relationship corresponding to the first-type register based on the first dependency release signal by using a first dependency release mechanism;
[0202] The dependency processing unit releases the second dependency relationship corresponding to the second-type register based on the second dependency release signal, including: the dependency processing unit processes and releases the second dependency relationship corresponding to the second-type register based on the second dependency release signal by using a second dependency release mechanism.
[0203] Among them, the first dependency release mechanism is a scoreboard mechanism, and the second dependency release mechanism is a sleep-wake mechanism.
[0204] Compared with the sleep-wake mechanism, the scoreboard mechanism has a smaller delay. Therefore, using the scoreboard mechanism to release the first dependency relationship can release the first-type register earlier. This is because under the scoreboard mechanism, the instructions whose dependencies have not been released after decoding will wait in the instruction queue and be directly issued to the subsequent units for execution after the dependencies are released; while the sleep-wake mechanism needs to be rescheduled after the dependencies are released and can be issued to the subsequent units for execution only after passing through the instruction fetch stage and the decoding stage in sequence. Therefore, using the scoreboard mechanism for the first-type register that is used up first and the sleep-wake mechanism for the second-type register that is used up later can make the first-type register be released earlier.
[0205] In some embodiments, the processor further includes a decoding unit; the above method further includes:
[0206] During the decoding stage, the decoding unit sets a first dependency and a second dependency for the instruction respectively.
[0207] In this embodiment, the decoding unit configures two dependencies for the instruction, enabling the same instruction to support two different dependencies, and then supporting the release of different types of registers at different times.
[0208] In some embodiments, the above method further includes:
[0209] When the decoding unit is unable to set the first dependency and the second dependency due to insufficient resources, it waits until the retransmission condition is met and then retransmits the address of the instruction to the program counter, which is used to store the address of the next instruction to be executed by the processor. Taking the first resource as the resource required for the first dependency and the second resource as the resource required for the second dependency, the retransmission condition is that the first resource required for the instruction is sufficient and the second resource is sufficient.
[0210] The case where only one kind of resource is insufficient:
[0211] In some embodiments, if the first resource required for the instruction is insufficient and the second resource is sufficient, the decoding unit sets the instruction to the sleep state and sets the wake-up condition as the available first resource being greater than or equal to the first resource required for the instruction; after the instruction is woken up again when the first resource is sufficient, it checks again whether the above retransmission condition is met, and if the retransmission condition is met, it retransmits the address of the instruction to the program counter. If the first resource required for the instruction is sufficient and the second resource is insufficient, the decoding unit sets the instruction to the sleep state and sets the wake-up condition as the available second resource being greater than or equal to the second resource required for the instruction; after the instruction is woken up again when the second resource is sufficient, it checks again whether the above retransmission condition is met, and if the retransmission condition is met, it retransmits the address of the instruction to the program counter.
[0212] In some other embodiments, if the first resource required for the instruction is insufficient and the second resource is sufficient, the decoding unit configures only the second dependency for the instruction and does not configure the first dependency; if the first resource required for the instruction is sufficient and the second resource is insufficient, the decoding unit configures only the first dependency for the instruction and does not configure the second dependency. This can enable the current instruction to run earlier.
[0213] The case where both resources are insufficient:
[0214] In some embodiments, if the first resource required by the instruction is insufficient and the second resource is also insufficient, the decoding unit sets the instruction to the sleep state. However, due to software and hardware limitations, the wake-up condition cannot be set to the first resource being sufficient and the second resource being sufficient (that is, it is impossible to set 2 conditions simultaneously, only 1 condition can be set). At the same time, since the release duration of the first dependency of the historical instructions running previously is generally longer than that of the second dependency, the wake-up condition set for this instruction is that the idle first resource is greater than or equal to the first resource required by the instruction. After the instruction is woken up again when the first resource is sufficient, it is detected again whether the above retransmission condition is satisfied. If the retransmission condition is satisfied, the address of the instruction is retransmitted to the program counter. If there is still one kind of resource insufficient, it can be processed according to the situation where only one kind of resource is insufficient as described above.
[0215] In the embodiments of the present application, when the resources for configuring two dependencies are insufficient, a sleep-wake mechanism is also used to temporarily put the first instruction in the sleep state, and then retransmit it to the program counter for execution when the resources are sufficient, ensuring the smooth execution of the first instruction.
[0216] In some embodiments, the instruction processing unit includes a load / store unit; the above method further includes:
[0217] When an internal exception occurs, the load / store unit ends the execution of the above instruction, and releases the first dependency and the second dependency.
[0218] In this embodiment, when an internal exception occurs, the load / store unit releases the first dependency and the second dependency simultaneously, which can implement the corresponding exception handling mechanism in the scenario where the same instruction supports two dependencies, ensuring the normal operation of the pipeline.
[0219] In some embodiments, the first type of register is the register corresponding to the control type of operand, and the second type of register is the register corresponding to the data type of operand.
[0220] On the other hand, the embodiments of the present application provide a graphics card, and the graphics card includes the processor described in each of the above embodiments. Optionally, the processor is a GPU.
[0221] On the other hand, the embodiments of the present application provide a computer device, and the computer device includes the processor described above. Optionally, the processor is a GPU. The computer device can be at least one of a portable computer, a desktop computer, a server, a server cluster, an artificial intelligence (AI) computing cluster, and a cloud computing cluster. Among them, the AI computing cluster can also be simply referred to as an intelligent computing cluster or a smart computing cluster.
[0222] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A processor, characterized in that, The processor includes: an instruction processing unit and a dependency processing unit, and at least two paths are provided between the instruction processing unit and the dependency processing unit, and the at least two paths include a first type of path and a second type of path; The instruction processing unit is configured to send a first dependency release signal to the dependency processing unit through the first type of path; the dependency processing unit is configured to release a first dependency relationship corresponding to a first type of register based on the first dependency release signal; The instruction processing unit is configured to send a second dependency release signal to the dependency processing unit through the second type of path; the dependency processing unit is configured to release a second dependency relationship corresponding to a second type of register based on the second dependency release signal; Wherein, both the first type of register and the second type of register are registers used by the same instruction; the first type of register is a register corresponding to a control type operand, and the second type of register is a register corresponding to a data type operand.
2. The processor according to claim 1, wherein The instruction processing unit includes a load / store unit; The first type of path includes: a first path between the load / store unit and the dependency processing unit; The second type of path includes: a second path between the load / store unit and the dependency processing unit.
3. The processor according to claim 2, wherein The instruction includes a store instruction; The load / store unit is configured to send the first dependency release signal to the dependency processing unit through the first path based on a first moment, and the first moment is the end moment of reading a first type of operand; The load / store unit is configured to send the second dependency release signal to the dependency processing unit through the second path based on a second moment, and the second moment is the end moment of issuing a memory write request; Wherein, the store instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and a memory access stage, the first moment is a moment in the execution stage, and the second moment is a moment in the memory access stage.
4. The processor according to claim 1, wherein The instruction processing unit includes a load / store unit and a write-back unit; The first type of path includes: a first path between the load / store unit and the dependency processing unit; The second type of path includes: a third path between the write-back unit and the dependency processing unit.
5. The processor according to claim 4, wherein The instruction includes a load instruction; The load / store unit is configured to send the first dependency release signal to the dependency processing unit through the first path based on a first moment, and the first moment is the end moment of reading a first type of operand; The write-back unit is configured to send the second dependency release signal to the dependency processing unit through the third path based on a third moment, and the third moment is the end moment of the write-back stage; Wherein, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and the write-back stage, and the first moment is a moment in the execution stage.
6. The processor according to claim 1, characterized in that, The instruction processing unit includes a load / store unit and a write-back unit; The first type of path includes: a second path between the load / store unit and the dependency processing unit; The second type of path includes: a third path between the write-back unit and the dependency processing unit.
7. The processor according to claim 6, wherein The instruction is a load instruction; The load / store unit is configured to send the first dependency release signal to the dependency processing unit through the second path based on a second moment, where the second moment is the end moment of issuing a memory read request; The write-back unit is configured to send the second dependency release signal to the dependency processing unit through the third path based on a third moment, where the third moment is the end moment of the write-back stage; Wherein, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes a memory access stage and the write-back stage, and the second moment is a moment in the memory access stage.
8. The processor according to any one of claims 1 to 7, characterized in that The dependency processing unit is configured to, based on the first dependency release signal, use a first dependency release mechanism to process and release the first dependency relationship corresponding to the first type of register; and based on the second dependency release signal, use a second dependency release mechanism to process and release the second dependency relationship corresponding to the second type of register.
9. The processor according to claim 8, wherein The first dependency release mechanism is a scoreboard mechanism, and the second dependency release mechanism is a sleep wake-up mechanism.
10. The processor according to any one of claims 1 to 7, characterized in that, The processor further includes a decoding unit; The decoding unit is configured to, in the decoding stage, respectively set the first dependency relationship and the second dependency relationship for the instruction.
11. The processor according to claim 10, characterized in that The decoding unit is further configured to, in the case where the first dependency relationship and the second dependency relationship cannot be set due to insufficient resources, wait until the retransmission condition is met and then retransmit the address of the instruction to the program counter, and the program counter is used to store the address of the next instruction to be executed by the processor.
12. The processor according to any one of claims 1 to 7, characterized in that, The instruction processing unit includes a load / store unit; The load / store unit is further configured to, in the case of an internal exception, end the execution of the instruction and release the first dependency relationship and the second dependency relationship.
13. A graphics card, characterized in that, The graphics card includes the processor according to any one of claims 1 to 12.
14. A computer device, characterized in that, The computer device includes the processor according to any one of claims 1 to 12.
15. A method for removing dependency, characterized in that: The method is executed by a processor, the processor includes: an instruction processing unit and a dependency processing unit, and at least two paths are provided between the instruction processing unit and the dependency processing unit, the at least two paths include a first type of path and a second type of path; the method includes: The instruction processing unit sends a first dependency release signal to the dependency processing unit through the first type of path; the dependency processing unit releases the first dependency relationship corresponding to the first type of register based on the first dependency release signal; The instruction processing unit sends a second dependency release signal to the dependency processing unit through the second type of path; the dependency processing unit releases the second dependency relationship corresponding to the second type of register based on the second dependency release signal; Among them, the first type of register and the second type of register are both registers used by the same instruction; the first type of register is the register corresponding to the control type operand, and the second type of register is the register corresponding to the data type operand.
16. The method according to claim 15, wherein The instruction processing unit includes a load / store unit; The first type of path includes: a first path between the load / store unit and the dependency processing unit; the second type of path includes: a second path between the load / store unit and the dependency processing unit; the instruction includes a store instruction; The instruction processing unit sends a first dependency release signal to the dependency processing unit through the first type of path, including: The load / store unit sends the first dependency release signal to the dependency processing unit through the first path based on a first moment, and the first moment is the end moment of reading the first type of operand; The instruction processing unit sends a second dependency release signal to the dependency processing unit through the second type of path, including: The load / store unit sends the second dependency release signal to the dependency processing unit through the second path based on a second moment, and the second moment is the end moment of issuing a memory write request; Among them, the store instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and a memory access stage, the first moment is a moment in the execution stage, and the second moment is a moment in the memory access stage.
17. The method according to claim 15, characterized in that, The instruction processing unit includes a load / store unit and a write-back unit; the first type of path includes: a first path between the load / store unit and the dependency processing unit; the second type of path includes: a third path between the write-back unit and the dependency processing unit; the instruction includes a load instruction; The instruction processing unit sends a first dependency release signal to the dependency processing unit through the first type of path, including: The load / store unit sends the first dependency release signal through the first path based on a first moment, and the first moment is the end moment of reading the first type of operand; The instruction processing unit sends a second dependency release signal to the dependency processing unit through the second type of path, including: The write-back unit sends the second dependency release signal to the dependency processing unit through the third path based on a third moment, and the third moment is the end moment of the write-back stage; Among them, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes an execution stage and the write-back stage, and the first moment is a moment in the execution stage.
18. The method according to claim 15, wherein The instruction processing unit includes a load / store unit and a write-back unit; the first type of path includes: a second path between the load / store unit and the dependency processing unit; the second type of path includes: a third path between the write-back unit and the dependency processing unit; the instruction is a load instruction; The instruction processing unit sends a first dependency release signal to the dependency processing unit through the first type of path, including: The load / store unit sends the first dependency release signal to the dependency processing unit based on a second moment through the second path, where the second moment is the end moment of issuing a memory read request; The instruction processing unit sends a second dependency release signal to the dependency processing unit through the second type of path, including: The write-back unit sends the second dependency release signal to the dependency processing unit based on a third moment through the third path, where the third moment is the end moment of the write-back stage; Wherein, the load instruction is executed using a multi-stage pipeline, the multi-stage pipeline includes a memory access stage and the write-back stage, and the second moment is a moment in the memory access stage.
19. The method according to any one of claims 15 to 18, characterized in that, The dependency processing unit releases a first dependency relationship corresponding to a first type of register based on the first dependency release signal, including: The dependency processing unit uses a first dependency release mechanism to process and release the first dependency relationship corresponding to the first type of register based on the first dependency release signal; The dependency processing unit releases a second dependency relationship corresponding to a second type of register based on the second dependency release signal, including: The dependency processing unit uses a second dependency release mechanism to process and release the second dependency relationship corresponding to the second type of register based on the second dependency release signal.
20. The method according to claim 19, wherein The first dependency release mechanism is a scoreboard mechanism, and the second dependency release mechanism is a sleep-wake mechanism.
21. The method according to any one of claims 15 to 18, characterized in that, The processor further includes a decoding unit; the method further includes: The decoding unit sets the first dependency relationship and the second dependency relationship for the instruction respectively in the decoding stage.
22. The method according to claim 21, wherein The method further includes: When the decoding unit is unable to set the first dependency relationship and the second dependency relationship due to insufficient resources, it waits until the retransmission condition is met and then retransmits the address of the instruction to the program counter, where the program counter is used to store the address of the next instruction to be executed by the processor.
23. The method according to any one of claims 15 to 18, characterized in that The instruction processing unit includes a load / store unit; the method further includes: When an internal exception occurs, the load / store unit ends the execution of the instruction and releases the first dependency relationship and the second dependency relationship.
Citation Information
Patent Citations
Apparatus and method for program order queue (POQ) to manage data dependencies in processor having multiple instruction queues
CN111752617A
Single Precision Vector Permute Immediate with "Word" Vector Write Mask
US20080100628A1