Instruction processing method, storage medium and electronic equipment
By introducing multiple pipelines into the processor and adjusting the order of instruction execution, the problem of idle bubbles in the superscalar processor due to instructions blockage is solved, and the execution efficiency of the processor is improved.
Patent Information
- Application Number
- CN202510058998.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
In the superscalar sequential execution processor, instructions are blocked in a certain stage due to data dependence or cache miss, resulting in multiple instructions being paused, affecting the efficiency of the processor's execution of instructions.
By introducing multiple pipelines into the processor and adjusting the execution order of instructions in the flow cycle, other instructions are allowed to continue to execute in the next flow stage, except for the paused instructions and their subsequent instructions in multiple instructions in the same flow stage.
This method reduces idle bubbles in the pipeline, improves the processor's execution efficiency, and ensures that instructions are executed in a preset order.
Smart Images

Figure CN119987866A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an instruction processing method, a storage medium and an electronic device. Background Art
[0002] With the continuous development of computer architecture, processor performance improvement increasingly depends on allowing more tasks to be performed simultaneously. Traditional processors use pipeline technology (Pipeline) in a single instruction stream to decompose instructions into multiple pipeline stages for execution, thereby improving the throughput of instructions. On this basis, superscalar technology introduces multi-issue pipelines, so that multiple instructions in the same cycle can flow through each pipeline stage in parallel, further improving processing performance. In superscalar sequential execution processors, the basic implementation method adopts the "lock-step" strategy (Lock-Step), that is, when there are multiple instructions in each pipeline stage at the same time, these instructions remain synchronized when entering and leaving the stage.
[0003] However, among multiple instructions running in lockstep, if one instruction is stalled (e.g., suspended) at a pipeline stage due to data dependency, cache miss, etc., the other instructions in the multiple instructions will also be suspended at the pipeline stage. After the one instruction resumes running, the multiple instructions will enter the next pipeline stage together, affecting the running of the other instructions in the multiple instructions except the one instruction, thus reducing the efficiency of the processor in executing instructions. Summary of the invention
[0004] In order to solve the problem of generating a large number of idle bubbles when instructions are executed by a superscalar in-order execution processor, the embodiments of the present application provide an instruction processing method, a storage medium and an electronic device, which are conducive to improving the efficiency of the processor in executing instructions.
[0005] In a first aspect, the present application provides an instruction processing method, which is applied to an electronic device, characterized in that a processor of the electronic device includes M pipelines, and the processor executes instructions including P pipeline stages, wherein M is an integer greater than 1, and P is an integer greater than 2; and the method includes:
[0006] In a first pipeline cycle, executing an i-th pipeline stage of P pipeline stages of the M first instructions through the M pipelines, where i is an integer greater than or equal to 1 and less than or equal to P;
[0007] In response to the kth first instruction executed in the kth pipeline being paused in the i-th pipeline stage in the second pipeline cycle, in the second pipeline cycle, the i+1-th pipeline stage of the 1st first instruction to the k-1th first instruction among the M first instructions is respectively executed through the 1st pipeline to the k-1th pipeline among the M pipelines, and the kth first instruction to the Mth first instruction among the M first instructions are respectively paused in the i-th pipeline stage of the kth pipeline to the Mth pipeline among the M pipelines, wherein the second pipeline cycle is the next pipeline cycle of the first pipeline cycle, and k is an integer greater than 1 and less than or equal to M.
[0008] In this way, during the parallel execution of M first instructions, after the kth first instruction in the same pipeline stage is paused in the i-th pipeline stage, the 1st to k-1th first instructions can continue to be executed in the i+1th pipeline stage, reducing the idle bubbles generated by the pause of the 1st to k-1th first instructions, thereby improving the execution efficiency of the processor.
[0009] In a possible implementation of the first aspect, in a first pipeline cycle, an i-1th pipeline stage of P pipeline stages of M second instructions is executed through M pipelines; in response to the kth first instruction executed on the kth pipeline in the second pipeline cycle being paused at the i-th pipeline stage, in the second pipeline cycle, the i-th pipeline stage of the first second instruction to the k-1th second instruction of the M second instructions is respectively executed through the first pipeline to the k-1th pipeline of the M pipelines;
[0010] And the kth second instruction to the Mth second instruction among the M second instructions are respectively paused in the i-1th pipeline stage of the kth pipeline to the Mth pipeline among the M pipelines.
[0011] In this way, when the kth first instruction cannot enter the i+1th pipeline stage in the second pipeline cycle, the 1st second instruction to the kth second instruction, the next instruction after the 1st first instruction to the kth first instruction in the second pipeline cycle, is allowed to enter the i-th pipeline stage, thereby reducing the idle bubbles generated by the pause of the 1st second instruction to the k-1th second instruction.
[0012] In a possible implementation of the first aspect above, the P pipeline stages of the M pipelines are respectively configured with control signals, and the values of the control signals are used to indicate the order of instructions in the corresponding pipeline stages. In the first pipeline cycle, the control signal of the i-th pipeline stage is a first value, and the first value indicates that the order of instructions executed in the i-th pipeline stage is a preset order.
[0013] When instructions are executed in a preset order, in the same pipeline stage, the execution order of the nth instruction on the M pipelines is before the execution order of the n+1th instruction, and the execution order of the nth instruction on the k-1th pipeline among the nth instructions is before the execution order of the nth instruction on the kth pipeline. Where n is an integer.
[0014] In this way, the preset order of instruction execution is determined by the configured control signal, ensuring that multiple instructions on the pipeline are executed in sequence.
[0015] In a possible implementation of the first aspect above, in the second pipeline cycle, in response to the kth first instruction executed by the kth pipeline being paused in the i-th pipeline stage, the control signal of the i-th pipeline stage is adjusted from a first value to a second value, and the second value indicates that the order of the kth first instruction in the i-th pipeline stage in the kth pipeline is before k-1 second instructions.
[0016] In this way, by changing the value of the control signal, it is ensured that for instructions in the same pipeline stage, the execution order of the kth first instruction is before the k-1 second instructions, avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0017] In a possible implementation of the first aspect, in response to the control signal of the i-th pipeline stage being a second value, when the k-th first instruction satisfies the requirement of executing the i+1-th pipeline stage in the third pipeline cycle, in the third pipeline cycle:
[0018] Execute the i+2th pipeline stage of the 1st first instruction to the k-1th first instruction among the M first instructions through the 1st pipeline to the k-1th pipeline among the M pipelines respectively;
[0019] Execute the i+1th pipeline stage of the kth first instruction to the Mth first instruction among the M first instructions respectively through the kth pipeline to the Mth pipeline among the M pipelines;
[0020] Execute the i+1th pipeline stage of the first second instruction to the k-1th second instruction of the M second instructions through the first pipeline to the k-1th pipeline of the M pipelines respectively;
[0021] The i-th pipeline stage of the k-th second instruction to the M-th second instruction among the M second instructions is executed respectively through the k-th pipeline to the M-th pipeline among the M pipelines.
[0022] In this way, an instruction on the pipeline is suspended in the i-th pipeline stage in the second pipeline cycle. After resuming execution in the third pipeline cycle, multiple instructions on the pipeline can enter the next pipeline stage in sequence according to the execution order of the second pipeline cycle.
[0023] In a possible implementation of the first aspect, in response to the control signal of the i-th pipeline stage being the second value, when the k-th first instruction is suspended from being executed at the i-th pipeline stage in the third pipeline cycle, in the third pipeline cycle:
[0024] Pause at the i+1th pipeline stage of the first instruction of the M first instructions to the first instruction of the k-1th through the first pipeline of the M pipelines respectively;
[0025] Pause at the i-th pipeline stage of the k-th first instruction to the M-th first instruction among the M first instructions respectively through the k-th pipeline to the M-th pipeline among the M pipelines;
[0026] Pause at the i-th pipeline stage of the first second instruction to the k-1-th second instruction among the M second instructions respectively through the first pipeline to the k-1-th pipeline among the M pipelines;
[0027] The kth pipeline to the Mth pipeline among the M pipelines are respectively paused at the i-1th pipeline stage of the kth second instruction to the Mth second instruction among the M second instructions.
[0028] In this way, when an instruction on the pipeline is suspended at the i-th pipeline stage in a certain pipeline cycle and is still suspended at the i-th pipeline stage in the next pipeline cycle, it is ensured that the instructions on multiple pipelines are executed in sequence, avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0029] In a possible implementation of the first aspect above, when P is 3, the P pipeline stages are, in sequence, an instruction fetch pipeline stage, a decoding pipeline stage, and an execution pipeline stage.
[0030] In a possible implementation of the first aspect above, when P is 5, the pipeline stages are the instruction fetch pipeline stage, the decoding pipeline stage, the execution pipeline stage, the memory access pipeline stage and the write back pipeline stage.
[0031] In a second aspect, the present application provides a processor, which executes the instruction processing method of any one of the first aspects.
[0032] In a third aspect, the present application provides an electronic device comprising the processor of the second aspect.
[0033] In a fourth aspect, the present application provides a readable storage medium having instructions stored thereon, and when the instructions are executed on an electronic device, the electronic device executes the instruction processing method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the accompanying drawings are described below.
[0035] Figure 1 According to some embodiments of the present application, a schematic diagram of an instruction execution process in an instruction processing method is shown;
[0036] Figure 2 According to some embodiments of the present application, a schematic diagram of an instruction execution process in an instruction processing method is shown;
[0037] Figure 3 According to some embodiments of the present application, a flowchart of an instruction processing method is shown;
[0038] Figure 4 According to some embodiments of the present application, a schematic diagram of an instruction execution process in an instruction processing method is shown;
[0039] Figure 5 According to some embodiments of the present application, a schematic diagram of an instruction execution process in an instruction processing method is shown;
[0040] Figure 6 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown. DETAILED DESCRIPTION
[0041] It should be noted that the method provided in the embodiment of the present application can be applied to any electronic device, including but not limited to a mobile station (MS), a mobile terminal (MT), etc. For example, the electronic device can be a mobile phone, a smart TV, a wearable device, a tablet computer (Pad), a desktop computer, a laptop computer, a virtual reality (VR) device, an augmented reality (AR) device, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical surgery, a terminal in a smart grid, a terminal in transportation safety, a terminal in a smart city, a terminal in a smart home, etc. The embodiment of the present application does not limit the specific form of the electronic device.
[0042] It should be noted that the instructions involved in the instruction execution method provided in the embodiment of the present application can be instructions in any computing task, including data processing, control flow management, computationally intensive operations, etc. The specific form of the instruction can include but is not limited to arithmetic operations (addition, subtraction, multiplication, division, etc.), logical operations (and, or, not, etc.), data load / store operations, conditional jump operations, bit operations, vector operations (such as dot multiplication, inner product, convolution, etc.), etc., to achieve complex computing tasks. These instructions support the processing of various data types, including integers, floating-point numbers, characters, strings, and more complex data structures (such as arrays and structures, etc.).
[0043] It should be noted that the data processed by the instructions involved in the embodiments of the present application may be image data, voice data, text data, etc., or other data, which is not limited here.
[0044] CPU instruction execution includes not only mathematical and logical operations, but also operations on memory data to ensure the correct transmission of data between registers and memory. Control flow instructions (such as conditional jumps and conditional no jumps) are used to manage program execution. In addition, the CPU also optimizes the execution efficiency of instructions through multi-core precision and quantitative technology, such as processing complex tasks such as large-scale data analysis, scientific computing, and multi-task batch processing.
[0045] During the instruction execution process, the CPU schedules different hardware resources (such as arithmetic logic unit (ALU), registers, cache, etc.) according to the operation type and execution conditions of the instruction, and based on the data dependency and control dependency between instructions, the CPU can process data in application scenarios such as single computing tasks and complex multi-threaded multi-tasking processing.
[0046] It will be understood that, as used herein, the term "module" may refer to or include an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and / or memory that executes one or more software or firmware programs, a combinational logic circuit, and / or other appropriate hardware components that provide the described functionality, or may be part of these hardware components.
[0047] It is understood that in each embodiment of the present application, the processor may be a microprocessor, a digital signal processor, a microcontroller, etc., and / or any combination thereof. According to another aspect, the processor may be a single-core processor, a multi-core processor, etc., and / or any combination thereof.
[0048] To facilitate understanding, the professional terms involved in the application are first explained.
[0049] (1) Pipeline
[0050] It is a technology used to improve the instruction execution efficiency of the processor. The execution process of an instruction can be divided into multiple independent, sequentially executed pipeline stages. Each pipeline stage uses different hardware resources to achieve parallel processing of multiple instructions. That is, at the same time (for example, the same clock cycle), the pipeline stages corresponding to multiple instructions are executed based on multiple different hardware resources.
[0051] Different pipeline stages in the pipeline are used to process a specific task of an instruction. The processing result of one pipeline stage can be used as the input of the next pipeline stage. The circuit used to implement all pipeline stages of an instruction in the processor can be called a pipeline.
[0052] The number of pipeline stages in a pipeline can be called the pipeline level. For example, based on the different pipeline stage data, the pipeline can have three-stage pipelines, five-stage pipelines, eight-stage pipelines, ten-stage pipelines, etc. Among them, the three-stage pipeline mainly includes three pipeline stages: instruction fetch (IF), decode (DE), and execute (EX); the five-stage pipeline mainly includes five pipeline stages: instruction fetch, decode, execute, memory access (MM), and write back (WB).
[0053] Taking a five-stage pipeline as an example, the following explains the content to be processed in each pipeline stage.
[0054] Instruction fetch pipeline stage: referred to as "instruction fetch", it is used to read the instructions to be executed from the instruction cache (or instruction queue) and save the instructions in the latch or register of the instruction fetch stage.
[0055] Decoding pipeline stage: referred to as "decoding", decodes the instruction to obtain the operand and operation code (used to indicate the type of operation corresponding to the instruction, such as shift, addition, subtraction, multiplication, etc.), and calculates or obtains the effective address of the operand.
[0056] Execution pipeline stage: referred to as "execution", it is used to execute different types of instructions through the corresponding hardware circuits, and write the execution results (such as calculation results, read data, etc.) into the destination register corresponding to the instruction. For example, if the instruction to be executed is an arithmetic and logic operation instruction, the corresponding operation is performed by the arithmetic and logic unit (ALU), and the operation result is written into the corresponding destination register. If the instruction to be executed is a load instruction, the address of the accessed memory is calculated by the load store unit (LSU), and data is read from the corresponding memory, and the data is written into the corresponding destination register. If the instruction to be executed is a storage instruction, the address of the accessed memory is calculated by the LSU, and the data is written into the corresponding memory.
[0057] Memory access pipeline stage: referred to as "memory access", used to read data from the memory or write data to the memory.
[0058] Write-back pipeline stage: abbreviated as "write-back", is used to write the instruction execution results (such as operation results, read data, etc.) into the destination register. In particular, the storage instruction has completed the writing of data in the memory access stage, and can be in an idle state or update status information in the write-back stage.
[0059] (2) Clock Cycle
[0060] It refers to the pulse time of the processor's internal clock. It is the shortest time unit required for the processor to perform a basic operation (such as reading a register, performing arithmetic and logical operations, etc.).
[0061] (3)Flow cycle
[0062] It refers to the time required for an instruction to pass through a pipeline stage in the pipeline. In a pipeline processor, the execution of instructions is broken down into multiple stages (such as instruction fetch, decoding, execution, memory access, write back, etc.), and each stage is completed within one pipeline cycle. Among them, a pipeline cycle can include one or more clock cycles, and the number of clock cycles included in the pipeline cycles corresponding to different pipeline stages can be different or the same.
[0063] As mentioned in the background technology, the processor improves the execution efficiency of instructions through pipeline technology. The pipeline decomposes the execution process of instructions into multiple pipeline stages (such as a five-stage pipeline including an instruction fetch pipeline stage, a decoding pipeline stage, an execution pipeline stage, a memory access pipeline stage, and a write-back pipeline stage), allowing multiple instructions to be processed in parallel at different pipeline stages, thereby improving the overall performance of the processor. However, in practical applications, due to data dependence, cache misses, etc., a certain instruction in lockstep execution may be hindered at a certain pipeline stage, which may cause the instructions on other pipelines to pause. Specifically, when a certain instruction cannot enter the next pipeline stage due to data dependence or cache misses at a certain pipeline stage, all pipelines may cause other instructions to wait and pause at this pipeline stage due to the blocking of the instruction, forming a pipeline idle bubble (Bubble). This is particularly prominent for multi-task and high-load data processing involving a large amount of data read and write operations, thereby affecting the continuous execution of instructions and reducing the execution efficiency of the processor.
[0064] The following describes the process of a processor executing instructions through a pipeline by taking a five-stage pipeline as an example in conjunction with the accompanying drawings.
[0065] For example, Figure 1 Figure 1 shows a schematic diagram of the execution process of two five-stage pipelines with multiple instructions executed in lockstep. Figure 1 As shown, the processor has two pipelines, pipeline 1 and pipeline 2, to perform data processing tasks. Between the 1st clock cycle and the 8th clock cycle, a total of five instructions, namely instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5, are executed. The execution of each instruction includes five pipeline stages: instruction fetch, decoding, execution, memory access, and write back. Instructions VL1, instruction VL3, and instruction VL5 are executed on pipeline 1, and instructions VL2 and instruction VL4 are executed on pipeline 2.
[0066] In the first clock cycle, instruction VL1 and instruction VL2 are in the instruction fetch pipeline stage;
[0067] In the second clock cycle, instructions VL1 and VL2 are in the decoding pipeline stage, and instructions VL3 and VL4 are in the instruction fetching pipeline stage. Instruction VL3 is the next instruction after instruction VL1, and instruction VL4 is the next instruction after instruction VL2.
[0068] In the third clock cycle, instruction VL1 is unable to enter the execution pipeline stage due to resource occupation (such as register resource occupation) and is forced to pause in the decoding pipeline stage. Instruction VL2 can enter the execution pipeline stage without waiting, but due to the lock-step execution of instruction VL1 and instruction VL2, instruction VL2 is also forced to pause with instruction VL1, that is, instruction VL1 and instruction VL2 need to pause for one clock cycle in the decoding pipeline stage. It can be understood that since instruction VL3 is the next instruction after instruction VL1 and instruction VL4 is the next instruction after instruction VL2, instruction VL3 and instruction VL4 also need to pause for one clock cycle in the instruction fetch pipeline stage.
[0069] In the fourth clock cycle, the resources of instruction VL1 are restored to idle (such as decoder resources are restored to idle), instruction VL1 is in the execution pipeline stage, instruction VL2 synchronizes instruction VL1 in the execution pipeline stage; at the same time, instruction VL3 and instruction VL4 are in the decoding pipeline stage, instruction VL5 is in the instruction fetch pipeline stage, where instruction VL5 is the next instruction after instruction VL3;
[0070] In the fifth clock cycle, instructions VL1 and VL2 are in the memory access pipeline stage, while instructions VL3 and VL4 are in the execution pipeline stage, and instruction VL5 is in the decoding pipeline stage;
[0071] In the sixth clock cycle, instructions VL1 and VL2 are in the write-back pipeline stage, while instructions VL3 and VL4 are in the memory access pipeline stage, and instruction VL5 is in the execution pipeline stage;
[0072] In the 7th clock cycle, instructions VL1 and VL2 are executed, while instructions VL3 and VL4 are in the write-back pipeline stage, and instruction VL5 is in the memory access pipeline stage;
[0073] In the 8th clock cycle, instruction VL3 and instruction VL4 are executed, and instruction VL5 is in the write-back pipeline stage;
[0074] In the above process, in the third clock cycle, due to the suspension of instruction VL1, instruction VL2, the next instruction VL3 of instruction VL1, the next instruction VL4 of instruction VL2, and the next instruction VL5 of instruction VL3, which do not need to wait in the same pipeline stage, are also suspended. Therefore, in the lock-step execution scheme between the above instructions, since instruction VL1 is suspended for one clock cycle, it takes a total of 8 clock cycles for instructions VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5 to complete execution.
[0075] It should be noted that, in some other embodiments, the time during which instruction VL1 is blocked may be one or more clock cycles. The more clock cycles during which instruction VL1 is paused, the more clock cycles that instruction VL2, the next instruction VL3 of instruction VL1, the next instruction VL4 of VL2, and the next instruction VL5 of VL3, which are executed in lockstep with instruction VL1, are delayed in execution. The paused execution of instructions in the pipeline causes a large number of bubbles to be generated in the pipeline, thereby reducing the efficiency of the processor in executing instructions.
[0076] In order to improve the efficiency of the processor in executing instructions, the present invention proposes an instruction processing method. When the processor runs instructions through multiple pipelines, if one of the multiple instructions in the same pipeline stage cannot enter the next pipeline stage in the next pipeline cycle due to data dependency, cache miss, etc., the processor can execute the next pipeline stage of at least one instruction in the execution order before the instruction in the multiple instructions in the next pipeline cycle. In addition, in the next pipeline cycle, the instruction and the instructions in the execution order after the instruction are suspended.
[0077] In this way, when the processor is executing multiple instructions in parallel, if a certain instruction cannot enter the next pipeline stage in the next pipeline cycle, except for the instruction and the instructions after the pipeline where the instruction is located, other instructions before the pipeline where the instruction is located can continue to execute and enter the next pipeline stage (reducing the idle bubbles of other instructions before the pipeline where the instruction is located), thereby improving the execution efficiency of the processor.
[0078] For example, for the aforementioned Figure 1 Instruction VL1 is paused for one clock cycle, which causes instructions VL2, VL3, VL4, and VL5 to be paused for one clock cycle, affecting the efficiency of the processor in executing multiple instructions. The instruction processing method provided in the embodiment of the present application can reduce the idle bubbles generated during the instruction execution process, which is conducive to improving the efficiency of the processor in executing instructions. For example, Figure 2 A schematic diagram of the execution process of a processor executing multiple instructions using the instruction processing method provided by an embodiment of the present application is shown.
[0079] like Figure 2As shown, the processor has two pipelines to perform data processing tasks. From the first clock cycle to the seventh clock cycle, there are five instructions, namely, instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5, which are executed. Each instruction execution includes five pipeline stages: instruction fetch, decoding, execution, memory access, and write back. Instructions VL1, instruction VL3, and instruction VL5 are executed on the first pipeline, and instructions VL2 and instruction VL4 are executed on the second pipeline. Each pipeline stage of the two pipelines is respectively configured with a control signal, and the value of the control signal is used to indicate the order of execution of the instructions in the corresponding pipeline stage. For example, when the value of the control signal is 0, it means that the instructions of pipeline 1 in this pipeline stage are executed first. When the value of the control signal is 1, it means that the instructions of pipeline 2 in this pipeline stage are executed first.
[0080] In the first clock cycle, instruction VL1 and instruction VL2 are in the instruction fetch pipeline stage, and the value of the instruction fetch pipeline stage control signal is 0.
[0081] In the second clock cycle, instructions VL1 and VL2 are in the decoding pipeline stage, and instructions VL3 and VL4 are in the instruction fetching pipeline stage. Instruction VL3 is the next instruction after instruction VL1, and instruction VL4 is the next instruction after instruction VL2. At this time, the value of the control signal of the instruction fetching pipeline stage is 0, and the value of the control signal of the decoding pipeline stage is 0.
[0082] In the third clock cycle, instruction VL2 is unable to enter the execution pipeline stage due to resource occupation (such as register resource occupation) and is forced to pause in the decoding pipeline stage. At this time, instruction VL1 is in the execution pipeline stage, while instruction VL3 is in the decoding pipeline stage and instruction VL5 is in the instruction fetch pipeline stage. Due to the pause of instruction VL2, instruction VL4 is also forced to pause in the instruction fetch pipeline stage; instruction VL5 is the next instruction after instruction VL3. At this time, the value of the instruction fetch pipeline stage control signal is 1, the value of the decoding pipeline stage control signal is 1, and the value of the execution pipeline stage control signal is 0.
[0083] In the 4th clock cycle, instruction VL1 is in the memory access pipeline stage, and instructions VL2 and VL3 are in the execution pipeline stage; at the same time, instructions VL4 and VL5 are in the decoding pipeline stage. At this time, the value of the decoding pipeline stage control signal is 1, the value of the execution pipeline stage control signal is 1, and the value of the memory access pipeline stage control signal is 0.
[0084] In the 5th clock cycle, instruction VL1 is in the write back pipeline stage, instructions VL2 and instruction VL3 are in the memory access pipeline stage, and at the same time, instructions VL4 and instruction VL5 are in the execution pipeline stage; at this time, the value of the execution pipeline stage control signal is 1, the value of the memory access pipeline stage control signal is 1, and the value of the write back pipeline stage control signal is 0.
[0085] In the 6th clock cycle, instruction VL1 is executed, instructions VL2 and VL3 are in the write back pipeline stage, and instructions VL4 and VL5 are in the memory access pipeline stage; at this time, the value of the memory access pipeline stage control signal is 1, and the value of the write back pipeline stage control signal is 1.
[0086] In the 7th clock cycle, instruction VL2 and instruction VL3 are executed, and at the same time, instruction VL4 and instruction VL5 are in the write-back pipeline stage. At this time, the value of the write-back pipeline stage control signal is 1.
[0087] During the above execution process, in the third clock cycle, since instruction VL2 is paused in the decoding pipeline stage, instruction VL1 in the same pipeline stage can enter the next pipeline stage (execution pipeline stage). At the same time, instruction VL3, the next instruction after instruction VL1, can also enter the same pipeline stage (decoding pipeline stage) as instruction VL2. Therefore, in the above instruction execution scheme, compared with the existing lock-step execution technical scheme, since one instruction pauses for one clock cycle, the processor needs a total of 7 clock cycles to complete the execution of instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5, which is 100% faster than the previous one. Figure 1 In the lock-step execution scheme shown, it takes a total of 8 clock cycles for the processor to execute instructions VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5. The instruction processing method provided in the embodiment of the present application saves one clock cycle.
[0088] The following introduces the technical solution of the present application by taking an example in which a processor includes M pipelines, each of which includes P pipeline stages.
[0089] Figure 3 According to some embodiments of the present application, a flowchart of an instruction processing method is shown, and the execution subject of the method is a processor, such as Figure 3 As shown, the method includes:
[0090] 101: In a first pipeline cycle, the i-th pipeline stage of M first instructions is executed through M pipelines.
[0091] In some embodiments, M is an integer greater than 1. For example, for a processor with two pipelines, M is 2.
[0092] In some embodiments, P is an integer greater than 2. For example, for a three-stage pipeline, P is 3, and for a five-stage pipeline, P is 5.
[0093] In some embodiments, i is an integer greater than or equal to 1 and less than or equal to P. For example, for a five-stage pipeline (mainly including five pipeline stages: instruction fetch, decoding, execution, memory access, and write back), i is 4, which represents the fourth pipeline stage (memory access pipeline stage) in the five-stage pipeline. For example, for a three-stage pipeline (mainly including three pipeline stages: instruction fetch, decoding, and execution), i is 3, which represents the third pipeline stage (execution pipeline stage) in the three-stage pipeline.
[0094] In some embodiments, in the first pipeline cycle, other pipeline stage processors before the i-th pipeline stage among the P pipeline stages may also execute other instructions, which are not limited here. For example, the i-1-th pipeline stage processor among the P pipeline stages may execute instruction A. For another example, the i-2-th pipeline stage processor among the P pipeline stages may execute instruction B.
[0095] In some embodiments, in the first pipeline cycle, other pipeline stage processors after the i-th pipeline stage among the P pipeline stages may also execute other instructions, which are not limited here. For example, the i+1-th pipeline stage processor among the P pipeline stages may execute the C instruction, and for another example, the i+2-th pipeline stage processor among the P pipeline stages may execute the D instruction.
[0096] For example, the first flow cycle can be the aforementioned Figure 2 In the second clock cycle, the M pipelines are as described above. Figure 2 There are two pipelines in (pipeline 1 and pipeline 2), the first instruction is the above Figure 2 Instruction VL1, instruction VL2, P pipeline stages as mentioned above Figure 2 There are five pipeline stages (including instruction fetch, decode, execute, access memory, and write back), the i-th pipeline stage is as mentioned above Figure 2 In the decoding pipeline stage. In the second clock cycle, instruction VL1 and instruction VL2 (M first instructions) are in the decoding pipeline stage (the i-th pipeline stage). Instruction VL3 and instruction VL4 are the next instruction (i.e., the second instruction) of instruction VL1 and instruction VL2 respectively, instruction VL5 is the next instruction (i.e., the third instruction) of instruction VL3, instruction VL1, instruction VL3, instruction VL5 are on the first pipeline, and instruction VL2 and instruction VL4 are on the second pipeline.
[0097] The M first instructions are executed in parallel in the i-th pipeline stage through multiple pipelines on the processor, which helps to improve the throughput of instructions in the processor.
[0098] In some embodiments, the P pipeline stages of the M pipelines may be respectively configured with control signals, and the values of the control signals are used to indicate the order in which instructions in the corresponding pipeline stages are executed.
[0099] With the aforementioned Figure 2 Taking the two pipelines as an example, when the value of the control signal is 0, it means that the instructions of pipeline 1 are executed first in this pipeline stage. When the value of the control signal is 1, it means that the instructions of pipeline 2 are executed first in this pipeline stage.
[0100] In this way, by configuring the P pipeline stages of the M pipelines with control signals respectively to indicate the execution order of instructions on multiple pipelines in the same pipeline stage, it is ensured that the instructions are executed in order.
[0101] In the first pipeline cycle, the control signal of the i-th pipeline stage is a first value, and the first value indicates that the order of instructions executed in the i-th pipeline stage is a preset order.
[0102] When instructions are executed in a preset order, in the same pipeline stage, the execution order of the nth instruction on the M pipelines on the processor is before the execution order of the n+1th instruction, and the execution order of the nth instruction on the k-1th pipeline of the nth instruction is before the execution order of the nth instruction on the kth pipeline, where n is an integer.
[0103] For example, for M = 2 (as mentioned above Figure 2 In the case of pipeline 1 and pipeline 2 in the pipeline, there are 5 instructions in total, namely instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5, among which instruction VL1 and instruction VL2 are the first instructions, instruction VL3 and instruction VL4 are the second instructions, and instruction VL5 is the third instruction. The preset execution order is that the first instruction (i.e. instruction VL1 and instruction VL2) takes precedence over the second instruction (i.e. instruction VL3 and instruction VL4), and the second instruction (i.e. instruction VL3 and instruction VL4) takes precedence over the third instruction (i.e. instruction VL5). Among them, instruction VL1, instruction VL3, and instruction VL5 are on pipeline 1, and instruction VL2 and instruction VL4 are on pipeline 2. Therefore, in the first instruction (i.e. instruction VL1 and instruction VL2), instruction VL1 on pipeline 1 takes precedence over instruction VL2 on pipeline 2. In the second instruction (i.e., instruction VL3 and instruction VL4), instruction VL3 on pipeline 1 is executed before instruction VL4 on pipeline 2. Therefore, the execution order of instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5 is instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5.
[0104] It can be understood that the control signal of the i-th pipeline stage is a first value, and the first value indicates that the order of instructions executed in the i-th pipeline stage is a preset order.
[0105] For example, for M = 2 (as mentioned above Figure 2 In the case of pipeline 1 and pipeline 2 in the processor, the control signal is a first value indicating that the order of instructions executed in the i-th pipeline stage in the processor is a preset order (such as pipeline 1 is executed first), and the control signal is a second value indicating that the order of instructions executed in the i-th pipeline stage is a non-preset order (such as pipeline 2 is executed first). The first value can be 0 and the second value can be 1; or the first value can be 1 and the second value can be 0.
[0106] In some embodiments, when M is other values (eg, 4, 8, 16, etc.), the value of the control signal may also be other values, and the number of bits of the control signal may also be other bits.
[0107] For example, when M is 4 (M pipelines include pipeline 1, pipeline 2, pipeline 3, and pipeline 4), that is, the processor contains four pipelines that execute instructions in parallel, the value of the control signal in the same pipeline stage can be 00, 01, 10, and 11.
[0108] Among them, 00 indicates that the order in which the processor executes instructions in the i-th pipeline stage is the preset order (ie, instruction VL1 on pipeline 1, instruction VL2 on pipeline 2, instruction VL3 on pipeline 3, instruction VL4 on pipeline 4).
[0109] 01 indicates that the order in which the processor executes instructions in the i-th pipeline stage is instruction VL2 on pipeline 2, instruction VL3 on pipeline 3, instruction VL4 on pipeline 4, and instruction VL1 on pipeline 1.
[0110] 10 indicates that the order in which the processor executes instructions in the i-th pipeline stage is instruction VL3 on pipeline 3, instruction VL4 on pipeline 4, instruction VL1 on pipeline 1, and instruction VL2 on pipeline 2.
[0111] 11 indicates that the order in which the processor executes instructions in the i-th pipeline stage is instruction VL4 on pipeline 4, instruction VL1 on pipeline 1, instruction VL2 on pipeline 2, and instruction VL3 on pipeline 3.
[0112] In addition, the value of the control signal of the same pipeline stage can also be expressed as other execution orders, which is not limited here. For example, 11 indicates that the order of instructions executed in the i-th pipeline stage is the preset order (that is, the execution order is instruction VL1 on pipeline 1, instruction VL2 on pipeline 2, instruction VL3 on pipeline 3, and instruction VL4 on pipeline 4).
[0113] In this way, by configuring the P pipeline stages of the M pipelines with control signals respectively, indicating the execution order of instructions on multiple pipelines in the same pipeline stage, it is ensured that the instructions are executed in order.
[0114] 102: In the second pipeline cycle, in response to the kth first instruction executed by the kth pipeline being paused at the i-th pipeline stage, the 1st pipeline to the k-1th pipeline respectively execute the i+1th pipeline stage of the 1st first instruction to the k-1th first instruction, and the kth first instruction to the Mth first instruction are respectively paused at the i-th pipeline stage of the kth pipeline to the Mth pipeline.
[0115] In some embodiments, the second pipeline cycle is the next pipeline cycle of the first pipeline cycle, and k is an integer greater than 1 and less than or equal to M.
[0116] In some embodiments, a control unit in the processor tracks the status of each instruction in multiple pipelines. In each clock cycle, the control unit can determine whether the register resources and data corresponding to each instruction are ready, and generate an indication signal whether the instruction can enter the next pipeline stage.
[0117] It can be understood that before the second pipeline cycle, the control unit in the processor detects that the register resources and data of the kth first instruction in the i+1th pipeline stage are not ready, but detects that the 1st first instruction to the k-1th first instruction are ready in the i+1th pipeline stage, that is, generates an indication signal indicating that the 1st first instruction to the k-1th first instruction can enter the i+1th pipeline stage. In the second pipeline cycle, the processor can execute the i+1th pipeline stage of the 1st first instruction to the k-1th first instruction through the 1st pipeline to the k-1th pipeline respectively.
[0118] For example, the second flow cycle can be the aforementioned Figure 2 In the third clock cycle, the first instruction is as mentioned above Figure 2Instructions VL1 and VL2 in the pipeline, in the second clock cycle, the control unit detects that the hardware resources (such as arithmetic logic unit resources, register files, etc.) and data required for instruction VL2 (the kth first instruction) on pipeline 2 in the execution pipeline stage (the i+1th pipeline stage) are not ready, but detects that the hardware resources and data required for instruction VL1 (the 1st first instruction to the k-1th first instruction) on pipeline 1 in the execution pipeline stage (the i+1th pipeline stage) are ready, and the control unit in the processor generates an indication signal for instruction VL1 on pipeline 1 to enter the execution pipeline stage (the i+1th pipeline stage). In the third clock cycle, the processor can execute the execution pipeline stage of instruction VL1.
[0119] In this way, after the kth first instruction in the same pipeline stage is paused in the i-th pipeline stage, the 1st to k-1th first instructions can continue to be executed in the i+1th pipeline stage, reducing the idle bubbles generated by the pause of the 1st to k-1th first instructions, thereby improving the execution efficiency of the processor.
[0120] At the same time, instruction VL2 (the first instruction of the kth pipeline) on pipeline 2 (as the kth pipeline) is suspended in the decoding pipeline stage (as suspended in the i-th pipeline stage) due to resource occupation (such as register resource occupation), pipeline 2 and the pipelines after pipeline 2 ( Figure 2 Only two pipelines are shown, not shown, as the first instruction ( Figure 2 Instruction VL2 in (VL2) is paused in the decoding pipeline stage (as the i-th pipeline stage).
[0121] In this way, after the kth first instruction in the same pipeline stage is paused in the ith pipeline stage, the first instructions on the kth pipeline to the Mth pipeline are also paused, avoiding multiple instructions on the pipeline from being executed out of order.
[0122] In some embodiments, it can be understood that in the first pipeline cycle, the i-1th pipeline stage of P pipeline stages of M second instructions is executed through M pipelines; in the second pipeline cycle, the kth first instruction executed by the kth pipeline is paused at the i-th pipeline stage, and in the second pipeline cycle, the i-th pipeline stage of the 1st second instruction to the k-1th second instruction is executed respectively through the 1st pipeline to the k-1th pipeline.
[0123] For example, the first flow cycle can be the aforementioned Figure 2 In the second clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2Instructions VL3 and VL4 in the second clock cycle. Instructions VL3 and VL4 (as M second instructions) are in the instruction fetch pipeline stage (as the i-1th pipeline stage). The second pipeline cycle is as mentioned above. Figure 2 In the third clock cycle, since instruction VL2 (the kth first instruction) is forced to pause in the decoding pipeline stage (as a pause in the i-th pipeline stage) due to resource occupation (such as register resources being occupied), in the third clock cycle, the processor can execute instruction VL3 (as the 1st second instruction to the k-1th second instruction) in the decoding pipeline stage (as the i-th pipeline stage) through pipeline 1 (as the 1st pipeline to the k-1th pipeline).
[0124] In this way, the kth first instruction is paused in the i-th pipeline stage, the 1st first instruction to the kth first instruction enter the i+1th pipeline stage, and the next instruction after the 1st first instruction to the kth first instruction (i.e., the 1st second instruction to the kth second instruction) can enter the i-th pipeline stage, allowing the next instruction of the non-paused instruction to be introduced into the idle slot of this pipeline stage, reducing the idle bubbles generated by the pause of the 1st second instruction to the k-1th second instruction.
[0125] In some embodiments, it can be understood that in the first pipeline cycle, the i-1th pipeline stage of P pipeline stages of M second instructions are executed through M pipelines; in response to the kth first instruction executed on the kth pipeline in the second pipeline cycle being paused at the i-1th pipeline stage, in the second pipeline cycle, the kth second instruction to the Mth second instruction are respectively paused at the i-1th pipeline stages of the kth pipeline to the Mth pipeline.
[0126] For example, the first flow cycle can be the aforementioned Figure 2 In the second clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2 Instructions VL3 and VL4 in the second clock cycle. Instructions VL3 and VL4 (as M second instructions) are in the instruction fetch pipeline stage (as the i-1th pipeline stage). The second pipeline cycle is as mentioned above. Figure 2 In the third clock cycle, because instruction VL2 (the first instruction of the kth order) is forced to pause in the decoding pipeline stage (as a pause in the i-th pipeline stage) due to resource occupation (such as register resource occupation), in the third clock cycle, pipeline 2 and the pipeline after pipeline 2 ( Figure 2 Only 2 pipelines are shown, not shown, as the second instruction ( Figure 2Instruction VL4) in is paused in the instruction fetch pipeline stage (as the i-1th pipeline stage).
[0127] In this way, after the kth first instruction in the same pipeline stage is paused in the i-th pipeline stage, the first instruction on the kth pipeline to the Mth pipeline is suspended, and the next instruction of the first instruction on the kth pipeline to the Mth pipeline (the second instruction on the kth pipeline to the Mth pipeline) is also suspended, avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0128] In some embodiments, in the second pipeline cycle, in response to the kth first instruction executed by the kth pipeline being paused in the i-th pipeline stage, the control signal of the i-th pipeline stage is adjusted from a first value to a second value, and the second value is used to indicate that the execution order of the kth first instruction in the i-th pipeline stage in the kth pipeline is before k-1 second instructions.
[0129] It can be understood that the instructions executed in the processor are executed in a sequential execution processor, and the processor executes one by one from the first instruction to the last instruction according to the instruction sequence in the program code. Since the kth first instruction executed by the kth pipeline is suspended at the ith pipeline stage, and there are the 1st second instruction to the k-1th second instruction in the ith pipeline stage, therefore, the kth first instruction to the Mth first instruction need to be executed before the k-1th second instruction.
[0130] For example, the second flow cycle can be the aforementioned Figure 2 In the third clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2 Instructions VL3 and VL4 in the decoding pipeline stage. In the decoding pipeline stage of the third clock cycle, instruction VL2 on pipeline 2 (the kth first instruction) is paused in the decoding pipeline stage (as a pause in the i-th pipeline stage) due to resource occupation, and instruction VL3 on pipeline 1 (the k-1 second instruction) is in the decoding pipeline stage (as the i-th pipeline stage). At this time, the control signal of the decoding pipeline stage is adjusted from 0 (the first value) to 1 (the second value), that is, the execution order of instruction VL2 on pipeline 2 (the kth first instruction) is before instruction VL3 on pipeline 1 (k-1 second instruction).
[0131] For another example, the second flow cycle can be the aforementioned Figure 2In the third clock cycle, in the instruction fetch pipeline stage, instruction VL4 (the kth first instruction) on pipeline 2 is forced to pause in the instruction fetch pipeline stage (as a pause in the i-th pipeline stage), and at the same time, instruction VL5 (the k-1 second instructions) on pipeline 1 is in the instruction fetch pipeline stage (as the i-th pipeline stage). At this time, the control signal of the instruction fetch pipeline stage is adjusted from 0 (the first value) to 1 (the second value), that is, the execution order of instruction VL4 (the kth first instruction) on pipeline 2 is before instruction VL5 (k-1 second instructions) on pipeline 1.
[0132] For another example, when M=4, and the values of the control signal are 00, 01, 10, and 11 as mentioned above:
[0133] Corresponding to k=1, the first value of the control signal can be the aforementioned 00 (preset order), and the execution order of the instructions in the 4 pipelines is the first instruction of the 1st pipeline, the first instruction of the 2nd pipeline, the first instruction of the 3rd pipeline, and the first instruction of the 4th pipeline.
[0134] Corresponding to k=2, when the control signal is converted from the first value 00 to the second value 01, the execution order of the instructions in the four pipelines is the first instruction of the second pipeline, the first instruction of the third pipeline, the first instruction of the fourth pipeline, and the second instruction of the first pipeline.
[0135] Corresponding to k=3, the control signal is converted from the first value 00 to the second value 10, and the execution order of the instructions in the 4 pipelines is the first instruction of the 3rd pipeline, the first instruction of the 4th pipeline, the second instruction of the 1st pipeline, and the second instruction of the 2nd pipeline.
[0136] Corresponding to k=4, the control signal is converted from the first value 00 to the second value 11, and the execution order of the instructions in the 4 pipelines is the first instruction of the 4th pipeline, the second instruction of the 1st pipeline, the second instruction of the 2nd pipeline, and the second instruction of the 3rd pipeline.
[0137] In this way, it is ensured that for instructions in the same pipeline stage in the processor, the execution order of the kth first instruction is before the k-1th second instruction, thereby avoiding the situation where multiple instructions on the processor are not executed in order.
[0138] 103: In response to the situation in which the i+1th pipeline stage of the kth first instruction is executed in the third pipeline cycle, in the third pipeline cycle, the 1st pipeline to the k-1th pipeline respectively execute the i+2th pipeline stage of the 1st first instruction to the k-1th first instruction.
[0139] In some embodiments, the control unit in the processor tracks the status of each instruction in multiple pipelines. In each clock cycle, the control unit can determine whether the register resources and data corresponding to each instruction are ready, and generate an indication signal whether the instruction can enter the next pipeline stage. For example, the control unit generates an indication signal for entering the i+1th pipeline stage when it confirms that the register resources and data required for the i+1th pipeline stage are ready; the control unit generates an indication signal for pausing entry into the i+1th pipeline stage when it detects that the register resources and data required for the i+1th pipeline stage are not ready (for example, the register resources are occupied by other instructions).
[0140] It can be understood that before the third pipeline cycle, the control unit in the processor determines that the register resources and data of the i+1th pipeline stage of the kth first instruction are ready, and generates an indication signal indicating that the kth first instruction can enter the i+1th pipeline stage. In the third pipeline cycle, the processor can execute the i+1th pipeline stage of the kth first instruction through the kth pipeline. At the same time, before the third pipeline cycle, the control unit in the processor determines that the register resources and data of the i+2th pipeline stage of the 1st to kth first instructions are ready, and generates an indication signal indicating that the i+2th pipeline stage can be entered. In the third pipeline cycle, the processor can execute the i+2th pipeline stages of the 1st first instruction to the k-1th first instruction respectively through the 1st pipeline to the k-1th pipeline.
[0141] For example, the third flow cycle can be the aforementioned Figure 2 In the 4th clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2 in the pipeline, in the third clock cycle, the control unit confirms that the register resources and data required for instruction VL2 (the kth first instruction) on pipeline 2 are ready for the execution pipeline stage (the i+1th pipeline stage), and the control unit in the processor generates an indication signal for instruction VL2 to enter the execution pipeline stage. In the fourth clock cycle, the processor can execute the execution pipeline stage of instruction VL2. At the same time, the control unit confirms that the register resources and data required for instruction VL1 (the 1st first instruction to the k-1th first instruction) on pipeline 1 are ready for the memory access pipeline stage (the i+2th pipeline stage), and the control unit in the processor generates an indication signal for entering the memory access pipeline stage. In the fourth clock cycle, the processor can execute the memory access pipeline stage of instruction VL1.
[0142] In some embodiments, the i-th pipeline stage of the k-th first instruction is the P-1-th pipeline stage of the instruction, that is, the k-th first instruction is executed in the i+1-th pipeline stage, and in the second pipeline cycle, the k-th first instruction executed by the k-th pipeline is paused at the i-th pipeline stage. In response to the i+1-th pipeline stage of executing the k-th first instruction being satisfied in the third pipeline cycle, in the third pipeline cycle, the instructions in the waiting queues from the 1st pipeline to the kth pipeline respectively enter the first pipeline stage of the 1st pipeline to the kth pipeline according to the preset order of the instructions in the aforementioned 101.
[0143] For example, taking a three-stage assembly line as an example, Figure 4 The schematic diagram of instruction execution is shown, and the processor includes two pipelines (pipeline 1 and pipeline 2. Among them, instruction VL1 and instruction VL2 are the first instructions, instruction VL3 and instruction VL4 are the second instructions, instruction VL5 and instruction VL6 are the third instructions, and instruction VL7 and instruction VL8 are the fourth instructions. Among them, instruction VL1, instruction VL3, instruction VL5, and instruction VL7 are on pipeline 1, and instruction VL2, instruction VL4, instruction VL6, and instruction VL8 are on pipeline 2.
[0144] In the first clock cycle, instructions VL1 and VL2 are in the instruction fetch pipeline stage, instructions VL3, VL5, and VL7 are in the waiting sequence of pipeline 1, and instructions VL4, VL6, and VL8 are in the waiting sequence of pipeline 2;
[0145] In the second clock cycle, instructions VL1 and VL2 are in the decoding pipeline stage, instructions VL3 and VL4 are in the instruction fetching pipeline stage, instructions VL5 and VL7 are in the waiting sequence of pipeline 1, and instructions VL6 and VL8 are in the waiting sequence of pipeline 2;
[0146] In the third clock cycle, instruction VL2 is suspended in the decoding pipeline stage due to resource occupation (such as register resource occupation), instruction VL1 is in the execution pipeline stage, instruction VL3 is in the decoding pipeline stage, instructions VL4 and VL5 are in the instruction fetch pipeline stage, instruction VL7 is in the waiting sequence of pipeline 1, and instructions VL6 and VL8 are in the waiting sequence of pipeline 2;
[0147] In the 4th clock cycle (the third pipeline cycle), instruction VL1 is executed, instructions VL2 and VL3 are in the execution pipeline stage, instructions VL4 and VL5 are in the decoding pipeline stage, and instruction VL6 is in the instruction fetch pipeline stage. Since instruction VL1 (the kth first instruction) is executed, instruction VL7 in the waiting queue in pipeline 1 (the 1st to kth pipelines) enters the instruction fetch pipeline stage (the first pipeline stage) on pipeline 1, and instruction VL8 is in the waiting sequence of pipeline 2.
[0148] Thus, after the kth first instruction is executed in the i+1th pipeline stage, in the next pipeline cycle, the instructions in the waiting queues in the 1st to kth pipelines can enter the first pipeline stage of execution in a preset order.
[0149] It can be understood that after the kth first instruction executed on the kth pipeline is paused at the i-th pipeline stage in the second pipeline cycle, in the third pipeline cycle, the kth first instruction satisfies the execution of the i+1-th pipeline stage, and the kth pipeline to the Mth pipeline respectively execute the i+1-th pipeline stage of the kth first instruction to the Mth first instruction.
[0150] For example, the third flow cycle can be the aforementioned Figure 2 In the 4th clock cycle, the first instruction is as mentioned above Figure 2 Instruction VL1 and instruction VL2 in the 3rd clock cycle, instruction VL2 (the first instruction of the kth order) pauses in the decoding pipeline stage (as paused in the i-th pipeline stage). In the 4th clock cycle, instruction VL2 releases resource occupation and is in the execution pipeline stage (the i+1-th pipeline stage). Pipeline 2 and the pipelines after pipeline 2 ( Figure 2 Only two pipelines are shown, not shown, as instruction VL2 (the kth first instruction to the Mth first instruction) on the kth pipeline to the Mth pipeline is in the execution pipeline stage (the i+1th pipeline stage).
[0151] In this way, it is ensured that the kth first instruction to the Mth first instruction can enter the i+1th pipeline stage in parallel, thereby improving the execution efficiency of the processor.
[0152] It can be understood that after the kth first instruction executed on the kth pipeline is paused in the i-th pipeline stage in the second pipeline cycle, in response to the kth first instruction satisfying the requirement to execute the i+1-th pipeline stage in the third pipeline cycle, in the third pipeline cycle: the i+1-th pipeline stage of the 1st second instruction to the k-1th second instruction is executed respectively through the 1st pipeline to the k-1th pipeline.
[0153] For example, the third flow cycle can be the aforementioned Figure 2In the 4th clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2 Instructions VL3 and VL4, in the 3rd clock cycle, instruction VL2 (the kth first instruction) is paused in the decoding pipeline stage due to resource occupation (as paused in the i-th pipeline stage). In the 4th clock cycle, instruction VL2 is in the execution pipeline stage (the i+1th pipeline stage) due to resource occupation release. In the 4th clock cycle, pipeline 1 (the 1st pipeline to the k-1th pipeline) executes the execution pipeline stage (the i+1th pipeline stage) of instruction VL3 (the 1st second instruction to the k-1th second instruction) due to resource occupation.
[0154] In this way, after the unpaused instructions (the 1st first instruction to the k-1th first instruction) enter the i+2th pipeline stage, the next instruction of the unpaused instructions (the 1st second instruction to the k-1th second instruction) enters the free slot of the current pipeline stage (i.e., the i+1th pipeline stage), further improving the efficiency of the processor in executing instructions.
[0155] It can be understood that after the kth first instruction executed in the kth pipeline is paused in the i-th pipeline stage in the second pipeline cycle, in response to the kth first instruction satisfying the requirement to execute the i+1-th pipeline stage in the third pipeline cycle, in the third pipeline cycle, the i-th pipeline stages of the kth second instruction to the Mth second instruction are respectively executed through the kth pipeline to the Mth pipeline.
[0156] For example, the third flow cycle can be the aforementioned Figure 2 In the 4th clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2 In the instructions VL3 and VL4, in the third clock cycle, instruction VL2 (the first instruction of the kth order) is paused in the decoding pipeline stage (paused in the i-th pipeline stage) due to resource occupation. In the fourth clock cycle, instruction VL2 is in the execution pipeline stage (the i+1-th pipeline stage) due to resource occupation release. At this time, pipeline 2 and the pipelines after pipeline 2 ( Figure 2 Only 2 pipelines are shown, not shown, as the k-th pipeline to the M-th pipeline) executing the decoding pipeline stage (the i-th pipeline stage) of instruction VL4 (the k-th second instruction to the M-th second instruction).
[0157] In this way, the kth first instruction in the pipeline is suspended in the i-th pipeline stage in the second pipeline cycle, and after resuming execution and entering the i+1-th pipeline stage in the third pipeline cycle, the next instruction of the kth first instruction enters the i+1-th pipeline stage, further improving the efficiency of the processor in executing instructions.
[0158] It can be understood that in the third pipeline cycle, there are the first instruction and the second instruction in the i-th pipeline stage. At this time, the execution order of the k-th first instruction in the i-th pipeline stage is before the k-1 second instructions.
[0159] For example, the third flow cycle can be the aforementioned Figure 2 In the 4th clock cycle, the first instruction is as mentioned above Figure 2 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 2 Instructions VL3 and VL4 in the pipeline are executed in the 4th clock cycle. Instruction VL2 on pipeline 2 (the kth first instruction) is in the execution pipeline stage (as paused in the i+1th pipeline stage), and instruction VL3 on pipeline 1 (the k-1 second instruction) is in the execution pipeline stage (as the i+1th pipeline stage). At this time, the control signal of the execution pipeline stage is 1 (the second numerical value), that is, the execution order of instruction VL2 (the kth first instruction) on processor pipeline 2 is before instruction VL3 (k-1 second instruction) on pipeline 1.
[0160] For another example, the third flow cycle can be as described above. Figure 2 In the 4th clock cycle, in the decoding pipeline stage, instruction VL4 (the kth first instruction) on pipeline 2 is in the decoding pipeline stage (as the ith pipeline stage), and at the same time, instruction VL5 (the k-1 second instructions) on pipeline 1 is in the decoding pipeline stage (as the ith pipeline stage). At this time, the control signal of the decoding pipeline stage is 1 (the second numerical value), that is, the execution order of instruction VL4 (the kth first instruction) on processor pipeline 2 is before instruction VL5 (k-1 second instructions) on pipeline 1.
[0161] In this way, it is ensured that for instructions in the same pipeline stage, the execution order of the kth first instruction is before the k-1th second instruction, thereby avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0162] 104: In response to the situation that the i+1th pipeline stage of the kth first instruction is not satisfied for execution in the third pipeline cycle, in the third pipeline cycle, the 1st pipeline to the k-1th pipeline are respectively paused at the i+1th pipeline stage of the 1st first instruction to the k-1th first instruction.
[0163] In some embodiments, the control unit in the processor tracks the status of each instruction in multiple pipelines. In each clock cycle, the control unit can determine whether the register resources and data corresponding to each instruction are ready, and generate an indication signal whether the instruction can enter the next pipeline stage. The control unit generates an indication signal for entering the i+1th pipeline stage when it confirms that the register resources and data required for the i+1th pipeline stage are ready; the control unit generates an indication signal for pausing to enter the i+1th pipeline stage when it detects that the register resources and data required for the i+1th pipeline stage are not ready (for example, the register resources are occupied by other instructions). In every other clock cycle, the control unit detects the status of the hardware resources and data, and determines whether the hardware resources (such as the register resources are released) and data detected by the control unit are ready.
[0164] It can be understood that before the third pipeline cycle, the control unit in the processor determines that the register resources and data of the i+1th pipeline stage of the kth first instruction are not ready (for example, the register resources are occupied by other instructions), and the control unit generates an indication signal to suspend the execution of the kth first instruction to enter the i+1th pipeline stage. In the third pipeline cycle, the processor can suspend the kth first instruction at the ith pipeline stage through the kth pipeline.
[0165] For example, the third flow cycle can be Figure 5 In the 4th clock cycle in the processor, there are two pipelines to perform data processing tasks. In the 1st to 8th clock cycles, there are five instructions, instruction VL1, instruction VL2, instruction VL3, instruction VL4, and instruction VL5, which are executed. Each instruction execution includes five pipeline stages: instruction fetch, decoding, execution, memory access, and write back. Instructions VL1, instruction VL3, and instruction VL5 are executed on the first pipeline, and instructions VL2 and instruction VL4 are executed on the second pipeline. Each pipeline stage of the two pipelines on the processor is respectively configured with a control signal, and the value of the control signal is used to indicate the order of instructions in the corresponding pipeline stage.
[0166] In the third clock cycle, the first instruction is the aforementioned Figure 5 Instructions VL1 and VL2, the second instruction is the aforementioned Figure 5Instructions VL3 and VL4 in the pipeline. In the third clock cycle, the control unit confirms that the hardware resources and data required by instruction VL2 (the kth first instruction) on pipeline 2 are not ready in the execution pipeline stage (the i+1th pipeline stage) (for example, the register resources are occupied by other instructions), and the control unit generates an indication signal to pause instruction VL2 from entering the execution pipeline stage. In the fourth clock cycle, instruction VL2 is paused in the decoding pipeline stage (paused in the i-th pipeline stage). Due to the sequential execution of instructions, instruction VL1 (the 1st first instruction to the k-1th first instruction) on processor pipeline 1 (the 1st pipeline to the k-1th pipeline) cannot execute the memory access pipeline stage (the i+2th pipeline stage) and is also paused in the execution pipeline stage (the i+1th pipeline stage).
[0167] In this way, when the kth first instruction is paused in the ith pipeline stage, the 1st to k-1th first instructions are also paused in the i+1th pipeline stage, avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0168] In some embodiments, in response to the situation that the i+1th pipeline stage of the kth first instruction is not satisfied in the third pipeline cycle, in the third pipeline cycle, the kth pipeline to the Mth pipeline are respectively paused at the i-th pipeline stage of the kth first instruction to the Mth first instruction.
[0169] For example, the third flow cycle can be the aforementioned Figure 5 In the 4th clock cycle, the first instruction is as mentioned above Figure 5 Instructions VL1 and VL2 in the 3rd clock cycle, instruction VL2 (the first instruction of the kth order) is paused in the decoding pipeline stage (paused in the i-th pipeline stage) due to resource occupation. In the 4th clock cycle, instruction VL2 is still in the decoding pipeline stage (the i-th pipeline stage) due to resource occupation. Pipeline 2 on the processor and the pipeline after pipeline 2 ( Figure 5 Only two pipelines are shown, not shown, as instructions VL2 (the kth first instruction to the Mth first instruction) on the kth pipeline to the Mth pipeline are also paused in the decoding pipeline stage (the i-th pipeline stage).
[0170] In this way, it is ensured that the kth first instruction to the Mth first instruction are respectively paused at the ith pipeline stage, thereby avoiding the situation where multiple instructions on the pipeline are not executed according to the execution order.
[0171] It can be understood that in response to the situation that the i+1th pipeline stage of the kth first instruction is not satisfied in the third pipeline cycle, in the third pipeline cycle, the processor pauses at the i-th pipeline stage of the 1st second instruction to the k-1th second instruction through the 1st pipeline to the k-1th pipeline respectively.
[0172] For example, the third flow cycle can be the aforementioned Figure 5 In the 4th clock cycle, the first instruction is as mentioned above Figure 5 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 5 Instructions VL3 and VL4 in the 3rd clock cycle, instruction VL2 (the kth first instruction) is paused in the decoding pipeline stage (paused in the i-th pipeline stage) due to resource occupation (such as register resources are occupied), and pipeline 1 (the 1st pipeline to the k-1th pipeline) executes the execution pipeline stage (the i-th pipeline stage) of instruction VL3 (the 1st second instruction to the k-1th second instruction). In the 4th clock cycle, instruction VL2 is in the decoding pipeline stage (the i-th pipeline stage) because resource occupation has not been released (such as register resources have not been occupied), and pipeline 1 (the 1st pipeline to the k-1th pipeline) is paused in the execution pipeline stage (the i-th pipeline stage) of instruction VL3 (the 1st second instruction to the k-1th second instruction).
[0173] In this way, it is ensured that the first second instruction to the k-1th second instruction are respectively paused at the i-th pipeline stage, thereby avoiding the situation where multiple instructions on the pipeline are not executed according to the execution order.
[0174] It can be understood that in response to the i+1th pipeline stage of the kth first instruction not being executed in the third pipeline cycle, in the third pipeline cycle, the processor pauses at the i-1th pipeline stage of the kth second instruction to the Mth second instruction through the kth pipeline to the Mth pipeline respectively.
[0175] For example, the third flow cycle can be the aforementioned Figure 5 In the 4th clock cycle, the first instruction is as mentioned above Figure 5 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 5 Instructions VL3 and VL4 in the 3rd clock cycle, instruction VL2 (the kth first instruction) is paused in the decoding pipeline stage due to resource occupation (as paused in the i-th pipeline stage). In the 4th clock cycle, instruction VL2 is still in the decoding pipeline stage (the i-th pipeline stage) because resource occupation has not been released. Since the kth second instruction to the Mth second instruction is the next instruction of the kth first instruction to the Mth first instruction, pipeline 2 and the pipelines after pipeline 2 ( Figure 5Only two pipelines are shown, not shown, as the kth pipeline to the Mth pipeline) are paused in the instruction fetch pipeline stage (the i-1th pipeline stage) of instruction VL4 (the kth second instruction to the Mth second instruction).
[0176] In this way, the kth first instruction to the Mth second instruction are respectively paused at the i-1th pipeline stage, thereby avoiding the situation where multiple instructions on the pipeline are not executed according to the execution order.
[0177] It can be understood that in the third pipeline cycle, there are the first instruction and the second instruction in the i-th pipeline stage. At this time, the execution order of the k-th first instruction in the i-th pipeline stage is before the k-1 second instructions.
[0178] For example, the third flow cycle can be the aforementioned Figure 5 In the 4th clock cycle, the first instruction is as mentioned above Figure 5 Instructions VL1 and VL2, the second instruction is as mentioned above Figure 5 In the decoding pipeline stage of the 4th clock cycle, instruction VL2 on pipeline 2 (the kth first instruction) is in the decoding pipeline stage (as paused at the i+1th pipeline stage), and at the same time, instruction VL3 on pipeline 1 (the k-1 second instruction) is in the decoding pipeline stage (as the i+1th pipeline stage). At this time, the control signal of the decoding pipeline stage is 1 (the second numerical value), that is, the execution order of instruction VL2 on pipeline 2 (the kth first instruction) is before instruction VL3 on pipeline 1 (the k-1 second instruction).
[0179] For another example, the third flow cycle can be the aforementioned Figure 5 In the 4th clock cycle, in the instruction fetch pipeline stage, instruction VL4 (the kth first instruction) on pipeline 2 is in the instruction fetch pipeline stage (as the i-1th pipeline stage), and at the same time, instruction VL5 (the k-1 second instruction) on pipeline 1 is in the instruction fetch pipeline stage (as the i-1th pipeline stage). At this time, the control signal of the instruction fetch pipeline stage is 1 (the second numerical value), that is, the execution order of instruction VL4 (the kth first instruction) on pipeline 2 is before instruction VL5 (k-1 second instruction) on pipeline 1.
[0180] In this way, it is ensured that for instructions in the same pipeline stage, the execution order of the kth first instruction is before the k-1th second instruction, thereby avoiding the situation where multiple instructions on the pipeline are not executed in order.
[0181] In summary, in the above-mentioned instruction running scheme, compared with the existing lock-step execution technical scheme, in the same pipeline stage, when there are at least two instructions and one of the instructions is suspended, all instructions in the pipeline stage are not forced to pause synchronously, and the instruction in front of the pipeline where the suspended instruction is located in the same pipeline stage is allowed to enter the next stage first. At the same time, after the execution of the non-paused instruction is completed, the next instruction of the non-paused instruction is introduced into the idle slot of the current pipeline stage. In this way, in the process of parallel execution of multiple instructions, in the same pipeline stage, except for the instruction and the instructions after the pipeline where the instruction is located, other instructions before the pipeline where the instruction is located in the processor can continue to execute and enter the next pipeline stage, reducing the idle bubbles of other instructions before the pipeline where the instruction is located, thereby improving the execution efficiency of the processor.
[0182] To facilitate understanding of the technical solution of the embodiment of the present application, the electronic device 100 is taken as an example to illustrate the structure of an electronic device to which the instruction processing method provided in the embodiment of the present application is applicable.
[0183] For example, Figure 6 According to some embodiments of the present application, a structural schematic diagram of an electronic device 100 is shown.
[0184] like Figure 6 As shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output (I / O) device 105, and a system control logic 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104 and the input / output device 105. Among them:
[0185] The processor 101 can be used to control the electronic device 100 to execute the instruction processing method of the present application. The processor 101 may include one or more processing units, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor or a processing module or processing circuit of a programmable logic device (field programmable gate array, FPGA). The processor 101 may include one or more single-core or multi-core processors.
[0186] The processor 101 can be used to control the electronic device 100 to execute the instruction processing method of the present application. The processor 101 may include one or more processing units, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor, or a processing module or processing circuit of a programmable logic device (field programmable gate array, FPGA). The processor 101 may include one or more single-core or multi-core processors. In some embodiments, the processor 101 can be used to execute the instruction processing method in the embodiments of the present invention (for example, the aforementioned) through multiple pipelines. Figure 2 and Figure 4 Instructions VL1 to VL5).
[0187] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions.
[0188] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as a flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card, etc.
[0189] In particular, the system memory 102 and the non-volatile memory 103 may respectively include: a temporary copy and a permanent copy of the instruction 107. The instruction 107 may include: when executed by the processor 101, the electronic device 100 implements the instruction processing method provided by each embodiment of the present application.
[0190] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, thereby communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the electronic device 100, for example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the electronic device 100 may communicate with other devices through the communication interface 104.
[0191] The input / output device 105 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. A user may interact with the electronic device 100 through the input / output device 105 .
[0192] System control logic 106 may include any suitable interface controller to provide any suitable interface with other modules of electronic device 100. For example, in some embodiments, system control logic 106 may include one or more memory controllers to provide interfaces to system memory 102 and non-volatile memory 103.
[0193] In some embodiments, at least one of the processors 101 may be packaged together with the logic of one or more controllers for the system control logic 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated on the same chip with the logic of one or more controllers for the system control logic 106 to form a SoC.
[0194] Understandably, Figure 6 The structure of the electronic device 100 shown is only an example. In other embodiments, the electronic device 100 may include more or fewer components than shown, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0195] In some embodiments, the embodiments of the present application also provide a computer-readable storage medium, which stores at least one computer program instruction, at least one program, code set or instruction set. The at least one computer program instruction, at least one program, code set or instruction set is loaded and executed by the model training system to implement the instruction processing method provided by the above-mentioned method embodiments.
[0196] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.
[0197] Program code may be applied to input instructions to perform the functions described herein and to generate output information. The output information may be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0198] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.
[0199] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0200] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0201] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0202] It should be noted that in the examples and description of this patent, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprises one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0203] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it will be apparent to those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application.
Claims
1. A command processing method, applied to an electronic device, characterized in that: The processor of the electronic device includes M pipelines, and the processor executes instructions including P pipeline stages, wherein M is an integer greater than 1, and P is an integer greater than 2; and the method includes: In a first pipeline cycle, executing an i-th pipeline stage of the P pipeline stages of the M first instructions through the M pipelines, where i is an integer greater than or equal to 1 and less than or equal to P; In response to the kth first instruction executed in the kth pipeline being paused in the i-th pipeline stage in the second pipeline cycle, in the second pipeline cycle, the i+1-th pipeline stage of the 1st first instruction to the k-1th first instruction among the M first instructions is respectively executed through the 1st pipeline to the k-1th pipeline among the M pipelines, and the kth first instruction to the Mth first instruction among the M first instructions are respectively paused in the i-th pipeline stage of the kth pipeline to the Mth pipeline among the M pipelines, wherein the second pipeline cycle is the next pipeline cycle of the first pipeline cycle, and k is an integer greater than 1 and less than or equal to M.
2. The method according to claim 1, characterized in that The method further comprises: In the first pipeline cycle, executing the i-1th pipeline stage of the P pipeline stages of the M second instructions through the M pipelines; In response to the kth first instruction executed in the kth pipeline in the second pipeline cycle being paused in the i-th pipeline stage, in the second pipeline cycle, the i-th pipeline stage of the first second instruction to the k-1th second instruction of the M second instructions is respectively executed through the first pipeline to the k-1th pipeline of the M pipelines, And pausing the kth second instruction to the Mth second instruction among the M second instructions in the i-1th pipeline stage of the kth pipeline to the Mth pipeline among the M pipelines respectively.
3. The method according to claim 2, characterized in that The P pipeline stages of the M pipelines are respectively configured with control signals, and the values of the control signals are used to indicate the order of instructions in the corresponding pipeline stages. In the first pipeline cycle, the control signal of the i-th pipeline stage is a first value, and the first value indicates that the order of instructions executed in the i-th pipeline stage is a preset order.
4. The method according to claim 3, characterized in that The method further comprises, In the second pipeline cycle, in response to the kth first instruction executed by the kth pipeline being paused in the i-th pipeline stage, the control signal of the i-th pipeline stage is adjusted from the first value to a second value, and the second value indicates that the order of the kth first instruction in the kth pipeline in the i-th pipeline stage is before the k-1 second instructions.
5. The method according to claim 4, characterized in that The method further comprises: In response to the control signal of the i-th pipeline stage being the second value, when the k-th first instruction satisfies the requirement of executing the i+1-th pipeline stage in the third pipeline cycle, in the third pipeline cycle: Execute the i+2th pipeline stage of the 1st first instruction to the k-1th first instruction among the M first instructions respectively through the 1st pipeline to the k-1th pipeline among the M pipelines; Executing the i+1th pipeline stage of the kth first instruction to the Mth first instruction among the M first instructions respectively through the kth pipeline to the Mth pipeline among the M pipelines; Executing the i+1th pipeline stage of the first second instruction to the k-1th second instruction among the M second instructions respectively through the first pipeline to the k-1th pipeline among the M pipelines; The i-th pipeline stage of the k-th second instruction to the M-th second instruction among the M second instructions is respectively executed through the k-th pipeline to the M-th pipeline among the M pipelines.
6. The method according to claim 5, characterized in that The method further comprises, In response to the control signal of the i-th pipeline stage being the second value, when the k-th first instruction is suspended at the i-th pipeline stage in the third pipeline cycle, in the third pipeline cycle: Pause at the i+1th pipeline stage of the first first instruction to the k-1th first instruction among the M first instructions respectively through the first pipeline to the k-1th pipeline among the M pipelines; Pause at the i-th pipeline stage of the k-th first instruction to the M-th first instruction among the M first instructions respectively through the k-th pipeline to the M-th pipeline among the M pipelines; Pause at the i-th pipeline stage of the first second instruction to the k-1-th second instruction among the M second instructions respectively through the first pipeline to the k-1-th pipeline among the M pipelines; The kth pipeline to the Mth pipeline among the M pipelines are respectively paused at the i-1th pipeline stage of the kth second instruction to the Mth second instruction among the M second instructions.
7. The method according to claim 1, characterized in that When P is 3, the P pipeline stages are the instruction fetch pipeline stage, the decoding pipeline stage and the execution pipeline stage in sequence.
8. The method according to claim 1, characterized in that When P is 5, the pipeline stages are the instruction fetch pipeline stage, the decoding pipeline stage, the execution pipeline stage, the memory access pipeline stage and the write-back pipeline stage.
9. A processor, characterized in that: The processor is used to execute the instruction processing method according to any one of claims 1 to 8.
10. An electronic device comprising the processor according to claim 9.
11. A readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the instruction processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Instruction processing method, processor, electronic equipment and storage medium
CN120578426A