Processing method and processing device of instruction pipeline, electronic equipment and storage medium
By skipping the search register mapping table in the processor and using the data dependency chain to judge the release of physical registers, the problem of unreasonable hardware resource utilization is solved and the processor's instruction processing performance is improved.
Patent Information
- Application Number
- CN202510345602.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-01
Smart Images

Figure CN120234047A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a processing method and apparatus for an instruction pipeline, an electronic device, and a storage medium. Background Art
[0002] With the gradual improvement of the performance of modern processors, it is required that the processor quickly execute multiple instructions. However, the correct processing of multiple instructions is often limited by the hardware resources of the processor. Therefore, the reasonable utilization of the processor hardware resources is one of the important factors for improving the instruction processing performance of modern processors. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a processing method for an instruction pipeline. The processing method for the instruction pipeline includes: obtaining multiple instructions to be issued in the same operation cycle; for an object instruction among the multiple instructions, in the case of skipping the search for the old physical register corresponding to the destination architecture register of the object instruction in the register mapping table, in response to there being a data dependency chain between at least two old instructions of the object instruction, determining whether to release the old physical register corresponding to the destination architecture register of the object instruction according to whether the destination architecture register of the object instruction is the same as the destination architecture register of at least one of the at least two old instructions; where the old instruction is an instruction whose instruction sequence is prior to that of the object instruction among the multiple instructions, the data dependency chain indicates that there is a data transfer relationship between the architecture registers of the corresponding multiple instructions, and the old physical register corresponding to the destination architecture register of the object instruction is the physical register mapped by the destination architecture register of the object instruction before the object instruction.
[0004] For example, the processing method for the instruction pipeline provided in at least some embodiments of the present disclosure further includes: in response to there being a data dependency chain between at least two old instructions, determining that the new physical registers of the destination architecture registers of the at least two old instructions are the same, and the new physical registers of the destination architecture registers of the at least two old instructions are the physical registers currently mapped by the destination architecture registers of the at least two old instructions.
[0005] For example, in the processing method for the instruction pipeline provided in at least some embodiments of the present disclosure, determining whether to release the old physical register corresponding to the destination architecture register of the object instruction includes: in response to the object instruction having a write-after-write relationship with at least one old instruction in the data dependency chain and the object instruction not having a write-after-write relationship with other old instructions except the at least one old instruction in the data dependency chain, determining that the old physical register corresponding to the destination architecture register of the object instruction cannot be released.
[0006] For example, in the processing method of an instruction pipeline provided by at least some embodiments of the present disclosure, at least two old instructions include a first old instruction and a second old instruction, and the instruction order of the first old instruction precedes that of the second old instruction. In response to a data dependence chain existing between the at least two old instructions, it includes: in response to the first old instruction and the second old instruction having a read-after-write relationship and the second old instruction being a data transfer cancellation type instruction, determining that there is a data dependence chain between the first old instruction and the second old instruction.
[0007] For example, in the processing method of an instruction pipeline provided by at least some embodiments of the present disclosure, determining whether to release the old physical register corresponding to the destination architecture register of an object instruction includes: in response to the object instruction having a write-after-write relationship with the first old instruction and not having a write-after-write relationship with the second old instruction, determining that the old physical register corresponding to the destination architecture register of the object instruction cannot be released.
[0008] For example, in the processing method of an instruction pipeline provided by at least some embodiments of the present disclosure, determining whether to release the old physical register corresponding to the destination architecture register of an object instruction includes: in response to the object instruction having a write-after-write relationship with the second old instruction and not having a write-after-write relationship with the first old instruction, determining that the old physical register corresponding to the destination architecture register of the object instruction cannot be released.
[0009] For example, in the processing method of an instruction pipeline provided by at least some embodiments of the present disclosure, the at least two old instructions further include a third old instruction, and the instruction order of the third old instruction is after that of the second old instruction. Determining whether to release the old physical register corresponding to the destination architecture register of an object instruction includes: in response to the object instruction having a write-after-write relationship with the second old instruction, not having a write-after-write relationship with the third old instruction, and the third old instruction not having a write-after-write relationship with the first old instruction, determining that the old physical register corresponding to the destination architecture register of the object instruction cannot be released.
[0010] For example, in the processing method of an instruction pipeline provided by at least some embodiments of the present disclosure, it further includes: in response to determining that the old physical register corresponding to the destination architecture register of an object instruction cannot be released, allocating a new physical register for the destination architecture register of the object instruction.
[0011] At least one embodiment of the present disclosure further provides a processing device for an instruction pipeline. The processing device for the instruction pipeline includes an acquisition unit and a determination unit. The acquisition unit is configured to acquire multiple instructions to be issued in the same operation cycle. The determination unit is configured to, for an object instruction among the multiple instructions, in a case of skipping to find an old physical register corresponding to the destination architecture register of the object instruction in a register mapping table, in response to there being a data dependency chain between at least two old instructions of the object instruction, determine whether to release the old physical register corresponding to the destination architecture register of the object instruction according to whether the destination architecture register of the object instruction is the same as the destination architecture register of at least one of the at least two old instructions; where an old instruction is an instruction whose instruction order is prior to that of the object instruction among the multiple instructions, a data dependency chain indicates that there is a data transfer relationship between the architecture registers of the corresponding multiple instructions, and the old physical register corresponding to the destination architecture register of the object instruction is the physical register to which the destination architecture register of the object instruction was mapped before the object instruction.
[0012] At least one embodiment of the present disclosure further provides a processing device for an instruction pipeline. The processing device for the instruction pipeline includes at least one storage unit and at least one processing unit. The at least one storage unit is configured to store computer-executable instructions. The at least one processing unit is configured to execute the computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processing unit, the instruction pipeline processing method provided in any embodiment of the present disclosure is implemented.
[0013] At least some embodiments of the present disclosure further provide an electronic device including the processing device for an instruction pipeline provided in any embodiment of the present disclosure.
[0014] At least some embodiments of the present disclosure further provide a non-transitory storage medium that non-transitorily stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by at least one processor, the processing method for an instruction pipeline provided in any embodiment of the present disclosure is implemented. Description of the Drawings
[0015] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0016] Figure 1A Shows an exemplary instruction pipeline of a scalar processor;
[0017] Figure 1B Shows an exemplary instruction pipeline of a superscalar processor;
[0018] Figure 2AShows a schematic diagram of a pipeline of a processor core;
[0019] Figure 2B Shows a schematic diagram of a pipeline of a processor core;
[0020] Figure 3 Shows a schematic diagram of the old / new physical register mapping of architectural registers after out-of-order instruction execution;
[0021] Figure 4 Shows a flowchart of a processing method for an instruction pipeline provided by at least one embodiment of the present disclosure;
[0022] Figure 5 Shows a block schematic diagram of a processing device for an instruction pipeline provided by at least one embodiment of the present disclosure;
[0023] Figure 6 Shows a block schematic diagram of a processing device for an instruction pipeline provided by at least one embodiment of the present disclosure;
[0024] Figure 7 Shows a schematic diagram of the structure of an electronic device provided by at least one embodiment of the present disclosure; and
[0025] Figure 8 Shows a schematic diagram of a non-transitory storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0026] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0027] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are merely used to distinguish different components. Similarly, terms such as "a", "an" or "the" do not denote a limitation of quantity, but rather indicate the presence of at least one. Words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0028] The following describes this disclosure through several specific embodiments. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of known functions and known components are omitted. When any component of the embodiments of this disclosure appears in more than one drawing, the component is denoted by the same or similar reference numerals in each drawing.
[0029] The terms used in this disclosure are those general terms that are currently widely used in the art in consideration of the functions of this disclosure, but these terms may vary according to the intentions of those of ordinary skill in the art, precedents, or new technologies in the art. In addition, specific terms may be selected by the applicant, and in such cases, their detailed meanings will be described in the detailed description of this disclosure. Therefore, the terms used in the specification should not be construed as mere names, but rather based on the meanings of the terms and the overall description of this disclosure.
[0030] Flowcharts are used in this disclosure to illustrate the operations performed by the systems according to the embodiments of this application. It should be understood that the operations before or below do not necessarily have to be performed precisely in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0031] Figure 1AShows the instruction pipeline of an exemplary scalar processor. The instruction pipeline includes a five-stage pipeline. Each instruction can be issued in each clock cycle and executed within a fixed time (e.g., 5 clock cycles). The execution of each instruction is divided into 5 steps: the instruction fetch (IF) stage 1001, the register read (RD) stage (or decoding stage) 1002, the arithmetic / logical unit (ALU) stage (or execution stage) 1003, the memory access (MEM) stage 1004, and the write-back (WB) stage 1005. In the IF stage 1001, the specified instruction is fetched from the instruction cache. A part of the fetched specified instruction is used to specify the source registers available for executing the instruction. In the RD stage 1002, it is decoded and control logic is generated to fetch the contents of the specified source registers. According to the control logic, arithmetic or logical operations are performed in the ALU stage 1003 using the fetched contents. In the MEM stage 1004, the memory in the instruction-readable / writable data cache is executed. Finally, in the WB stage 1005, the value obtained by executing the instruction can be written back to a certain register.
[0032] As Figure 1A Shown in the conventional scalar pipeline, the average number of instructions executed per clock cycle is less than or equal to 1, that is, its instruction-level parallelism is less than or equal to 1. Superscalar refers to a method of executing multiple instructions in parallel in one cycle. A processor with increased instruction-level parallelism that can process multiple instructions in one cycle is called a superscalar processor. A superscalar processor adds additional resources on the basis of a normal scalar processor, creates multiple pipelines, and each pipeline executes the instructions assigned to it to achieve parallelization. Figure 1B Shows the instruction pipeline of an exemplary superscalar processor. Four instructions can be input in parallel at each stage in the pipeline. For example, instructions 1 - 4 are processed in parallel, instructions 5 - 8 are processed in parallel, and instructions 9 - 12 are processed in parallel. As Figure 1B Shown in the conventional scalar pipeline, the average number of instructions executed per clock cycle is greater than 1, that is, its instruction-level parallelism is greater than 1.
[0033] For example, a superscalar processor can further support out-of-order execution. Out-of-order execution refers to a technique in which the CPU allows multiple instructions to be sent to the corresponding circuit units for processing in an order different from the order specified by the program. Out-of-order execution involves many algorithms, and these algorithms are basically designed based on reservation stations. The core idea of the reservation station is to save the instructions after decoding according to their respective instruction types in their respective reservation stations. If all the operands of the instruction are ready, out-of-order issue can start.
[0034] AsFigure 1B The pipeline shown can execute a simple instruction sequence smoothly, with one instruction executed per clock cycle. However, the instruction sequences in a program are not always simple sequences. There are usually dependencies between instructions, which may lead to pipeline execution conflicts or even errors. Currently, the main factors affecting the pipeline can be divided into three categories: resource conflicts (structural dependencies), data conflicts (data dependencies), and control conflicts (control dependencies).
[0035] In a program, if two instructions access the same register or memory address, and at least one of these two instructions is a write instruction, then there is a data dependency between these two instructions. Data dependencies can be divided into three cases according to the order of read and write in the conflicting access: RAW (read after write), WAW (write after write), and WAR (write after read). For RAW (read after write), the subsequent instruction needs to use the data written by the previous instruction, which is also called a true dependency; for WAW (write after write), two instructions write to the same target address, which is also called an output dependency; for WAR (write after read), the subsequent instruction overwrites the target address read by the previous instruction, which is also called an anti-dependency. The possible pipeline conflicts of WAW (write after write) and WAR (write after read) can be resolved through register renaming technology.
[0036] Figure 2AA schematic diagram of a pipeline of a processor core is shown. The dashed lines with arrows in the figure represent the redirected instruction stream. As shown, the processor core of a single-core processor or a multi-core processor (such as a CPU core) improves the instruction-level parallelism through pipeline technology. Inside the processor core, there are multiple pipeline stages. For example, after the program counters from various sources are fed into the pipeline and the next program counter (PC) is selected through a multiplexer (Mux), the instruction corresponding to the program counter has to go through various stages of processing such as branch prediction, instruction fetch, instruction decode, instruction dispatch and rename, instruction execution, and instruction retire. Between each pipeline stage, waiting queues are set up as needed, and these queues are usually first-in-first-out (FIFO) queues. For example, after the branch prediction unit, there is a branch prediction (BP) FIFO queue to store the branch prediction results; after the instruction fetch unit, there is an instruction cache (IC) FIFO to cache the fetched instructions; after the instruction decode unit, there is a decode (DE) FIFO to cache the decoded instructions; after the instruction dispatch and rename unit, there is a retire (RT) FIFO to cache the instructions waiting for confirmation of retirement after execution. At the same time, the pipeline of the processor core also includes an instruction queue to cache the instructions waiting for the instruction execution unit to execute after instruction dispatch and rename. To support a high operating frequency, each pipeline stage may in turn contain multiple pipeline levels (clock cycles). Although each pipeline level performs limited operations, in this way each clock can be made the shortest, and the performance of the processor core is improved by increasing the operating frequency of the processor core. Each pipeline level can also further improve the performance of the processor core by accommodating more instructions (i.e., superscalar technology).
[0037] Figure 2B A schematic diagram of a pipeline of a processor core is shown, where the instruction fetch, decode, and rename stages are the early stages of instruction processing. In the rename stage, through the register mapping table (also called the register renaming mapping table), the architectural register finds the corresponding mapped physical register, and this process involves the available list of physical registers to ensure the effective allocation of registers.
[0038] Subsequently, the instructions enter the out-of-order dispatch stage, where they can be reordered to optimize the execution efficiency and interact with the physical register file to obtain or store data in preparation for the execution stage.
[0039] After out-of-order scheduling, the instructions are sent to the execution units for out-of-order execution. For example, the actual computational execution is carried out through the Arithmetic Logic Unit (ALU), Multiplier (MUL) or Divider (DIV). At this stage, the execution results of the instructions are calculated and stored in the physical register file.
[0040] The instructions that have completed execution enter the In-Order Retirement stage. Although the instructions are executed out-of-order during the execution stage, they are completed in the logical order of the program during the In-Order Retirement stage. This stage manages the status of physical registers through the physical register availability list to ensure that the status of physical registers is correctly updated when the instructions retire and to release the physical registers that are no longer in use.
[0041] Figure 3 A schematic diagram showing the old / new physical register mapping of architectural registers after out-of-order execution of instructions is shown. Figure 3 The left side shows the instruction queue to be issued, that is, multiple instructions to be issued in the instruction queue. In the figure, three instructions are taken as examples, namely Instr0, Instr1, and Instr2. Among these instructions, Instr0 is the oldest, and these instructions need to be processed for register renaming. Figure 3 On the right side is the old / new physical register mapping table of architectural registers after instruction execution. The physical register before the slash ( / ) is the mapped old physical register, and the physical register after the slash ( / ) is the new physical register. As shown in this old / new physical register mapping table: after Instr0 is executed, the architectural register a is mapped from having no old physical register to the new physical register P1, while the architectural register d remains mapped to the physical register P0 unchanged (the new physical register and the old physical register are the same). After Instr1 is executed, the architectural register a continues to be mapped to the physical register P1, while the architectural register d is updated from being mapped to the old physical register P0 to the new physical register P1. After Instr2 is executed, the architectural register a is updated from being mapped to the old physical register P1 to the new physical register P2, while d continues to be mapped to the physical register P1.
[0042] As Figure 3As shown, in the prior instruction Instr0, the destination architectural register a is renamed to the new physical register P1. In the subsequent instruction Instr1, the destination architectural register d is renamed to the new physical register P1. Since the destination architectural registers of instruction Instr0 and instruction Instr2 are the same, both being a. Therefore, for the processing of instruction Instr2, the old physical register corresponding to the destination architectural register a (i.e., the physical register mapped due to the register renaming process of the previous instruction a) is P1, and the physical register P1 is used by the destination architectural register d of the prior instruction Instr1. Thus, a register renaming process needs to be performed for the destination architectural register a of instruction Instr2. At this time, a physical register in the unallocated idle state can be selected from the physical register available list, such as physical register P2.
[0043] The above mapping process of physical registers ensures that even in the case of out-of-order execution, each instruction can correctly access the register values it needs, while avoiding data conflicts. When a register renaming process needs to be performed on an instruction, the processor can select an unoccupied physical register from the physical register available list for mapping. The selected physical register will be removed from the physical register available list, and when the instruction execution is completed and committed, the originally occupied physical register can be released and added back to the physical register available list for subsequent instructions to use.
[0044] Therefore, during the out-of-order execution of instructions, determining when to release a mapped and occupied physical register is an issue that needs to be considered in register renaming.
[0045] During the out-of-order execution of some instructions, there are usually many register-to-register data transfer instructions in the instruction sequence, such as in scenarios involving parameter passing. In the register renaming stage, the destination architectural register of such a data transfer instruction can be mapped to the physical register that the source architectural register of this instruction already corresponds to, thus completing the function of the data transfer instruction in the renaming stage. This method cancels the execution stage and the read / write operand stage of the data transfer instruction, improving the performance of the processor core while saving power. This feature is called data transfer instruction cancellation, and such data transfer instructions can also be called "data transfer cancellation instructions", for example, the mov instruction.
[0046] However, the above register-to-register data transfer instruction cancellation feature also introduces a problem, that is, multiple architectural registers may be mapped to the same physical register, complicating the judgment of releasing the old physical register corresponding to the destination architectural register of this instruction.
[0047] For a situation where there is one or more instructions waiting to be issued in the same operation cycle, determining whether the old physical register corresponding to the destination architecture register of a certain instruction among these instructions can be released requires considering three aspects:
[0048] (1) In the register renaming mapping table, query whether the old physical register corresponding to the destination architecture register of the current instruction hits the new physical register corresponding to the destination architecture register of any instruction older than the current instruction. If it hits and there is no write-after-write (WAW) relationship between the current instruction and other prior instructions, then the old physical register corresponding to the destination architecture register of the current instruction cannot be released.
[0049] (2) In the register renaming mapping table, query whether the old physical register corresponding to the destination architecture register of each current instruction hits the physical registers corresponding to all architecture registers. If the query hits and the query hit is valid, then the old physical register corresponding to the destination architecture register of the current instruction cannot be released.
[0050] (3) If the current instruction is a data transfer cancellation instruction, then when the physical register mapped by its source architecture register is the same as the old physical register of the destination architecture register, it means that this old physical register will continue to be used, and then the old physical register of the destination architecture register of the current instruction cannot be released.
[0051] However, if the above methods are used together, there will be disadvantages such as increasing the usage of additional logic resources and wiring resources, reducing the circuit timing performance, and the more instructions waiting to be issued in the instruction queue, the more architecture registers need to be judged, which has a greater negative impact on the circuit area and timing.
[0052] The inventors of the present disclosure noticed that in the scenario where it is determined that the old physical register corresponding to the destination architecture register of the current instruction cannot be released, one is that it can be judged that the physical registers mapped by the architecture registers of two adjacent instructions must be the same through the WAW and RAW relationships between instructions, and the other is that it is necessary to query in the register renaming mapping table whether the old physical register corresponding to the destination architecture register of the current instruction hits the new physical registers corresponding to all architecture registers, and then, it is also necessary to judge whether the query result is valid through additional logic. If a hybrid processing judgment is to be performed to cover both scenarios, it will additionally increase the usage of pipeline logic resources and wiring resources.
[0053] At least one embodiment of the present disclosure provides a method for processing an instruction pipeline, which includes obtaining multiple instructions to be issued in the same operation cycle; for the target instruction among the multiple instructions, in the case of skipping the search for the old physical register corresponding to the target architecture register of the target instruction in the register mapping table, in response to there being a data dependency chain between at least two old instructions of the target instruction, determining whether to release the old physical register corresponding to the target architecture register of the target instruction according to whether the target architecture register of the target instruction is the same as the target architecture register of at least one of the at least two old instructions; wherein, the old instruction is an instruction whose instruction sequence is prior to the target instruction among the multiple instructions, the data dependency chain indicates that there is a data transfer relationship between the architecture registers of the corresponding multiple instructions, and the old physical register corresponding to the target architecture register of the target instruction is the physical register mapped by the target architecture register of the target instruction before the target instruction.
[0054] In the method for processing an instruction pipeline according to the above embodiment of the present disclosure, for the first aspect among the three aspects that need to be considered regarding whether the old physical register can be released, since multiple instructions to be issued in the same operation cycle can be used to determine whether to release the old physical register corresponding to the target architecture register of the target instruction according to the data dependency chain between at least two old instructions of the target instruction and whether the target architecture register of the target instruction is the same as the target architecture register of at least one of the at least two old instructions with a data dependency chain, it is possible to skip the search for the old physical register corresponding to the target architecture register of the target instruction in the register mapping table and determine whether to release the old physical register corresponding to the target architecture register of the target instruction, avoiding the increase in the number of register mapping tables and the resulting large selection logic caused by searching the register mapping table for each instruction, and also avoiding the increase in the usage of circuit logic and routing resources.
[0055] The method for processing an instruction pipeline according to the above embodiment simplifies the process of determining the release of the physical register corresponding to the architecture register, reduces the usage of relevant circuit logic and routing resources, and improves the timing performance of the circuit as the number of circuit logic levels decreases. Moreover, the above effect is more obvious when the number of instructions operating in the same operation cycle and the number of architecture registers are larger in the method for processing an instruction pipeline according to the above embodiment.
[0056] Next, each embodiment of the present disclosure will be described in conjunction with specific examples.
[0057] As Figure 4 shown, in some embodiments of the present disclosure, the method for processing an instruction pipeline includes step S30 and step S31.
[0058] In step S30, multiple instructions to be issued in the same operation cycle can be obtained.
[0059] In step S31, for the target instruction among multiple instructions, in the case of skipping the search for the old physical register corresponding to the target architecture register of the target instruction in the register mapping table, in response to a data dependence chain existing between at least two old instructions of the target instruction, it is determined whether to release the old physical register corresponding to the target architecture register of the target instruction according to whether the target architecture register of the target instruction is the same as the target architecture register of at least one of the at least two old instructions.
[0060] Here, the "target instruction" is the instruction regarded as the description object among multiple instructions. For example, it can be any instruction among multiple instructions that meets the requirements of the above processing method scenario; the old instruction is the instruction whose instruction sequence is prior to the target instruction among multiple instructions, the data dependence chain indicates that there is a data transfer relationship between the architecture registers of the corresponding multiple instructions, and the old physical register corresponding to the target architecture register of the target instruction is the physical register to which the target architecture register of the target instruction was mapped before the target instruction.
[0061] During the out-of-order execution of some instruction pipelines, the processor can issue multiple instructions in the same operation cycle. For example, the multiple instructions can be the instructions being executed in the instruction pipeline, and these instructions are also called "on-flight" instructions. For example, during the out-of-order execution in the instruction pipeline, the processor can perform out-of-order execution of multiple instructions in one operation cycle to improve the instruction throughput and the efficiency of the processor.
[0062] For example, the processor in the embodiments of the present disclosure can be any processor that supports out-of-order execution, such as a central processing unit (CPU), a microcontroller unit (MCU), or a digital signal processor (DSP), etc., and the present disclosure does not make any limitations.
[0063] For example, these processors can support a single-threaded working mode or a multi-threaded working mode, and the present disclosure does not make any limitations.
[0064] In the present disclosure, the operation cycle can be a clock cycle. For example, the state of the processor can be updated or basic logical operations can be performed within one clock cycle. Alternatively, the operation cycle can be a machine cycle or an instruction cycle. For example, it can be the machine cycle or instruction cycle required to execute a complete instruction. For example, one machine cycle or instruction cycle can include multiple clock cycles.
[0065] The register mapping table (hereinafter also simply referred to as the "mapping table") is used to store the mapping relationship between architectural registers and the corresponding allocated physical registers, and can be used during the register renaming process.
[0066] The "old physical register" po of a certain architectural register (such as the destination architectural register) Ar of an instruction Ic refers to the physical register to which the architectural register Ar was mapped in the register mapping table before the instruction Ic (when other instructions before it performed register renaming mapping). For example, it can be the physical register to which the destination architectural register of the target instruction was mapped before the target instruction. The "new physical register" pn of the architectural register Ar refers to the physical register to which the architectural register Ar is currently mapped when the instruction Ic performs register renaming mapping. For example, it can be the physical register to which the destination architectural register of the old instruction is currently mapped.
[0067] Moreover, when referring to a certain instruction a hitting or being regarded as hitting another instruction b in the content addressable memory (CAM) matching of the physical register in the register mapping table (also expressed as instruction a "CAM" instruction b hitting or being regarded as hitting), it means that the old physical register of the destination architectural register of a certain instruction a is the same as the new physical register of the destination architectural register of another instruction b. This matching result indicates that the old physical register of the destination architectural register of instruction a is still in use and cannot be released to the physical register available list.
[0068] For example, to further accelerate the processing speed of the instruction pipeline, step S31 can be executed in the decoding stage of the instruction pipeline.
[0069] In some embodiments of the present disclosure, the above instruction pipeline processing method further includes step S32.
[0070] In step S32, in response to at least two old instructions having a data dependency chain, it can be determined that the new physical registers of the destination architectural registers of the at least two old instructions are the same. The new physical registers of the destination architectural registers of the at least two old instructions are the physical registers to which the destination architectural registers of the at least two old instructions are currently mapped.
[0071] A data dependency chain indicates a data transfer relationship between architectural registers of at least two older instructions. For example, in the pipeline of a processor core that supports out-of-order execution (such as a single-core processor or a multi-core processor), multiple instructions can be issued in each cycle. For example, in the instruction queue cache, there are four exemplary instructions to be issued that have undergone register renaming processing:
[0072] Instruction instr0;
[0073] Instruction instr1;
[0074] Instruction instr2;
[0075] Instruction instr3;
[0076] The time of these four instructions decreases in sequence from instruction instr0 to instruction instr3, that is, from older to newer in terms of time.
[0077] For example, the object instruction is the current instruction. For example, if the object instruction is instruction instr2, it has at least two older instructions, namely instruction instr0 and instruction instr1.
[0078] For example, instruction instr0 is mov a,c, the source architectural register of instruction instr0 is c, and the destination architectural register is a; instruction instr1 is mov b,a, the source architectural register of instruction instr1 is a, and the destination architectural register is b. Therefore, the data transfer relationship between the architectural registers can be determined as c→a→b, and instruction instr0 and instruction instr1 have a data dependency chain.
[0079] For example, based on the fact that there is a data dependency chain between at least two older instructions (such as instruction instr0 and instruction instr1), it can be determined that the new physical registers of the destination architectural registers of at least two older instructions are the same.
[0080] For example, instruction instr0 and instruction instr1 have a data dependency chain, where instruction instr1 is a data transfer cancellation type. According to the characteristic of the data transfer cancellation type instruction in the register renaming process that the source architectural register and the destination architectural register point to the same physical register, the destination architectural register b of instruction instr1 and the source architectural register a of instruction instr1 point to the same physical register. Then, the new physical register of the destination architectural register a of instruction instr0 is the same as the new physical register of the destination architectural register b of instruction instr1.
[0081] In the above embodiments, according to the fact that there is a data dependence chain between at least two old instructions, it can be determined that the new physical registers of the destination architecture registers of at least two old instructions are the same, so that there is no need to search the register mapping table, and according to the data transfer relationship (i.e., the data dependence chain) between instructions, the relationship between the new physical registers corresponding to the destination architecture registers of at least two old instructions can be determined.
[0082] In some embodiments of the present disclosure, step S31 in the above method for processing an instruction pipeline further includes step S310.
[0083] In step S310, it can be determined that the old physical register corresponding to the destination architecture register of the object instruction cannot be released in response to the object instruction having a write-after-write relationship with at least one old instruction in the data dependence chain and the object instruction not having a write-after-write relationship with other old instructions except the at least one old instruction in the data dependence chain.
[0084] As described above, write-after-write (WAW) is a data correlation between instructions, which describes the situation where two consecutive instructions both perform write operations on the same architecture register. Among them, the second write operation needs to wait for the first write operation to complete before it can be executed to ensure data consistency and correctness.
[0085] If the object instruction has a write-after-write relationship with a certain old instruction, the old physical register of the destination architecture register of the object instruction must be the same as the new physical register of the destination architecture register of the old instruction, that is, CAM hit. However, this hit is an invalid hit for the first aspect among the three aspects to be considered for whether the old physical register can be released. However, in at least one example of the embodiments of the present disclosure, this write-after-write relationship can be utilized in combination with the data dependence chain to perform CAM hit judgment between the object instruction and other instructions, and further determine whether the old physical register of the destination architecture register of the object instruction can be released.
[0086] By determining the data correlation between the object instruction and at least two old instructions in the data dependence chain, it can be determined whether the old physical register corresponding to the destination architecture register of the object instruction is used by an old instruction whose instruction order is prior to the object instruction, so that the situation where the old physical register corresponding to the destination architecture register of the object instruction cannot be released can be determined without searching the register mapping table, simplifying the process of determining the release of the physical register corresponding to the architecture register, reducing the circuit logic and the occupied area of the circuit, and improving the timing performance of the circuit.
[0087] In some embodiments of the present disclosure, the at least two old instructions include a first old instruction and a second old instruction, and the instruction order of the first old instruction is prior to that of the second old instruction. Step S32 in the above method for processing an instruction pipeline further includes step S320.
[0088] In step S320, in response to the first old instruction and the second old instruction having a write-after-read relationship and the second old instruction being a data transfer cancellation type instruction, it can be determined that there is a data dependency chain between the first old instruction and the second old instruction.
[0089] As described above, the write-after-read relationship (RAW) is a type of data dependency that describes the dependency relationship where an instruction needs to read a corresponding architectural register after another instruction writes to it.
[0090] For example, the above at least two old instructions include a first old instruction (instruction instr0) and a second old instruction (instruction instr1). For example, if instruction instr0 and instruction instr1 have a RAW relationship (also denoted as raw0to1) and the instruction type of instruction instr1 is a data transfer cancellation type instruction (also denoted as move[1]), then it can be determined that there is a data dependency chain between the first old instruction and the second old instruction.
[0091] By determining the WAW relationship and RAW relationship between multiple instructions, it can be determined whether there is a data dependency chain among at least two old instructions of the target instruction, and further determine whether the old physical register corresponding to the destination architectural register of the target instruction can be released.
[0092] In some embodiments of the present disclosure, step S31 in the above method for processing an instruction pipeline further includes step S311.
[0093] In step S311, in response to the target instruction having a write-after-write relationship with the first old instruction and not having a write-after-write relationship with the second old instruction, it can be determined that the old physical register corresponding to the destination architectural register of the target instruction cannot be released.
[0094] For example, the object instruction is instruction instr2, and at least two old instructions, i.e., the first old instruction (instruction instr0) and the second old instruction (instruction instr1), have a data dependency chain. If instruction instr2 has a WAW relationship with instruction instr0 (also denoted as waw0to2) and instruction instr2 does not have a WAW relationship with instruction instr1 (also denoted as ~waw1to2), then it can be determined that the destination architecture registers of instruction instr2 and instruction instr0 are the same, but the destination architecture registers of instruction instr2 and instruction instr1 are different. And since the new physical registers of the destination architecture registers of at least two old instructions are the same, that is, the new physical registers of the destination architecture registers of instruction instr1 and instruction instr0 are the same, then instruction instr2 hits instruction instr1, the old physical register corresponding to the destination architecture register of instruction instr2 is the same as the new physical register corresponding to the destination architecture register of instruction instr1, and the old physical register corresponding to the destination architecture register of instruction instr2 is being used by instruction instr1. It can be determined that the old physical register corresponding to the destination architecture register of the object instruction (instruction instr2) cannot be released.
[0095] It should be noted that by determining that the object instruction (instruction instr2) does not have a WAW relationship with the second old instruction (instruction instr1), it can be determined that the destination architecture registers of the object instruction and the second old instruction are different. The hit between the object instruction and the second old instruction is not an invalid hit for the first aspect among the three aspects to be considered for whether the above-mentioned old physical register can be released (hereinafter, whether the hit is valid or invalid is based on the first aspect among the three aspects to be considered for whether the above-mentioned old physical register can be released).
[0096] For example, the object instruction (instruction instr2) has a WAW relationship with the first old instruction (instruction instr0) and the object instruction (instruction instr2) does not have a WAW relationship with the second old instruction (instruction instr1), and there is a data dependency chain between the first old instruction and the second old instruction (instruction instr0 and instruction instr1 have a RAW relationship and instr1 is a data transfer cancellation type instruction). At this time, the object instruction (instruction instr2) satisfies the hit formula waw0to2&~waw1to2&move[1]&raw0to1, then the old physical register of the destination architecture register of instruction instr2 is the same as the new physical register of the destination architecture register of instruction instr1 (the hit of instruction instr2 CAM instruction instr1 is valid).
[0097] For the specific scenario analysis that meets the hit formula waw0to2 & ~waw1to2 & move[1] & raw0to1, the following Table 1 can be referred to.
[0098] Table 1: Scenario Analysis of Instruction Type of Instruction instr0
[0099]
[0100] From Scenario 0 and Scenario 1 shown in Table 1, it can be known that in the case of meeting the hit formula, the old physical register P2 corresponding to the destination architecture register a of the object instruction (instruction instr2) in Scenario 0 is the same as the old physical register P2 corresponding to the destination architecture register b of the second-oldest instruction (instruction instr1), and the old physical register P4 corresponding to the destination architecture register a of the object instruction (instruction instr2) in Scenario 1 is the same as the old physical register P4 corresponding to the destination architecture register b of the second-oldest instruction (instruction instr1). Whether the object instruction (instruction instr2) hits the second-oldest instruction (instruction instr1) is valid and has nothing to do with the instruction type of the first-oldest instruction (instruction instr0).
[0101] It should be noted that in Scenario 0 and Scenario 1 of the above Table 1, whether the object instruction (instruction instr2) hits the second-oldest instruction (instruction instr1) has nothing to do with the instruction type of the object instruction. For example, the object instruction (instruction instr2) can also be a data transfer cancellation instruction mov a,d, and the present disclosure does not make any restrictions.
[0102] In some embodiments of the present disclosure, step S312 is further included in step S31 of the above instruction pipeline processing method.
[0103] In step S312, it can be determined that the old physical register corresponding to the destination architecture register of the object instruction cannot be released in response to the object instruction having a write-after-write relationship with the second-oldest instruction and the object instruction not having a write-after-write relationship with the first-oldest instruction.
[0104] For example, the object instruction is instruction instr2, and at least two old instructions, namely the first old instruction (instruction instr0) and the second old instruction (instruction instr1), have a data dependence chain. If instruction instr2 has a WAW relationship with instruction instr1 (also denoted as waw1to2) and instruction instr2 does not have a WAW relationship with instruction instr0 (also denoted as ~waw0to2), it can be determined that the destination architecture registers of instruction instr2 and instruction instr1 are the same, but the destination architecture registers of instruction instr2 and instruction instr0 are different. And since the new physical registers of the destination architecture registers of at least two old instructions are the same, that is, the new physical registers of the destination architecture registers of instruction instr1 and instruction instr0 are the same, then instruction instr2 hits instruction instr0. The old physical register corresponding to the destination architecture register of instruction instr2 is the same as the new physical register corresponding to the destination architecture register of instruction instr0, and the old physical register corresponding to the destination architecture register of instruction instr2 is being used by instruction instr0. It is determined that the old physical register corresponding to the destination architecture register of the object instruction (instruction instr2) cannot be released.
[0105] It should be noted that by determining that the object instruction (instruction instr2) does not have a WAW relationship with the first old instruction (instruction instr0), it can be determined that the destination architecture registers of the object instruction and the first old instruction are different, and the hit between the object instruction and the first old instruction is not an invalid hit.
[0106] For example, the object instruction (instruction instr2) has a WAW relationship with the second old instruction (instruction instr1) and the object instruction (instruction instr2) does not have a WAW relationship with the first old instruction (instruction instr0), and there is a data dependence chain between the first old instruction and the second old instruction (instruction instr0 and instruction instr1 have a RAW relationship and instr1 is a data transfer cancellation type instruction). At this time, the object instruction (instruction instr2) satisfies the hit formula waw1to2&~waw0to2&move[1]&raw0to1, then the old physical register of the destination architecture register of instruction instr2 is the same as the new physical register of the destination architecture register of instruction instr0 (the hit of instruction instr2 to instruction instr0 is valid).
[0107] For the specific scenario analysis of the hit formula waw1to2&~waw0to2&move[1]&raw0to1, please refer to Table 2 below.
[0108] Table 2: Scenario Analysis of Instruction Type of Instruction instr0
[0109]
[0110] As can be seen from Scenario 2 and Scenario 3 shown in Table 2, when the hit formula is satisfied, the old physical register P3 corresponding to the destination architecture register b of the object instruction (instruction instr2) in Scenario 2 is the same as the old physical register P3 corresponding to the destination architecture register a of the first old instruction (instruction instr0). The old physical register P2 corresponding to the destination architecture register b of the object instruction (instruction instr2) in Scenario 3 is the same as the old physical register P2 corresponding to the destination architecture register a of the first old instruction (instruction instr0). Whether the object instruction (instruction instr2) hits the first old instruction (instruction instr0) is valid and has nothing to do with the instruction type of the first old instruction (instruction instr0).
[0111] It should be noted that in Scenario 2 and Scenario 3 of Table 2 above, whether the object instruction (instruction instr2) hits the first old instruction (instruction instr0) also has nothing to do with the instruction type of the object instruction. For example, the object instruction (instruction instr2) can also be a data transfer cancellation instruction mov b,d, and the present disclosure does not make any restrictions.
[0112] In some embodiments of the present disclosure, at least two old instructions further include a third old instruction, and the instruction sequence of the third old instruction is after that of the second old instruction. Step S31 in the above method for processing the instruction pipeline further includes step S313.
[0113] In step S313, it can be determined that the old physical register corresponding to the destination architecture register of the object instruction cannot be released in response to the object instruction having a write-after-write relationship with the second old instruction, the object instruction not having a write-after-write relationship with the third old instruction, and the third old instruction not having a write-after-write relationship with the first old instruction.
[0114] For example, the object instruction is instruction instr3, the third old instruction can be instruction instr2 whose instruction sequence is after that of the second old instruction (instruction instr1) but before the object instruction, and the first old instruction is instruction instr0. As shown in the specific scenario analysis in Table 3.
[0115] Table 3: Hit Scenario Analysis of 4 Instructions
[0116]
[0117] In Scenario 4, there are two hit scenarios. One is that the third old instruction (instruction instr2) hits the first old instruction (instruction instr0). If the third old instruction in Scenario 4 is regarded as the object instruction, it is the same as the situation in Scenario 3 of the above embodiment and will not be elaborated here.
[0118] Another is that the object instruction (instruction instr3) in Scenario 4 hits the first old instruction (instruction instr0). For the object instruction (instruction instr3), the old instructions of instruction instr3 include instruction instr0, instruction instr1, and instruction instr2. Among them, instruction instr0 and instruction instr1 have a RAW relationship and instr1 is a data transfer cancellation type instruction, and instruction instr0 and instruction instr1 have a data dependency chain.
[0119] Based on the fact that instruction instr0 and instruction instr1 have a data dependency chain, according to the fact that the object instruction (instruction instr3) has a WAW relationship with the second old instruction (instruction instr1) (also denoted as waw1to3), the object instruction (instruction instr3) has no WAW relationship with the third old instruction (instruction instr2) (also denoted as ~waw2to3), and the third old instruction (instruction instr2) has no WAW relationship with the first old instruction (instruction instr0) (also denoted as ~waw0to2), it can be determined that the object instruction (instruction instr3) satisfies the hit formula ~waw0to2 & waw1to3 & ~waw2to3 & raw0to1 & move[1], and it can be determined that the object instruction (instruction instr3) hits the first old instruction (instruction instr0).
[0120] It should be noted that the Scenarios 0 to 4 shown in the above Tables 1 to 3 are only exemplary scenario analyses, and do not represent limitations on the instruction types and the number of instructions of multiple instructions. Based on the present disclosure, those skilled in the art can deduce and enumerate more instruction scenarios and instruction hit scenarios in the case of more instruction numbers through a limited number of experiments according to the technical solutions of the present disclosure.
[0121] Through the above at least one embodiment, it is possible to determine whether the old physical register corresponding to the destination architecture register of the object instruction is in use based on the data correlation between the object instruction and at least two old instructions having a data dependency chain, so as to determine whether the old physical register corresponding to the destination architecture register of the object instruction can be released in the case of skipping the search for the old physical register corresponding to the destination architecture register of the object instruction in the register mapping table.
[0122] It should be noted that the present disclosure is not intended to cover all hit scenarios, but rather to be able to determine whether the old physical register corresponding to the destination architecture register of the target instruction can be released in the case of skipping the search for the register mapping table for the hit scenarios that can be derived and exhausted through the technical solutions of the present disclosure and limited experiments based on the technical solutions of the present disclosure, so as to simplify the process of determining the release of the physical register corresponding to the architecture register and reduce the relevant circuit logic.
[0123] For example, when determining whether the target instruction CAM hits the old instruction, the scenarios that cannot be judged by the embodiments of the present disclosure can be grouped together and implemented by using the method of comparing the physical register numbers by searching the register mapping table; the scenarios that can be judged by the embodiments of the present disclosure can be grouped together and implemented by any of the above-mentioned implementation manners of the present disclosure.
[0124] In some embodiments of the present disclosure, the above method for processing the instruction pipeline further includes step S33.
[0125] In step S33, in response to determining that the old physical register corresponding to the destination architecture register of the target instruction cannot be released, a new physical register can be allocated for the destination architecture register of the target instruction.
[0126] By allocating a new physical register for the destination architecture register of the target instruction when the old physical register corresponding to the destination architecture register of the target instruction cannot be released, the pseudo-dependence caused by the same name of the physical register corresponding to the destination architecture register is eliminated, thereby improving the efficiency of the instruction pipeline.
[0127] At least one embodiment of the present disclosure further provides a processing device for an instruction pipeline. Figure 5 The block diagram of the processing device for the instruction pipeline provided by at least one embodiment of the present disclosure is shown. As Figure 5 shown, the processing device 300 for the instruction pipeline includes an acquisition unit 310 and a determination unit 320.
[0128] The acquisition unit 310 is configured to acquire multiple instructions to be issued in the same operation cycle.
[0129] The determination unit 320 is configured to, for the target instruction among the multiple instructions, in the case of skipping the search for the old physical register corresponding to the destination architecture register of the target instruction in the register mapping table, in response to there being a data dependence chain between at least two old instructions of the target instruction, determine whether to release the old physical register corresponding to the destination architecture register of the target instruction according to whether the destination architecture register of the target instruction is the same as the destination architecture register of at least one of the at least two old instructions.
[0130] The old instruction is an instruction whose instruction order is prior to that of the target instruction among multiple instructions. The data dependence chain indicates that there is a data transfer relationship between the architectural registers of the corresponding multiple instructions. The old physical register corresponding to the destination architectural register of the target instruction is the physical register to which the destination architectural register of the target instruction was mapped before the target instruction.
[0131] For example, the determination unit 320 is further configured to determine that the new physical registers of the destination architectural registers of at least two old instructions are the same in response to there being a data dependence chain between at least two old instructions, and the new physical registers of the destination architectural registers of at least two old instructions are the physical registers currently mapped by the destination architectural registers of at least two old instructions.
[0132] For example, the determination unit 320 is further configured to determine that the old physical register corresponding to the destination architectural register of the target instruction cannot be released in response to the target instruction having a write-after-write relationship with at least one old instruction in the data dependence chain and the target instruction not having a write-after-write relationship with other old instructions except the at least one old instruction in the data dependence chain.
[0133] In some embodiments of the present disclosure, the at least two old instructions include a first old instruction and a second old instruction, and the instruction order of the first old instruction is prior to that of the second old instruction. The determination unit 320 is further configured to determine that there is a data dependence chain between the first old instruction and the second old instruction in response to the first old instruction and the second old instruction having a read-after-write relationship and the second old instruction being a data transfer cancellation type instruction.
[0134] For example, the determination unit 320 is further configured to determine that the old physical register corresponding to the destination architectural register of the target instruction cannot be released in response to the target instruction having a write-after-write relationship with the first old instruction and the target instruction not having a write-after-write relationship with the second old instruction.
[0135] For example, the determination unit 320 is further configured to determine that the old physical register corresponding to the destination architectural register of the target instruction cannot be released in response to the target instruction having a write-after-write relationship with the second old instruction and the target instruction not having a write-after-write relationship with the first old instruction.
[0136] In some embodiments of the present disclosure, the at least two old instructions further include a third old instruction, and the instruction order of the third old instruction is posterior to that of the second old instruction. The determination unit 320 is further configured to determine that the old physical register corresponding to the destination architectural register of the target instruction cannot be released in response to the target instruction having a write-after-write relationship with the second old instruction, the target instruction not having a write-after-write relationship with the third old instruction, and the third old instruction not having a write-after-write relationship with the first old instruction.
[0137] In some embodiments of the present disclosure, the processing device 300 of the instruction pipeline further includes an allocation unit 330, which is configured to allocate a new physical register for the destination architectural register of the target instruction in response to determining that the old physical register corresponding to the destination architectural register of the target instruction cannot be released.
[0138] At least one embodiment of the present disclosure further provides a processing device for an instruction pipeline. Figure 6 FIG. shows a block diagram of a processing device for an instruction pipeline provided by at least one embodiment of the present disclosure. The processing device 800 of the instruction pipeline includes at least one storage unit 810 and at least one processing unit 820.
[0139] At least one storage unit 810 is configured to store computer-executable instructions.
[0140] At least one processing unit 820 is configured to execute the computer-executable instructions. When the computer-executable instructions are executed by the at least one processing unit, the instruction pipeline processing method provided by any embodiment of the present disclosure is implemented.
[0141] The technical effects of the processing device of the instruction pipeline in the above embodiments of the present disclosure are the same as those of the above instruction pipeline processing method, and thus will not be elaborated herein.
[0142] At least one embodiment of the present disclosure further provides an electronic device. The electronic device includes the processing device of the instruction pipeline described in the above at least one embodiment. For example, the electronic device may be a processor or a device including the processor. Embodiments of the present disclosure do not limit the specifications of the processor and the microarchitecture employed. For example, the processor may be a single-core processor or a multi-core processor, a single-threaded processor or a multi-threaded processor, and may adopt an X86 microarchitecture, a RISC-V microarchitecture, an ARM microarchitecture, a MIPS microarchitecture, etc.; in addition to the above processing device of the instruction pipeline, the processor may further include components for branch prediction, instruction decoding, instruction execution, one-level or more caches, etc. Embodiments of the present disclosure do not limit the specific structures, functions, and implementation manners of these components.
[0143] The technical effects of the electronic device in the above embodiments of the present disclosure are the same as those of the above processing device of the instruction pipeline, and thus will not be elaborated herein.
[0144] Figure 7 FIG. is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.
[0145] The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The illustrated electronic device 1000 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0146] For example, referring to Figure 7 , in some examples, the electronic device 1000 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. For example, the processing device 1001 includes a processing device of the instruction pipeline in any embodiment of the present disclosure, such as the above-exemplified processor. In the RAM 1003, various programs and data required for the operation of the computer system are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other through an interconnection network 1004. An input / output (I / O) interface 1005 is also connected to the interconnection network 1004.
[0147] For example, the following components may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; a communication device 1009 including, for example, a network interface card such as a LAN card, a modem, etc. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. A driver 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 1010 as needed so that a computer program read from it can be installed into the storage device 1008 as needed. Although Figure 7 the electronic device 1000 including various devices is illustrated, it should be understood that it is not required to implement or include all the illustrated devices. Instead, more or fewer devices may be implemented or included.
[0148] For example, the electronic device 1000 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 1009 may communicate with the network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0149] For example, the electronic device 1000 may be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, a server, etc., or may be a combination of an operating device and hardware of any instruction pipeline processing device. The embodiments of the present disclosure are not limited thereto.
[0150] At least one embodiment of the present disclosure also provides a non-transitory storage medium that non-transitorily stores computer-executable instructions. For example, when the computer-executable instructions are executed by a processor, the processing method of the instruction pipeline provided by at least one embodiment of the present disclosure is implemented.
[0151] Figure 8 is a schematic diagram of a non-transitory storage medium provided by some embodiments of the present disclosure. As Figure 8 shown, the non-transitory storage medium 900 may non-transitorily store computer-executable instructions 910, and the computer-executable instructions 910 implement the processing method of the instruction pipeline provided by any embodiment of the present disclosure when executed by a computer.
[0152] Regarding the present disclosure, the following points need to be noted:
[0153] (1) In the drawings of the embodiments of the present disclosure, only the structures related to the embodiments of the present disclosure are involved, and other structures may refer to the general design.
[0154] (2) Without conflict, the features in the same embodiment and different embodiments of the present disclosure may be combined with each other.
[0155] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A method for processing an instruction pipeline, comprising: Fetch multiple instructions to be issued in the same operation cycle; For a target instruction among the plurality of instructions, in a case where a search for an old physical register corresponding to a destination architectural register of the target instruction in a register mapping table is skipped, in response to at least two old instructions of the target instruction having a data dependency chain, determining whether to release the old physical register corresponding to the destination architectural register of the target instruction according to whether the destination architectural register of the target instruction is the same as the destination architectural register of at least one of the at least two old instructions; Among them, the old instruction is an instruction in the multiple instructions whose instruction order is earlier than the object instruction, the data dependency chain indicates that there is a data transfer relationship between the corresponding architecture registers of the multiple instructions, and the old physical register corresponding to the destination architecture register of the object instruction is the physical register to which the destination architecture register of the object instruction is mapped before the object instruction.
2. The processing method of the instruction pipeline according to claim 1, further comprising: In response to a data dependency chain between the at least two old instructions, it is determined that the new physical registers of the destination architectural registers of the at least two old instructions are the same, and the new physical registers of the destination architectural registers of the at least two old instructions are the physical registers currently mapped to the destination architectural registers of the at least two old instructions.
3. The processing method of the instruction pipeline according to claim 2, wherein: The determining whether to release the old physical register corresponding to the destination architecture register of the object instruction includes: In response to the object instruction having a write-after-write relationship with at least one old instruction in the data dependency chain and the object instruction not having a write-after-write relationship with other old instructions in the data dependency chain except the at least one old instruction, it is determined that the old physical register corresponding to the destination architectural register of the object instruction cannot be released.
4. The processing method of the instruction pipeline according to claim 2, wherein: The at least two old instructions include a first old instruction and a second old instruction, wherein the instruction sequence of the first old instruction precedes that of the second old instruction, The step of responding to the at least two old instructions having a data dependency chain comprises: In response to the first old instruction and the second old instruction having a read-after-write relationship and the second old instruction being a data transfer cancel type instruction, it is determined that there is a data dependency chain between the first old instruction and the second old instruction.
5. The processing method of the instruction pipeline according to claim 4, wherein: The determining whether to release the old physical register corresponding to the destination architecture register of the object instruction includes: In response to the object instruction and the first old instruction having a write-after-write relationship and the object instruction and the second old instruction not having a write-after-write relationship, it is determined that the old physical register corresponding to the destination architectural register of the object instruction cannot be released.
6. The processing method of the instruction pipeline according to claim 4, wherein: The determining whether to release the old physical register corresponding to the destination architecture register of the object instruction includes: In response to the object instruction and the second old instruction having a write-after-write relationship and the object instruction and the first old instruction not having a write-after-write relationship, it is determined that the old physical register corresponding to the destination architectural register of the object instruction cannot be released.
7. The processing method of the instruction pipeline according to claim 4, wherein: The at least two old instructions further include a third old instruction, the instruction sequence of the third old instruction is later than the second old instruction, and the determining whether to release the old physical register corresponding to the destination architecture register of the object instruction includes: In response to the object instruction and the second old instruction having a write-after-write relationship, the object instruction and the third old instruction not having a write-after-write relationship, and the third old instruction and the first old instruction not having a write-after-write relationship, it is determined that the old physical register corresponding to the destination architecture register of the object instruction cannot be released.
8. The processing method of the instruction pipeline according to any one of claims 1 to 7, further comprising: In response to determining that the old physical register corresponding to the destination architectural register of the target instruction cannot be released, a new physical register is allocated for the destination architectural register of the target instruction.
9. A processing device for an instruction pipeline, comprising: a fetch unit configured to fetch a plurality of instructions to be issued in the same operation cycle; a determining unit configured to, for a target instruction among the plurality of instructions, in a case where a search for an old physical register corresponding to a destination architectural register of the target instruction in a register mapping table is skipped, determine whether to release the old physical register corresponding to the destination architectural register of the target instruction in response to a data dependency chain between at least two old instructions of the target instruction and according to whether the destination architectural register of the target instruction is the same as the destination architectural register of at least one of the at least two old instructions; Among them, the old instruction is an instruction in the multiple instructions whose instruction order is earlier than the object instruction, the data dependency chain indicates that there is a data transfer relationship between the corresponding architecture registers of the multiple instructions, and the old physical register corresponding to the destination architecture register of the object instruction is the physical register to which the destination architecture register of the object instruction is mapped before the object instruction.
10. A processing device for an instruction pipeline, comprising: at least one storage unit configured to store computer executable instructions; as well as At least one processing unit is configured to execute the computer executable instructions, wherein the computer executable instructions, when executed by the at least one processing unit, implement the instruction pipeline processing method according to any one of claims 1-8.
11. An electronic device comprising the instruction pipeline processing device according to claim 9 or 10.
12. A non-transitory storage medium that non-transitorily stores computer-executable instructions, wherein: When the computer executable instructions are executed by at least one processor, the instruction pipeline processing method according to any one of claims 1-8 is implemented.
Citation Information
Cited By
Instruction processing method
CN120508321A
An instruction processing method
CN120508321B