Instruction Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium
By allocating a cache unit to the instructions after the instruction processing process is executed, and decoupling the instruction sorting and exit cache functions of the ROB module, the problem of waste of cache resources in the superscalar processor is solved, and the processor's instruction processing efficiency is improved.
Patent Information
- Application Number
- CN202210784294.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-28
AI Technical Summary
In existing hyperscalar processors, the reorder cache module (ROB) allocates cache resources in advance during instruction processing, resulting in waste of resources, reducing resource utilization.
The instruction is assigned to the instruction after the instruction processing process is completed, and the instruction sorting function and the exit instruction cache function are decoupled to reduce the cache resource usage and increase the number of pending instructions.
By reducing cache resource usage, the number of pending instructions on the instruction processing pipeline is improved, and the processor's instruction processing efficiency is improved.
Smart Images

Figure CN115080121B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies. Specifically, this application relates to an instruction processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, processor designs usually adopt a superscalar design method. A superscalar design processor supports multiple-issue out-of-order execution, that is, in a superscalar design processor, instructions have the characteristics of out-of-order execution and in-order retirement. Most superscalar design processors sort instructions through a Reorder Buffer (ROB) module to ensure the in-order retirement of instructions. However, in the ROB design method, ROB cache resources are usually allocated during the instruction processing process, resulting in waste of cache resources. Summary of the Invention
[0003] The objective of this application aims to solve at least one of the above technical defects, especially the technical defect of waste of cache resources.
[0004] According to one aspect of this application, an instruction processing method is provided. The method includes:
[0005] Receiving first indication information; wherein, the first indication information indicates that the instruction processing process corresponding to a first instruction in a processor is completed;
[0006] In response to the first indication information, allocating a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the instruction generation order.
[0007] Optionally, after allocating the target cache unit for the first instruction in the first cache, the method further includes:
[0008] Storing the first instruction into the target cache unit.
[0009] Optionally, the method further includes:
[0010] If a second instruction generated before the first instruction is stored in the first cache, deleting the first instruction and the second instruction in the instruction generation order.
[0011] Optionally, after receiving an instruction generation message, the method further includes:
[0012] Adding an identity identifier for the first instruction, where the identity identifier indicates the instruction generation order of the first instruction.
[0013] Optionally, allocating the target cache unit for the first instruction in the first cache includes:
[0014] Determine the target cache unit corresponding to the identity identifier according to the correspondence between the identity identifier and the cache unit identifier
[0015] According to another aspect of the present application, there is provided an instruction processing apparatus, which includes:
[0016] A receiving module, configured to receive first indication information; wherein, the first indication information indicates that the instruction processing process corresponding to the first instruction in the processor is completed;
[0017] An allocation module, configured to, in response to the first indication information, allocate a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the instruction generation order.
[0018] Optionally, the apparatus further includes:
[0019] A storage module, configured to store the first instruction to the target cache unit after allocating the target cache unit for the first instruction in the first cache.
[0020] Optionally, the apparatus further includes:
[0021] An instruction deletion module, configured to delete the first instruction and the second instruction in the instruction generation order when the second instruction generated before the first instruction is stored in the first cache.
[0022] According to another aspect of the present application, there is provided an electronic device, which includes:
[0023] One or more processors;
[0024] A memory;
[0025] One or more application programs, wherein the one or more application programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs are configured to: execute the instruction processing method according to any one of the first aspects of the present application.
[0026] For example, in the third aspect of the present application, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0027] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the instruction processing method as shown in the first aspect of the present application.
[0028] According to another aspect of the present application, there is provided a computer-readable storage medium, and when the computer program is executed by a processor, the instruction processing method described in any one of the first aspects of the present application is implemented.
[0029] For example, in the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the instruction processing method shown in the first aspect of the present application is implemented.
[0030] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the above-mentioned first aspect.
[0031] The beneficial effects brought by the technical solution provided by the present application are:
[0032] In the embodiment of the present application, by receiving first indication information; in response to the first indication information, a target cache unit is allocated for the first instruction in a first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed; that is to say, in the embodiment of the present application, when receiving the first indication information that the instruction processing flow corresponding to the first instruction has been completed, a target cache unit can be allocated for the first instruction in the first cache; compared with the prior art in which the ROB cache unit is allocated for the instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of the present application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.
[0034] Figure 1 It is a schematic flowchart of an instruction processing method provided by an embodiment of the present application;
[0035] Figure 2 It is a schematic diagram of an application scenario of an instruction processing method provided by an embodiment of the present application;
[0036] Figure 3 It is a schematic diagram of an application scenario of an instruction processing method provided by an embodiment of the present application;
[0037] Figure 4Schematic diagram of an application scenario of an instruction processing method provided by an embodiment of the present application;
[0038] Figure 5 Schematic diagram of the structure of an instruction processing device provided by an embodiment of the present application;
[0039] Figure 6 Schematic diagram of the structure of an electronic device for instruction processing provided by an embodiment of the present application. Detailed implementation manners
[0040] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0041] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0042] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0043] With the advent of the big data era, people's requirements for the processing speed of computers are getting higher and higher. To improve the performance of the processor, people have proposed to process instructions in a pipeline manner. Among them, the pipeline technology is a technology that decomposes each instruction into multiple stages and overlaps the operations of each stage, so as to achieve parallel processing of several instructions. Here, the instructions in the program are still executed sequentially one by one, but several instructions can be pre-obtained, and when the current instruction has not been executed yet, some other stages of the subsequent instructions can be started in advance, so that the processing efficiency of the processor can be significantly improved.
[0044] In a Central Processing Unit (CPU), there is more than one pipeline, and more than one instruction can be completed within each clock cycle. This design is called superscalar technology. Currently, superscalar processors are mainly adopted. In a superscalar processor, a processor core can perform a type of parallel operation with instruction-level parallelism. In this way, a superscalar processor can achieve a higher processor throughput at the same processor main frequency. For a superscalar processor, the instruction pipeline can include stages such as the Fetch stage, Decoder stage, Dispatch stage, Issue stage, Execute stage, and Retire stage, etc. Among them, in the Fetch stage, multiple instructions can be fetched from the instruction cache (I-Cache); in the Decoder stage, multiple instructions can be decoded sequentially within the same clock cycle, and then issued out-of-order in the Issue stage, enabling the Execute stage to execute these multiple instructions out-of-order. To enable the results of the out-of-order executed instructions to be retired in order, a Reorder Buffer (ROB) module is introduced in the instruction pipeline. The role of the ROB module is to reorder the results of the out-of-order executed instructions. After the instructions are executed out-of-order, the execution results will be submitted to the ROB module out-of-order. However, the ROB module will retire the out-of-order submitted execution results in the order of the cache entries.
[0045] That is to say, currently, superscalar technology is adopted in most processors. Such processors support the characteristics of multi-issue out-of-order execution. However, in high-performance superscalar out-of-order execution processors, the characteristics of "out-of-order execution and in-order retirement" are generally retained, making the internal execution details of the processor not perceptible to software, greatly reducing the software debugging difficulty. It should be noted that in most processor designs, it is usually the ROB module that sorts the instructions to ensure the in-order retirement of the instructions. This ROB design needs to allocate ROB resources at the previous stage before the instructions are out-of-order in the pipeline (usually the Dispatch stage). Among them, the time point for the ROB module to allocate resources is set at the Dispatch stage (specifically, allocate ROB cache resources. At the time point of the Dispatch stage, determine the quantity and location of the ROB resources to be occupied and reserve them in advance so that they cannot be occupied by other instructions anymore). This is because the ROB module needs to record the original order of the instruction stream before the instruction stream is out-of-order in order to achieve the instruction sorting function. However, the allocated ROB resources will not be used until the instructions go through stages such as issue, execution, and retirement. This increases resource waste to a certain extent.
[0046] Based on this, the embodiments of the present application provide an instruction processing method, which includes receiving first indication information; in response to the first indication information, allocating a target cache unit for the first instruction in a first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed; that is to say, the embodiments of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction has been completed; compared with the prior art in which the ROB cache unit is allocated for the instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiments of the present application achieve reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency.
[0047] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0048] See Figure 1 , the embodiments of the present application provide an instruction processing method. Optionally, this method is applied to an electronic device. For ease of description, the following takes this method applied to a processor as an example to introduce the embodiments of the present application; for example, the processor can be a CPU or a Graphics Processing Unit (GPU), etc.
[0049] In the embodiments of the present application, the processor is a core component for operations such as fetching instructions, decoding, issuing, and executing. It should be noted that, unless otherwise specified, the processor described in this article can specifically be a superscalar processor.
[0050] In a superscalar processor, it usually has the characteristics of "out-of-order execution and in-order retirement". Correspondingly, a superscalar processor is equipped with a ROB module. By sorting the instruction stream through the ROB module, the in-order retirement of instructions can be ensured. Specifically, in the design of the ROB module, the ROB module has multiple functions, which can include functions such as instruction sorting and retired instruction caching. In some architectures, it also undertakes the work of register renaming. However, this method not only increases the design complexity of the ROB module, but also causes resource waste to a certain extent and reduces resource utilization. Based on this, in the embodiments of the present application, the instruction sorting function and the retired instruction caching function are decoupled. In addition, it should be noted that other functions in the ROB module can also be decoupled from the ROB module. For example, the register renaming function can be decoupled from the ROB module, and the register renaming function can be completed through an independent register renaming module to release the renaming task of the ROB module; the embodiments of the present application do not make specific limitations on this, and the following will take the decoupling of the instruction sorting function from the ROB module as an example for detailed description.
[0051] It can be understood that in a superscalar processor, multiple instructions in the instruction stream sequence are sequential operations in the instruction fetch stage, decoding stage, and distribution stage; then from the emission stage, it enters out-of-order operation, so that these multiple instructions can be executed out of order in the execution stage. Thus, in order to implement the instruction sorting function, the embodiments of the present application need to record the original order of the instruction stream before the instruction stream becomes out-of-order. For the related technology, the ROB module needs to allocate resources in advance to record the original order of the instruction stream; in this way, considering from the perspective that the ROB module allocates resources in advance due to the need for instruction sorting, the related technology starts to allocate ROB resources at the previous stage (usually the distribution stage) of the instruction out-of-order in the pipeline, but it causes resource waste to a certain extent. In the embodiments of the present application, when receiving the first indication information indicating that the instruction processing flow corresponding to the first instruction in the processor has been completed, a target cache unit can be allocated for the first instruction in the first cache; it realizes reducing the occupation of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency. Specifically, the method can include the following steps:
[0052] S101: Receive the first indication information. Wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed.
[0053] Optionally, the embodiments of the present application can be applied to the field of computer technology, and specifically can be applied to the processing operation scenario of storing the first instruction in a re-order buffer (ROB; hereinafter can be simply referred to as ROB cache).
[0054] The first instruction may include any instruction in the processor. For example, the first instruction may include a data read instruction, a data write instruction, a data operation instruction, a target program start instruction, a target program exit instruction, and so on.
[0055] The first indication information is the indication information indicating that the instruction processing flow corresponding to the first instruction has been completed.
[0056] Specifically, it can be determined that the instruction processing flow of the first instruction has been completed in the following ways:
[0057] Way 1: Determine by generating the processing result of the first instruction. For example, the first instruction is a data operation instruction; when the data operation result corresponding to the first instruction is generated, it can be determined that the instruction processing flow of the first instruction has been completed.
[0058] Way 2: Determine by the completion of the processing steps of the instruction processing flow. For example, the instruction processing flow of the first instruction includes 5 processing steps executed in the execution order. When the 5th processing step is completed, it can be determined that the instruction processing flow of the first instruction has been completed. As an example, in combination with Figure 2 , the instruction processing flow of the embodiment of the present application will be described: The instruction processing flow shown in the figure includes 5 processing steps, which are the instruction fetch step, the decoding step, the dispatching step, the issuing step, and the execution step respectively.
[0059] Instruction Fetch Unit (IFU): Execute to obtain an instruction, and use the value of the Program Counter Register (PC) register as the address to fetch the instruction from the I-Cache.
[0060] Decoder: Execute to decode the fetched instruction and read the register file according to the decoding result to obtain the source operands of the instruction.
[0061] Dispatcher: Execute to send the decoded instruction to the issuing module in the original order specified in the program.
[0062] Issue: Execute to send the instruction in the issue queue to the execution module. Specifically, during the execution of the instruction pipeline, the instruction after fetching, decoding, and dispatching will be pushed onto the issuing module and cached in the issue queue of the issuing module.
[0063] Execute or lsu: Execute to execute the instruction according to the decoding result.
[0064] S102: In response to the first indication information, allocate a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the order of instruction generation.
[0065] Wherein, the first cache includes a cache for storing instructions that have completed the instruction processing flow; optionally, in the embodiments of the present application, the first cache may be a ROB cache.
[0066] In a high-performance superscalar processor, instruction out-of-order execution is usually adopted to obtain better instruction-level parallelism (ILP), so as to improve the processor performance. Among them, in order to support functions such as precise exceptions, hardware speculation, and register renaming in instruction out-of-order execution, a reorder buffer (ROB) mechanism is usually adopted, that is, the instruction sequential exit pipeline of the instruction processing flow is realized through the reorder buffer ROB, so as to ensure the correctness of program out-of-order execution.
[0067] For example, in an actual processing scenario, a ROB cache unit can be allocated for an instruction after instruction decoding to complete register renaming and instruction order preservation; in the instruction processing completion stage, the instruction that has completed the instruction processing flow is stored in the ROB cache; then, the instructions exit the ROB cache in the order of generation; after the instructions exit, the general-purpose registers (GPRs) and memories are updated sequentially, ensuring the implementation of precise exceptions and hardware speculation.
[0068] However, in the above processing scenario, in the instruction decoding stage, that is, when the instruction processing has not been completed, a ROB cache unit has already been allocated for the instruction, and it is not until the instruction processing flow has been completed that the instruction is stored in the ROB cache; in this way, when the instruction processing flow has not been completed, pre-allocating a ROB cache unit in the ROB cache will occupy cache resources and reduce the number of instructions to be processed flowing on the instruction processing pipeline.
[0069] To reduce the occupation of cache resources, increase the number of instructions to be processed on the instruction processing pipeline, and improve the instruction processing efficiency, the embodiments of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the instruction (i.e., the first instruction of the present application) has been completed.
[0070] Wherein, the target cache unit is the cache unit corresponding to the first instruction. In the first cache, for example, in the ROB cache, the stored instructions are arranged in the order of instruction generation.
[0071] Optionally, in an actual processing scenario, since the instructions stored in the ROB cache are arranged in the order of instruction generation, embodiments of the present application can allocate a target cache unit for the first instruction through the correspondence between the identity document (ID) of the first instruction and the cache unit identifier.
[0072] Specifically, instruction IDs can be established for each instruction after instruction generation to identify the order of instruction generation; and record the cache unit identifier of the cache unit where the first instruction is located in the ROB cache; then determine the target cache unit corresponding to the first instruction through the correspondence between the instruction ID of the first instruction and its corresponding cache unit identifier.
[0073] For example, the instruction ID of the first generated instruction (such as the first generated instruction within a pre-designed time period) can be "0", that is, the first generated instruction is instruction 0; the cache unit identifier of the cache unit where instruction 0 is located in the ROB cache is 10, that is, instruction 0 corresponds to cache unit 10; since the instructions stored in the ROB cache are arranged in the order of instruction generation, then, the instructions generated after instruction 0 can be stored in the cache units after cache unit 10. For example, instruction 1 is stored in cache unit 11, instruction 2 is stored in cache unit 12, instruction 3 is stored in cache unit 13, and so on. In addition, in some embodiments, the cache units corresponding to instructions may not be adjacent; for example, instruction 1 is stored in cache unit 11, instruction 2 is stored in cache unit 13, instruction 3 is stored in cache unit 15, and so on.
[0074] Embodiments of the present application receive first indication information; in response to the first indication information, allocate a target cache unit for the first instruction in the first cache; where the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed; that is to say, embodiments of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction has been completed; compared with the prior art method of allocating a ROB cache unit for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; embodiments of the present application achieve reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency.
[0075] In an embodiment of the present application, after allocating the target cache unit for the first instruction in the first cache, the method further includes:
[0076] Store the first instruction into the target cache unit.
[0077] After determining the target cache unit corresponding to the first instruction, store the first instruction in the target cache unit. It can be understood that in the embodiments of the present application, the first instruction stored in the target cache unit may be the instruction information corresponding to the first instruction. For example, the ID of the first instruction, the processing object of the first instruction, the processing status of the first instruction, and so on.
[0078] In one embodiment of the present application, the method further includes:
[0079] If the second instruction generated before the first instruction is stored in the first cache, delete the first instruction and the second instruction in the order of instruction generation.
[0080] Specifically, the second instruction includes the instructions generated before the first instruction.
[0081] When the second instruction is stored in the first cache, that is, when the instruction processing flow corresponding to the second instruction is completed and the second instruction is stored in the first cache, the first instruction and the second instruction can be deleted in the order of instruction generation.
[0082] It can be understood that in the embodiments of the present application, the instructions whose instruction processing flow is completed are stored in the ROB cache; then, the instructions are withdrawn from the ROB cache in the order of generation.
[0083] In one embodiment of the present application, after receiving the instruction generation message, the method further includes:
[0084] Add an identity identifier to the first instruction, and the identity identifier indicates the instruction generation order of the first instruction.
[0085] In one embodiment of the present application, the step of allocating a target cache unit for the first instruction in the first cache includes:
[0086] Determine the target cache unit corresponding to the identity identifier according to the correspondence between the identity identifier and the cache unit identifier.
[0087] Optionally, in an actual processing scenario, since the instructions stored in the ROB cache are arranged in the order of instruction generation, in the embodiments of the present application, the target cache unit can be allocated for the first instruction through the correspondence between the identity identifier (Identity document, ID) of the first instruction and the cache unit identifier.
[0088] Specifically, after the instruction is generated, an instruction ID can be established for each instruction to identify the sequence of instruction generation; and the cache unit identifier of the cache unit where the first instruction is stored in the ROB cache is recorded; then, through the correspondence between the instruction ID of the first instruction and its corresponding cache unit identifier, the target cache unit corresponding to the first instruction is determined.
[0089] For example, the instruction ID of the first generated instruction (such as the first generated instruction within a pre-designed time period) can be "0", that is, the first generated instruction is instruction 0; the cache unit identifier of the cache unit where instruction 0 is stored in the ROB cache is 10, that is, the cache unit corresponding to instruction 0 is 10; since the instructions stored in the ROB cache are arranged in the order of instruction generation, then, the instructions generated after instruction 0 can be stored in the cache units after cache unit 10. For example, instruction 1 is stored in cache unit 11, instruction 2 is stored in cache unit 12, instruction 3 is stored in cache unit 13, and so on. In addition, in some embodiments, the cache units corresponding to the instructions may not be adjacent; for example, instruction 1 is stored in cache unit 11, instruction 2 is stored in cache unit 13, instruction 3 is stored in cache unit 15, and so on.
[0090] The embodiment of the present application receives first indication information; in response to the first indication information, a target cache unit is allocated for the first instruction in the first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed; that is to say, the embodiment of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction has been completed; compared with the prior art method of allocating a ROB cache unit for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of the present application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency.
[0091] The instruction retirement scenario of the embodiment of the present application is described below through two examples:
[0092] Example 1, in combination with Figure 3 , the instruction retirement scenario is described.
[0093] The instruction processing flow shown in the figure includes 6 processing steps, and the 6 processing steps are respectively an instruction fetching step, a decoding step, a distribution step, an issuing step, an execution step, and a retirement step.
[0094] Instruction Fetch Unit (IFU): Fetch instructions. Use the value of the Program Counter Register (PC) as the address to fetch instructions from the I-Cache.
[0095] Decoder: Decode the fetched instructions and read the register file according to the decoding results to obtain the source operands of the instructions.
[0096] Dispatcher: Send the decoded instructions to the issue module in the original order specified in the program.
[0097] Issue: Send the instructions in the issue queue to the execution module. Specifically, during the execution of the instruction pipeline, the instructions after fetching, decoding, and dispatching are pushed onto the issue module and cached in the issue queue of the issue module.
[0098] Execute or lsu: Execute the instructions according to the decoding results.
[0099] Retire: Retire the instructions after execution is completed.
[0100] As Figure 3 shown, the identifiers of the cache units in the ROB cache start from 10, and then increase sequentially from top to bottom, such as 11, 12, 13, 14, etc. Assume that the instructions with instruction IDs 0, 1, 2, 3, and 4 are flowing in the instruction pipeline.
[0101] At time t0, the instruction processing flow of the above instructions has not been completed (that is, the instruction processing status shown in the figure is 0); at time t1, instruction 1 completes the instruction processing flow (that is, the instruction processing status shown in the figure is 1), so instruction 1 can be stored in the cache unit; since the oldest instruction generated first (that is, instruction 0 corresponding to the oldest id in the figure) can be stored in cache unit 10 (that is, the oldest instruction in the figure), entry corresponding to cache unit 10), then, according to the correspondence between the instruction identifier and the cache unit identifier, instruction 1 can be stored in cache unit 11; at time t2, instruction 3 completes the instruction processing flow, so instruction 3 can be stored in the cache unit, and according to the correspondence between the instruction identifier and the cache unit identifier, instruction 3 can be stored in cache unit 13; at time t3, instruction 2 completes the instruction processing flow, so instruction 2 can be stored in the cache unit, and according to the correspondence between the instruction identifier and the cache unit identifier, instruction 2 can be stored in cache unit 12; at time t4, instruction 0 completes the instruction processing flow, so instruction 0 can be stored in cache unit 10; at this time, instruction 0, instruction 1, and instruction 2 generated before instruction 3 have all been stored in the cache unit, then instruction 0, instruction 1, instruction 2, and instruction 3 can be retired, that is, instruction 0, instruction 1, instruction 2, and instruction 3 can be deleted from the ROB cache. Correspondingly, the oldest id is updated to 4, and the oldest rob entry is updated to 14.
[0102] It should also be noted that, in the embodiment of the present application, the oldest instruction (also referred to as the "oldest instruction") refers to the instruction that is ranked first in the original order of the instruction stream among all instructions that have not yet been retired, and the value of the oldest instruction identification information carried by the oldest instruction is the smallest. In other words, the smaller the value of the instruction id, the older the instruction is, and the higher it is ranked in the original order of the instruction stream; conversely, the larger the value of the instruction id, the newer the instruction is, and the later it is ranked in the original order of the instruction stream. Normally, the value of the oldest instruction id carried by the oldest instruction is zero, and the value of the corresponding oldest cache identifier in the ROB cache is also zero, but this is not specifically limited.
[0103] Example 2: Combination Figure 4 , describes the retirement scenario of instructions when an exception occurs during the instruction execution process.
[0104] like Figure 4 The instruction processing flow shown includes 6 processing steps, which are instruction fetch step, decoding step, distribution step, emission step, execution step, and retirement step.
[0105] like Figure 4As shown, the identifiers of the cache units in the ROB cache start from 10, and then 11, 12, 13, 14, etc., increasing sequentially from top to bottom. Assume that the instructions with instruction IDs 0, 1, 2, 3, and 4 are flowing in the instruction pipeline.
[0106] At time t0, the instruction processing flows of the above instructions are not all completed (i.e., the instruction processing status shown in the figure is 0); at time t1, instruction 1 completes the instruction processing flow (i.e., the instruction processing status shown in the figure is 1). Therefore, instruction 1 can be stored in the cache unit; since the earliest generated instruction (i.e., instruction 0 corresponding to the oldest id in the figure) can be stored in cache unit 10 (i.e., cache unit 10 corresponding to the oldest rob entry in the figure), then, according to the correspondence between the instruction identifier and the cache unit identifier, instruction 1 can be stored in cache unit 11; at time t2, instruction 3 completes the instruction processing flow. Therefore, instruction 3 can be stored in the cache unit. According to the correspondence between the instruction identifier and the cache unit identifier, instruction 3 can be stored in cache unit 13; at time t3, instruction 2 completes the instruction processing flow. Therefore, instruction 2 can be stored in the cache unit. According to the correspondence between the instruction identifier and the cache unit identifier, instruction 2 can be stored in cache unit 12. However, there is an exception during the instruction processing of instruction 2, and this instruction can be marked as an abnormal instruction; at time t4, instruction 0 completes the instruction processing flow. Therefore, instruction 0 can be stored in cache unit 10; at this time, the instructions 0, 1, and 2 generated before instruction 3 have all been stored in the cache unit. Also, since instruction 2 is an abnormal instruction, only instructions 0 and 1 can be retired, that is, instructions 0 and 1 can be deleted from the ROB cache. In this case, due to the exception of instruction 2, the instructions generated after instruction 2 can also be considered abnormal instructions. Then, instruction 2 and the instructions 3 and 4 generated after instruction 2 can be deleted; then, the newly generated instructions are identified starting from instruction ID "2".
[0107] The embodiment of the present application receives the first indication information; in response to the first indication information, allocates a target cache unit for the first instruction in the first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor is completed; that is to say, the embodiment of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction is completed; compared with the prior art method of allocating a ROB cache unit for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of the present application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction pipeline, and improving the instruction processing efficiency.
[0108] In short, in the embodiments of the present application, the instruction sorting function and the retired instruction cache function included in the traditional ROB module are decoupled, so that the decoupled ROB module will only retain the retired instruction cache function, and the instruction sorting function is provided by a separate tagging module. Since the instruction sorting function is decoupled from the traditional ROB module, the processing of allocating cache units can be postponed until after the instruction execution is completed. Generally, the cache module and caches such as the issue queue, load queue, and store queue are in a serial and upstream / downstream relationship, rather than the coexistence mode in the related art, effectively improving the resource utilization rate.
[0109] An embodiment of the present application provides an instruction processing device, as Figure 5 shown. The instruction processing device 50 may include: a receiving module 501 and an allocation module 502. Among them,
[0110] The receiving module 501 is configured to receive first indication information; wherein, the first indication information indicates that the instruction processing process corresponding to the first instruction in the processor is completed;
[0111] The allocation module 502 is configured to, in response to the first indication information, allocate a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the instruction generation order.
[0112] In an embodiment of the present application, the device further includes:
[0113] A storage module, configured to store the first instruction in the target cache unit after allocating the target cache unit for the first instruction in the first cache.
[0114] In an embodiment of the present application, the device further includes:
[0115] An instruction deletion module, configured to delete the first instruction and the second instruction in the instruction generation order when the second instruction generated before the first instruction is stored in the first cache.
[0116] In an embodiment of the present application, the device further includes:
[0117] An identification adding module, configured to, after receiving the instruction generation message,
[0118] add an identity identifier to the first instruction, and the identity identifier indicates the instruction generation order of the first instruction.
[0119] In one embodiment of the present application, the allocation module is specifically configured to determine a target cache unit corresponding to the identity identifier according to the correspondence between the identity identifier and the cache unit identifier.
[0120] In short, compared with the related art, the instruction processing device of the embodiment of the present application mainly includes:
[0121] (1) An identity addition module is added. This module can label each executed instruction to indicate the sequence of each instruction, and this label will flow in the pipeline with the instruction until the instruction is retired. As an independent module, this module can perform operations at any stage before the out-of-order stage in the instruction pipeline. For example, in the IFU stage, when an instruction is fetched from the I-Cache, the labeling module follows the PC value to determine the instruction address corresponding to the instruction to complete the labeling of the instruction, that is, add the corresponding instruction identity information.
[0122] (2) The decoupled allocation module only retains the function of retiring instruction caching. In addition, the implementation logics of other exception marking, pipeline flushing, etc. are basically the same as those of the ROB module before decoupling. However, after the function of the ROB module is decoupled, the allocation of ROBentry will be performed when the instruction execution is completed.
[0123] (3) One implementation method: According to the current oldest identity register, this register is used to record the instruction identity information of the oldest instruction among the instructions that have not been retired yet. Exemplarily, after a typical power-on reset, the register value is 0, that is, the instruction numbered 0. In this way, this register will assist the allocation module to complete the allocation process of writing the corresponding instruction information of the executed instruction into the ROBentry.
[0124] (4) Another implementation method: According to the mapping between the instruction identity information and the cache unit identity information, maintain the correspondence between the instruction identity information carried by the current instruction and the cache unit identity information of the ROB entry where it is located, which will also assist the allocation module to complete the allocation process of writing the corresponding instruction information of the executed instruction into the ROB entry of the cache module.
[0125] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function descriptions of the modules of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details will not be repeated here.
[0126] An embodiment of the present application receives first indication information; in response to the first indication information, a target cache unit is allocated for the first instruction in a first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor is completed; that is to say, the embodiment of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction is completed; compared with the prior art in which a ROB cache unit is allocated for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of the present application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing on the instruction processing pipeline, and improving the instruction processing efficiency.
[0127] An embodiment of the present application provides an electronic device, which includes: a memory and a processor; at least one program, stored in the memory, when executed by the processor, can achieve compared with the prior art: the embodiment of the present application receives first indication information; in response to the first indication information, a target cache unit is allocated for the first instruction in a first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor is completed; that is to say, the embodiment of the present application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction is completed; compared with the prior art in which a ROB cache unit is allocated for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of the present application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing on the instruction processing pipeline, and improving the instruction processing efficiency.
[0128] In an optional embodiment, an electronic device is provided, as Figure 6 shown, Figure 6 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiment of the present application.
[0129] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0130] The bus 4002 can include a path for transmitting information among the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0131] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0132] The memory 4003 is used to store the application program code (computer program) for executing the solution of this application, and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0133] Among them, the electronic device includes but is not limited to: mobile phone, laptop, multimedia player, desktop computer, etc.
[0134] The embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0135] The embodiment of this application receives first indication information; in response to the first indication information, a target cache unit is allocated for the first instruction in the first cache; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor has been completed; that is to say, the embodiment of this application can allocate a target cache unit for the first instruction in the first cache when receiving the first indication information that the instruction processing flow corresponding to the first instruction has been completed; compared with the prior art method of allocating a ROB cache unit for an instruction before the instruction processing is completed, and the instruction is stored in the ROB cache only after the instruction processing flow is completed; the embodiment of this application realizes reducing the occupancy of cache resources, increasing the number of instructions to be processed flowing in the instruction processing pipeline, and improving the instruction processing efficiency.
[0136] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than that shown or described in words.
[0137] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0138] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. An instruction processing method, characterized in that Including: Receiving first indication information; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor is completed; In response to the first indication information, allocating a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the instruction generation order.
2. The instruction processing method according to claim 1, wherein After allocating the target cache unit for the first instruction in the first cache, the method further includes: Storing the first instruction into the target cache unit.
3. The instruction processing method according to claim 1, characterized in that The method further includes: If a second instruction generated before the first instruction is stored in the first cache, deleting the first instruction and the second instruction in the instruction generation order.
4. The instruction processing method according to any one of claims 1 to 3, characterized in that, The method further includes: Adding an identity identifier to the first instruction, and the identity identifier indicates the instruction generation order of the first instruction.
5. The instruction processing method according to claim 4, wherein The allocating a target cache unit for the first instruction in the first cache includes: Determining the target cache unit corresponding to the identity identifier according to the correspondence between the identity identifier and the cache unit identifier.
6. An instruction processing device, characterized in that, Including: A receiving module, configured to receive first indication information; wherein, the first indication information indicates that the instruction processing flow corresponding to the first instruction in the processor is completed; An allocating module, configured to, in response to the first indication information, allocate a target cache unit for the first instruction in a first cache; wherein, the instructions stored in the first cache are arranged in the instruction generation order.
7. The instruction processing apparatus according to claim 6, wherein The apparatus further includes: A storing module, configured to store the first instruction into the target cache unit after allocating the target cache unit for the first instruction in the first cache.
8. The instruction processing device according to claim 6, wherein, The apparatus further includes: An instruction deleting module, configured to delete the first instruction and the second instruction in the instruction generation order when a second instruction generated before the first instruction is stored in the first cache.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the instruction processing method according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the instruction processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Processor and instruction operation method
CN110806898A
Device and method for submitting instructions out of order
CN114217859A