Instruction Processing Method, Apparatus, Electronic Device, Storage Medium, and Program Product
By managing the cache of dependency instructions and its result data in a multi-level pipeline, the problem of inefficiency of processors when processing dependency instructions is solved, and more efficient and accurate instruction processing is achieved.
Patent Information
- Application Number
- CN202410458137.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-04-16
AI Technical Summary
In the arithmetic logic unit that adopts a multi-stage pipeline mechanism, how to process instructions with dependencies more efficiently and accurately to improve the instruction processing efficiency within the same time.
By determining the current processing level of the later execution instruction that has a dependency relationship with the first execution instruction in the multi-stage pipeline, and when appropriate, the result data written back to the level is stored in the cache space of the later execution instruction, and then passing it until the instruction is executed to the target processing level.
This method improves the processor's execution efficiency and accuracy of multiple instructions with dependencies, and avoids the reduction in efficiency due to strict alignment of instruction execution cycles.
Smart Images

Figure CN118276950B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of instruction processing, specifically to the technical fields of processors, arithmetic logic units, multi-stage pipelines, data bypass, etc., and particularly to an instruction processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] An Arithmetic and Logic Unit (ALU) is a basic logic operation unit that constitutes a processor and is also a combinational logic circuit system that implements multiple sets of arithmetic operations and logic operations.
[0003] The ALU in a processor is usually implemented in a multi-stage pipeline mechanism or structure. That is, when the ALU executes an instruction passed in by the processor, it first reads one or more source operands to be processed indicated by the instruction from the storage space, then sends the source operands to the data processing stage for corresponding calculations, and finally writes the calculation result back to the storage space.
[0004] In an ALU adopting a multi-stage pipeline mechanism, how to process instructions with dependency relationships more efficiently and accurately to improve the instruction processing efficiency within the same time is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] Embodiments of the present disclosure propose an instruction processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which are used to enable the arithmetic logic unit adopting a multi-stage pipeline mechanism in a processor to process instructions with dependency relationships more efficiently and accurately, so as to improve the instruction processing efficiency within the same time.
[0006] In a first aspect of the embodiments of the present disclosure, an instruction processing method is proposed, including: in response to a first instruction being executed to the data write-back stage in the multi-stage pipeline of the arithmetic logic unit, determining the current processing stage to which a second instruction having a dependency relationship with the first instruction is executed in the multi-stage pipeline, and the multiple data processing stages constituting the multi-stage pipeline include multiple data reading stages, multiple data calculation stages, and one data write-back stage arranged in sequence; in response to the current processing stage being a pre-processing stage of a target processing stage, storing the result data to be written back at the data write-back stage into the cache space corresponding to the second instruction, and the target processing stage is determined based on the instruction processing logic or dependency relationship of the second instruction and is one of the multiple data reading stages or multiple data calculation stages; in response to the second instruction being executed to the target processing stage in the multi-stage pipeline, transferring the result data stored in the cache space to the target processing stage.
[0007] In some other embodiments of the first aspect, the method further includes:
[0008] Determine that there is a dependency relationship between the second instruction and the first instruction according to the instruction association information passed in together with the second instruction; or
[0009] Perform instruction parsing on the second instruction at the instruction parsing stage of the multi-stage pipeline, and determine that there is a dependency relationship between the second instruction and the first instruction according to the obtained instruction parsing result. The multi-stage pipeline further includes an instruction parsing stage located before the first data reading stage.
[0010] In some other embodiments of the first aspect, storing the result data to be written back at the data write-back stage into the cache space corresponding to the second instruction includes:
[0011] Store the result data into the cache space corresponding to the second instruction in a preset data processing stage of the multi-stage pipeline through the first data bypass. The first data bypass exists as a data transmission path connecting the data write-back stage and the preset data processing stage.
[0012] In some other embodiments of the first aspect, the dependency relationship is read-after-write related, the target processing stage is the first data calculation stage, and the preset data processing stage is the last data reading stage or the stage before the first data calculation stage.
[0013] In some other embodiments of the first aspect, the method further includes:
[0014] In response to the second instruction entering the multi-stage pipeline, create or allocate a corresponding cache space for the second instruction;
[0015] In response to the result data being successfully acquired by the target processing stage, release the cache space.
[0016] In some other embodiments of the first aspect, the number of levels of the data reading stage is N, where N is a positive integer not less than 1. Creating or allocating a corresponding cache space for the second instruction includes:
[0017] Determine the target number of the cache space according to the number of source operands to be read by the second instruction through N data reading stages;
[0018] Determine the target depth of the cache space according to the number of levels of N data reading stages;
[0019] Create or allocate a cache space for the second instruction with the target number and the target depth.
[0020] In some other embodiments of the first aspect, the target number is the number of source operands to be read by the second instruction through N data reading stages; wherein, different source operands are independently stored in different cache spaces.
[0021] In some other embodiments of the first aspect, the target depth is N.
[0022] In some other embodiments of the first aspect, the cache space includes a first-in, first-out (FIFO) queue with a queue length of N.
[0023] In some other embodiments of the first aspect, the method further includes:
[0024] Using the FIFO queue to temporarily store the source operands obtained through N data reading stages.
[0025] In some other embodiments of the first aspect, storing the result data to be written back by the data write-back stage into the cache space corresponding to the second instruction includes:
[0026] Storing the result data to be written back by the data write-back stage into the corresponding position of the first-in, first-out queue according to the polling rule.
[0027] In some other embodiments of the first aspect, the method further includes:
[0028] In response to the result data to be written back by the data write-back stage having been stored in the cache space corresponding to the second instruction, attaching a tag indicating that there is valid data to the cache space.
[0029] In some other embodiments of the first aspect, the method further includes:
[0030] In response to the result data being successfully obtained by the target processing stage, clearing the tag attached to the cache space.
[0031] In some other embodiments of the first aspect, the dependency relationship is write-after-read related, the number of stages of the data calculation stage is M, where M is a positive integer not less than 1, and the second instruction enters the multi-stage pipeline after the M-th execution cycle when the first instruction enters the multi-stage pipeline.
[0032] In some other embodiments of the first aspect, the method further includes:
[0033] In response to the first instruction entering the multi-stage pipeline, creating or allocating a corresponding cache space for the first instruction;
[0034] Temporarily storing the source operands obtained when the first instruction is executed to the data reading stage into the cache space corresponding to the first instruction;
[0035] In response to the source operands being successfully obtained when the first instruction is executed to the data calculation stage, releasing the cache space corresponding to the first instruction.
[0036] In some other embodiments of the first aspect, the method further includes:
[0037] Enter a multi - stage pipeline in response to a third instruction, and create or allocate a corresponding cache space for the third instruction;
[0038] Temporarily store the source operands obtained when the third instruction is executed to the data reading stage in the cache space corresponding to the third instruction;
[0039] In response to the successful acquisition of the source operands when the third instruction is executed to the data calculation stage, release the cache space corresponding to the third instruction.
[0040] In some other embodiments of the first aspect, the method further includes:
[0041] In response to the current processing stage being the target processing stage, directly transfer the result data that should be written back at the data write - back stage to the target processing stage.
[0042] In some other embodiments of the first aspect, directly transferring the result data that should be written back at the data write - back stage to the target processing stage includes:
[0043] Directly transfer the result data that should be written back at the data write - back stage to the target processing stage through a second data bypass, and the second data bypass exists as a data transmission path connecting the data write - back stage and the target processing stage.
[0044] In a second aspect, an instruction processing apparatus is proposed according to an embodiment of the present disclosure, including: a current processing stage determination module configured to determine the current processing stage at which a second instruction having a dependency relationship with a first instruction is executed in a multi - stage pipeline in response to the first instruction being executed to the data write - back stage in the multi - stage pipeline of an arithmetic logic unit. The multiple data processing stages constituting the multi - stage pipeline include multiple data reading stages, multiple data calculation stages, and one data write - back stage arranged in sequence; a result data cache module configured to store the result data that should be written back at the data write - back stage in the cache space corresponding to the second instruction in response to the current processing stage being the pre - processing stage of the target processing stage. The target processing stage is determined based on the instruction processing logic or dependency relationship of the second instruction and is one of the multiple data reading stages or multiple data calculation stages; a result data transfer module configured to transfer the result data stored in the cache space to the target processing stage in response to the second instruction being executed to the target processing stage in the multi - stage pipeline.
[0045] In some other embodiments of the second aspect, the apparatus further includes:
[0046] A dependency relationship first determination module configured to determine that the second instruction has a dependency relationship with the first instruction according to the instruction association information passed in together with the second instruction; or
[0047] The second dependency determination module is configured to perform instruction parsing on a second instruction at the instruction parsing stage of a multi-stage pipeline, and determine that there is a dependency relationship between the second instruction and a first instruction according to the obtained instruction parsing result. The multi-stage pipeline further includes an instruction parsing stage located before the first data reading stage.
[0048] In some other embodiments of the second aspect, the result data caching module is further configured to:
[0049] Store the result data into a cache space corresponding to the second instruction in a preset data processing stage of the multi-stage pipeline through a first data bypass. The first data bypass exists as a data transmission path connecting the data write-back stage and the preset data processing stage.
[0050] In some other embodiments of the second aspect, the dependency relationship is read-after-write related, the target processing stage is the first data calculation stage, and the preset data processing stage is the last data reading stage or the stage before the first data calculation stage.
[0051] In some other embodiments of the second aspect, the apparatus further includes:
[0052] The first cache space processing module is configured to create or allocate a corresponding cache space for the second instruction in response to the second instruction entering the multi-stage pipeline;
[0053] The first release module is configured to release the cache space in response to the result data being successfully acquired by the target processing stage.
[0054] In some other embodiments of the second aspect, the number of data reading stages is N, where N is a positive integer not less than 1. The first cache space processing module is further configured to:
[0055] Determine the target number of cache spaces according to the number of source operands to be read by the second instruction through N data reading stages;
[0056] Determine the target depth of the cache space according to the number of stages of N data reading stages;
[0057] Create or allocate a cache space for the second instruction with the target number and the target depth.
[0058] In some other embodiments of the second aspect, the target number is the number of source operands to be read by the second instruction through N data reading stages; wherein, different source operands are independently stored in different cache spaces.
[0059] In some other embodiments of the second aspect, the target depth is N.
[0060] In some other embodiments of the second aspect, the cache space includes a first-in, first-out (FIFO) queue with a queue length of N.
[0061] In some other embodiments of the second aspect, the apparatus further includes:
[0062] A first source operand staging module, configured to stage the source operands obtained through N data read levels by using the FIFO queue.
[0063] In some other embodiments of the second aspect, the result data cache module is further configured to:
[0064] Store the result data to be written back by the data write-back stage into the corresponding positions of the first-in, first-out queue according to a polling rule.
[0065] In some other embodiments of the second aspect, the apparatus further includes:
[0066] A tag addition module, configured to add a tag storing valid data to the cache space in response to the result data to be written back by the data write-back stage being stored in the cache space corresponding to the second instruction.
[0067] In some other embodiments of the second aspect, the apparatus further includes:
[0068] A tag clearing module, configured to clear the tag added to the cache space in response to the result data being successfully obtained by the target processing stage.
[0069] In some other embodiments of the second aspect, the dependency is read-after-write related, the number of levels of the data calculation stage is M, where M is a positive integer not less than 1, and the second instruction enters the multi-stage pipeline after the Mth execution cycle when the first instruction enters the multi-stage pipeline.
[0070] In some other embodiments of the second aspect, the apparatus further includes:
[0071] A second cache space processing module, configured to create or allocate a corresponding cache space for the first instruction in response to the first instruction entering the multi-stage pipeline;
[0072] A second source operand staging module, configured to stage the source operands obtained when the first instruction is executed to the data read level into the cache space corresponding to the first instruction;
[0073] A second release module, configured to release the cache space corresponding to the first instruction in response to the source operands being successfully obtained when the first instruction is executed to the data calculation level.
[0074] In some other embodiments of the second aspect, the apparatus further includes:
[0075] The third processing module for cache space is configured to enter a multi - stage pipeline in response to a third instruction and create or allocate a corresponding cache space for the third instruction;
[0076] The third temporary storage module for source operands is configured to temporarily store the source operands obtained when the third instruction is executed to the data reading stage in the cache space corresponding to the third instruction;
[0077] The third release module is configured to release the cache space corresponding to the third instruction in response to the successful acquisition of the source operands when the third instruction is executed to the data calculation stage.
[0078] In some other embodiments of the second aspect, the apparatus further includes:
[0079] The direct transfer module is configured to directly transfer the result data to be written back at the data write - back stage to the target processing stage in response to the current processing stage being the target processing stage.
[0080] In some other embodiments of the second aspect, the direct transfer module is further configured to:
[0081] Directly transfer the result data to be written back at the data write - back stage to the target processing stage through a second data bypass, and the second data bypass exists as a data transmission path connecting the data write - back stage and the target processing stage.
[0082] In a third aspect, an embodiment of the present disclosure provides an arithmetic logic unit, including: a multi - stage pipeline structure including a plurality of sequentially arranged data reading stages, a plurality of data calculation stages, and a data write - back stage; a cache controller for determining the current processing stage to which a second instruction having a dependency relationship with the first instruction is executed in response to the first instruction being executed to the data write - back stage; storing the result data to be written back at the data write - back stage in the cache space corresponding to the second instruction in response to the current processing stage being the pre - processing stage of the target processing stage; and transferring the result data stored in the cache space to the target processing stage in response to the second instruction being executed to the target processing stage, where the target processing stage is determined based on the instruction processing logic of the current instruction or the dependency relationship with a previous instruction in the arithmetic logic unit and is one of the plurality of data reading stages or the plurality of data calculation stages.
[0083] In some other embodiments of the third aspect, the arithmetic logic unit further includes:
[0084] A first data bypass, where the first data bypass is connected between the data write - back stage and the pre - processing stage;
[0085] The cache controller is specifically configured to: in response to the current processing stage being the pre - processing stage, store the result data to be written back at the data write - back stage in the cache space corresponding to the second instruction in the pre - processing stage through the first data bypass.
[0086] In some other embodiments of the third aspect, the arithmetic logic unit further includes:
[0087] A second data bypass, which is connected between the data write-back stage and the target processing stage;
[0088] The cache controller is further configured to: in response to the current processing stage being the target processing stage, directly pass the result data to be written back by the data write-back stage to the target processing stage through the second data bypass.
[0089] In some other embodiments of the third aspect, the first data bypass and the second data bypass include: wires and data transmission devices on the wires.
[0090] In a fourth aspect, an embodiment of the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the instruction processing method described in any implementation manner of the first aspect.
[0091] In a fifth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to implement the instruction processing method described in any implementation manner of the first aspect.
[0092] In a sixth aspect, an embodiment of the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the computer program is enabled to implement the steps of the instruction processing method described in any implementation manner of the first aspect.
[0093] In the instruction processing solution provided in this embodiment, when the first instruction that first enters the multi-stage pipeline is executed to the data write-back stage, it is determined to which data processing stage (collectively referred to as the current processing stage hereinafter) the second instruction that later enters the multi-stage pipeline and has a dependency relationship with the first instruction is currently executed, and when it is found that the current processing stage is a preprocessing stage corresponding to the data write-back stage, the result data is temporarily stored in the cache space corresponding to the second instruction, so that when the second instruction is subsequently executed to the target processing stage, the result data stored in the cache space is passed to the target processing stage.
[0094] That is, as can be seen from the above instruction processing solution provided in this embodiment, when the first instruction has been executed up to the final data write-back stage, by storing the result data to be written back in the cache space pre-created or allocated for the second instruction, the second instruction can obtain it from the cache space at any time when it is executed up to the target processing stage where the result data is required, without strictly adhering to the requirement that the target processing stage and the data write-back stage must be in the same execution cycle to transfer the result data. This is equivalent to relaxing the timing for the second instruction to enter the multi-stage pipeline and start execution, thus optimizing, expanding, and improving the existing data bypass technology, and further enhancing the execution efficiency and accuracy of the processor for multiple instructions with dependency relationships.
[0095] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent:
[0097] Figure 1 is an exemplary system architecture in which the present disclosure can be applied;
[0098] Figure 2 is a flowchart of an instruction processing method provided by an embodiment of the present disclosure;
[0099] Figure 3 is a specific structural schematic diagram of a multi-stage pipeline structure adopted by an ALU provided by an embodiment of the present disclosure;
[0100] Figure 4 is a flowchart of another instruction processing method provided by an embodiment of the present disclosure;
[0101] Figure 5 is a flowchart of a method for creating or allocating cache space provided by an embodiment of the present disclosure;
[0102] Figure 6 For the embodiment of the present disclosure Figure 3 is a schematic diagram of the first data bypass and the second data bypass set for the multi-stage pipeline structure of the arithmetic logic unit shown;
[0103] Figures 7-1 to 7-7 are respectively execution timing relationship diagrams or hardware working diagrams representing instructions provided by an embodiment of the present disclosure in an application scenario;
[0104] Figure 8 is a structural block diagram of an instruction processing device provided by an embodiment of the present disclosure;
[0105] Figure 9 A structural schematic diagram of an electronic device suitable for executing an instruction processing method provided by an embodiment of the present disclosure. Detailed implementation manners
[0106] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0107] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0108] Figure 1 An exemplary system architecture 100 is shown in which embodiments of the instruction processing method, apparatus, electronic device, and computer-readable storage medium of the present disclosure can be applied.
[0109] As Figure 1 shown, the system architecture 100 may include a user 101, various forms of terminal devices 102, and a processor 1021 that is a main functional component in the terminal device 102. The user 101 can interact with an application or program running on the terminal device 102 using the information input and output components that make up the terminal device 102, or can use the terminal device 102 as a transfer station to interact with a remote server through a network.
[0110] The terminal device 102 and the server can be hardware or software. When the terminal device 102 is hardware, it can be various electronic devices with a display screen, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and the like; when the terminal device 102 is software, it can be installed in the above-listed electronic devices, and it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited herein. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited herein.
[0111] Whether it is the terminal device 102 or the remote server, when processing the received request, it is processed by its internal processor. Taking the processor 1021 that constitutes the terminal device 102 as an example, it is usually composed of multiple functional units including memory, cache, controller, arithmetic logic unit (i.e., ALU), and other functional units. Among them, the ALU is mainly used to process the operation instructions including arithmetic operations and logical operations received and split by the processor 1021. Nowadays, the ALU usually adopts a multi-stage pipeline mechanism to process instructions. The following will explain the subsequent solutions for the ALU of the multi-stage pipeline specifically composed of multiple sequentially arranged data reading levels, multiple data calculation levels, and one data write-back level:
[0112] Assume that for two instructions with a dependency relationship successively issued by the processor 1021 to the ALU (for example, the subsequent instruction depends on the processing result of the previous instruction), the controller used as the execution entity can control the ALU to process as follows: First, when the first instruction has been executed to the data write-back level in the multi-stage pipeline, determine the current processing level to which the second instruction that enters the multi-stage pipeline later and has a dependency relationship with the first instruction is executed; Then, when the current processing level is the pre-processing level of the target processing level, store the result data that should be written back at the data write-back level into the cache space corresponding to the second instruction, and the target processing level is determined based on the instruction processing logic of the second instruction or the dependency relationship; Next, when the second instruction is subsequently executed to the target processing level in the multi-stage pipeline, transfer the result data stored in the cache space to the target processing level, so that the second instruction can directly obtain the result data at the target processing level for use by subsequent processing levels.
[0113] It should be understood that Figure 1 the number of users, terminal devices, processors, and ALUs in
[0114] Please refer to Figure 2 , Figure 2 which is a flowchart of an instruction processing method provided by an embodiment of the present disclosure. The process 200 includes the following steps:
[0115] Step 201: In response to the first instruction being executed to the data write-back level in the multi-stage pipeline of the arithmetic logic unit, determine the current processing level to which the second instruction having a dependency relationship with the first instruction is executed in the multi-stage pipeline.
[0116] This step aims to be performed by the execution entity of the instruction processing method (such as Figure 1The controller in the processor 1021 shown determines, when it is confirmed that the first instruction that first enters the ALU adopting a multi-stage pipeline mechanism has been executed to the data write-back stage (the multiple data processing stages constituting the multi-stage pipeline include multiple data read stages, multiple data calculation stages, and one data write-back stage arranged in sequence, and the data write-back stage is used to write back the result calculated by the data calculation stage to the target location), to which data processing stage (for the convenience of reference, it will be referred to as the current processing stage hereinafter) the second instruction that enters the multi-stage pipeline after the first instruction and has a dependency relationship with the first instruction is currently being executed. A schematic structural diagram of the multi-stage pipeline of an exemplary ALU adopting a multi-stage pipeline mechanism can be seen in Figure 3 , it can be seen that in addition to specifically including 3 data read stages, 4 data calculation stages, and 1 data write-back stage, there is also an instruction parsing stage for parsing the instructions entering the ALU before the first data read stage (of course, this instruction parsing stage may not be included or may not appear in the multi-stage pipeline structure).
[0117] It should be clear that there may be more than one second instruction having a dependency relationship with the first instruction, and the timing of multiple second instructions entering the multi-stage pipeline may also be different. Further, there may also be more than one first instruction having a dependency relationship with the second instruction, that is, there is a situation where a unique second instruction depends on the result data respectively calculated by multiple first instructions.
[0118] Among them, the dependency relationships between instructions specifically include the following three types:
[0119] 1) Read-after-write (WAR) dependency, which refers to the situation where an instruction attempts to write to a register after another instruction reads the same register. This situation may lead to incorrect or uncertain results. For example, there are two instructions. Instruction 1 writes the data in address a to register b, and Instruction 2 writes the content in register b back to address a. At this time, if Instruction 2 is executed in advance, then uncertain content will be written to address a, resulting in the content of register b being incorrect. However, if sequential execution is required, there will be a blockage, reducing the execution efficiency;
[0120] 2) Write-after-write (WAW) dependency, which means that both instructions write to a register. However, it is possible that the result of the first instruction is depended on by other instructions, but it is overwritten by the result of the second instruction before it is read, which may cause incorrect results. If sequential execution is required, it may be necessary to specifically wait for all instructions depending on the first instruction to finish executing before the second instruction can be executed, which will reduce the efficiency.
[0121] 3) Read After Write (RAW) refers to a situation in a pipelined processor where the current instruction needs to read the value of a register that was written by a previous instruction. This data dependency can cause problems because if the current instruction needs to read the data before the previous instruction has written it, then it may read incorrect data or data that has not been updated yet.
[0122] Considering that WAR and WAW can be easily resolved through register renaming, while the resolution of RAW depends on the bus. Therefore, although WAR and WAW can also be solved using the core inventive concept of the solution provided in this application, the solution provided in this application is mainly used to solve the problem of low instruction processing efficiency existing in RAW-type instruction dependencies (i.e., the data calculation indicated by the second instruction depends on the result data obtained by the first instruction through calculation).
[0123] Specifically, for the recognition of the above-mentioned dependencies, according to the instruction association information passed into the ALU along with the second instruction by the above-mentioned execution entity, the instruction association information can be used to determine which instruction existing in the multi-stage pipeline of the ALU has a dependency relationship with the second instruction; it can also be, under the multi-stage pipeline structure as Figure 3 shown, the above-mentioned execution entity parses the second instruction at the instruction parsing stage of the multi-stage pipeline, and determines which instruction existing in the ALU multi-stage pipeline has a dependency relationship with the second instruction according to the obtained instruction parsing result.
[0124] Step 202: In response to the current processing stage being the preprocessing stage of the target processing stage, store the result data that should be written back by the data write-back stage into the cache space corresponding to the second instruction;
[0125] Based on step 201, this step aims to have the above-mentioned execution entity discover that the current processing stage is specifically the preprocessing stage of the target processing stage in the second instruction (i.e., each data processing stage that executes before the target processing stage in the multi-stage pipeline. Still taking the specific structure of the multi-stage pipeline as Figure 3 shown as an example, assuming that the target processing stage is the first data calculation stage at the 5th level, then the first 4 levels are all preprocessing stages of the target processing stage. That is, no matter which level among the first 4 levels, it belongs to the situation corresponding to this step), and store the result data that should be written back by the data write-back stage into the cache space pre-allocated or created for the second instruction.
[0126] Among them, the target processing stage is determined based on the instruction processing logic of the second instruction or the dependency relationship, and should be one of multiple data reading stages or multiple data calculation stages. That is to say, at which data processing stage of the second instruction are some of the reading results or calculation results of the first instruction required, and this data processing stage is the target processing stage. Also, the data processing stage of the write-back stage for matching data is determined as the target processing stage according to the dependency relationship. For ease of understanding, an example is given here:
[0127] Suppose the first instruction is used to complete the calculation logic of the right-side expression: R[3] = R[0] + R[1], and the second instruction is used to complete the calculation logic of the right-side expression: R[8] = R[3] + R[6]. It can be seen from this example that the dependency relationship between the second instruction and the first instruction is write-after-read related, that is, the calculation result R[3] of the first instruction is used as a source operand for the second instruction to calculate R[8]. Then, when the first instruction is executed to the data write-back stage and passes the result R[3] to the second instruction, since the above calculation logic is only a simple addition operation between two calculation objects, the data processing stage that matches this write-after-read related can be either a certain data reading stage in the data reading stage or a certain data calculation stage in the data calculation stage after the data reading stage. Of course, when the calculation logic is more complex, there may not be multiple alternative target processing stages.
[0128] Among them, this cache space is a storage space set by this application to better transfer the result data generated by the first instruction to the target processing stage for the second instruction to use. Since it is the second instruction that needs to use it when being executed to the target processing stage, this cache space should correspond to the second instruction, rather than the first instruction. It should be understood that if this cache space corresponds to the first instruction, then when the first instruction is executed to the data write-back stage and should normally complete this instruction, in order to maintain the availability of the cache space corresponding to the first instruction and wait for the second instruction to be executed to the target processing stage, the first instruction still needs to continue to exist or extend for several execution cycles until the second instruction is executed to the target processing stage, which is obviously inappropriate. Therefore, the result data written back by the data write-back stage should be stored in the cache space corresponding to the second instruction, so as not to delay the normal end of the first instruction nor the subsequent use of the result data by the second instruction during its execution process.
[0129] Specifically, in combination with Figure 1From the perspective of the internal structure division of the processor 1021, this cache space can come from the cache, memory, or any other part that can perform logical processing and calls. The manifestation form of this cache space and the corresponding relationship with the second instruction can also be flexibly selected, as long as it can satisfy the temporary storage of result data and can issue the result data to the target processing stage or make it available for the target processing stage to read at an appropriate time. There are no excessive restrictions on other aspects.
[0130] In terms of practical implementation, to store the result data in the cache space corresponding to the second instruction, the result data can also be stored in the cache space corresponding to the second instruction in the preset data processing stage of the multi-stage pipeline by setting a first data bypass at the hardware level of the arithmetic logic unit (ALU). That is, the first data bypass exists as an entity data transmission path connecting the data write-back stage and the preset data processing stage. Taking the read-after-write dependency as an example, the target processing stage can be the first data calculation stage. Correspondingly, the preset data processing stage can be the last data reading stage or the stage before the first data calculation stage.
[0131] Step 203: In response to the second instruction being executed to the target processing stage in the multi-stage pipeline, transfer the result data stored in the cache space to the target processing stage.
[0132] Based on step 202, this step aims to have the above-mentioned execution entity transfer the result data stored in the cache space to the target processing stage when it discovers that the second instruction is subsequently executed to the target processing stage in the multi-stage pipeline. Among them, the transfer of the result data can specifically include two forms. One is to actively issue the result data stored in the cache space to the target processing stage under the control of the above-mentioned execution entity. The other is to make the second instruction actively read the result data stored in the cache space when it is executed to the target processing stage under the control of the above-mentioned execution entity. Which method to use can be flexibly selected according to the actual situation and is not specifically limited here.
[0133] In contrast to the situation provided in step 202: If the current processing stage is the target processing stage, the above-mentioned execution entity can directly transfer the result data that the data write-back stage should write back to the target processing stage, that is, it can transfer the result data to the target processing stage within the same execution cycle without using the cache space according to the solution provided by the prior art.
[0134] That is, in terms of practical implementation, the result data that the data write-back stage should write back can be directly transferred to the target processing stage through a second data bypass set at the hardware level of the arithmetic logic unit. That is, the second data bypass exists as an entity data transmission path connecting the data write-back stage and the target processing stage.
[0135] It should be noted that there are various reasons that may cause two instructions with a dependency relationship to not be strictly aligned in execution cycles. For example, instruction scheduling delays may result in the ALU receiving the second instruction relatively late, or instruction read data conflicts may cause the second instruction to take several cycles of delay to read all source operands from the storage space, thereby causing the second instruction to start execution later than the specified time and miss the cycle when the bypass data is valid (that is, in the case of only the second data bypass, it is necessary to require the data write-back stage and the target processing stage to execute in the same execution cycle to make the bypass data valid). In the existing solution, once the second instruction misses the cycle when it can use the bypass data, to solve the above problem, the second instruction needs to be delayed by several cycles until the first instruction completes data write-back before the second instruction can start execution. That is to say, the start time of the second instruction must be precisely arranged, resulting in increased scheduling difficulty. And once the second instruction misses the cycle of early scheduling, it cannot use the bypass (i.e., the second data bypass) and must wait for the first instruction to complete execution in a non-bypass manner before it can execute. When executing complex code, the possibility that the second instruction misses the scheduling cycle for using the bypass is very high, which will cause the bypass design to not play its role as expected.
[0136] Therefore, in view of the problems existing in the above prior art, in the instruction processing method provided in this embodiment, when the first instruction has been executed to the final data write-back stage, by storing the result data to be written back in the cache space pre-created or allocated for the second instruction, the second instruction can obtain it from the cache space at any time when it is executed to the target processing stage that needs to use the result data, without strictly adhering to the requirement that the target processing stage and the data write-back stage must be in the same execution cycle to transfer the result data. This is equivalent to relaxing the timing for the second instruction to enter the multi-stage pipeline and start execution, optimizing, expanding, and improving the existing data bypass technology, and further enhancing the execution efficiency and accuracy of the processor for multiple instructions with a dependency relationship.
[0137] It should be noted that in the case where the dependency relationship is write-after-read related and the number of levels of this data calculation stage is M (a positive integer not less than 1), it is also necessary to ensure that the second instruction must enter the multi-stage pipeline after the Mth execution cycle when the first instruction enters the multi-stage pipeline to normally trigger the above solution and thus ensure the improvement of the instruction execution efficiency.
[0138] To further strengthen the understanding of the solution provided in the above embodiment, this application also Figure 4 provides a flowchart of another instruction processing method, where process 400 includes the following steps.
[0139] Step 401: Create or allocate a first cache space for the first instruction that first enters the multi-stage pipeline, and create or allocate a second cache space for the second instruction that later enters the multi-stage pipeline.
[0140] In this embodiment, this step aims to reflect that the above-mentioned execution entity allocates or creates a dedicated cache space corresponding to each instruction entering the multi-stage pipeline. In this embodiment, it will be specifically manifested as: creating or allocating a first cache space for the first instruction that first enters the multi-stage pipeline, and creating or allocating a second cache space for the second instruction that later enters the multi-stage pipeline. It should be understood that the cache space in this application is at least used to temporarily store the instruction execution products from the prior instructions that may be used by the corresponding instruction later. In this embodiment, the second cache space is used to temporarily store the result data that should be written back by the first instruction at the data write-back stage and will be used when the second instruction is executed to the target processing stage.
[0141] Step 402: Temporarily store the source operands obtained when the first instruction is executed to the data reading stage into the first cache space;
[0142] That is, this step describes that when the first instruction has not been executed to the data write-back stage but only to the data reading stage, the first cache space can also temporarily store the source operands obtained when it is executed to the data reading stage for use when it is later executed to the data calculation stage.
[0143] Step 403: In response to the successful acquisition of the source operands when the first instruction is executed to the data calculation stage, release the first cache space;
[0144] Based on step 402, since the first cache space is also used to temporarily store the source operands, and all the source operands should be completely fetched when executed to the first data calculation stage, and the first instruction has no dependency relationship with other prior instructions. Therefore, when the source operands stored in the first cache space are successfully acquired when the first instruction is executed to the data calculation stage, the first cache space that has lost its function can be released, so that the first instruction can be normally completed without an associated or bound cache space after being executed to the data write-back stage, thus fundamentally avoiding the situation where the execution of the corresponding instruction is forced to be extended by several execution cycles because the cache space has not been released after the corresponding instruction has completed the last data write-back stage.
[0145] The "release" described in this step mainly refers to disconnecting or clearing the association or correspondence between the first cache space and the first instruction. Specifically, the association or correspondence can be cleared by initializing the first cache space to make it eligible for reallocation to subsequent new instructions, or by more thoroughly destroying the first cache space. Which method to choose can be flexibly determined according to the actual situation and is not specifically limited here.
[0146] Step 404: In response to the first instruction reaching the data write-back stage, determine the current execution stage of the second instruction that has a dependency relationship with the first instruction.
[0147] Step 405: In response to the current execution stage being the pre-processing stage of the target processing stage in the second instruction, store the result data that should be written back at the data write-back stage into the second cache space.
[0148] Furthermore, when the result data that should be written back at the data write-back stage has been stored in the cache space corresponding to the second instruction, a tag indicating that valid data is stored can also be attached to the cache space, so as to quickly determine whether the cache space stores the result data required by the second instruction using this tag, preventing duplicate data storage and resource waste.
[0149] Step 406: In response to the second instruction subsequently reaching the target processing stage in the multi-stage pipeline, obtain the result data from the second cache space.
[0150] Steps 404, 405, and 406 respectively correspond to Figure 2 Steps 201, 202, and 203 in Process 200, with the only difference being that through the additional Step 401, it is clear that corresponding cache spaces are created or allocated for the first instruction and the second instruction when they enter the multi-stage pipeline, and specifically, in Step 404, the implementation method is used in which the above-mentioned execution entity controls the second instruction to actively obtain the result data stored in the second cache space. For the same parts, refer to the relevant elaborations in Steps 201, 202, and 203 and will not be elaborated here.
[0151] Step 405: In response to the result data being successfully obtained by the target processing stage, release the second cache space.
[0152] Based on Step 404, the purpose of this step is that when the result data is successfully obtained by the target processing stage by the above-mentioned execution entity, since the second cache space has also completed its most important function (temporarily storing all the data required when being executed to the data calculation stage), it also meets the condition of being "released". For the detailed description of the relevant release methods, refer to the description under Step 403.
[0153] Further, before releasing the second cache space, the tag attached to the cache space can be cleared first, so as to reflect that the result data required by the second instruction is no longer stored in the cache space by clearing the tag. Of course, the clearing of the tag can be used as an intermediate means for the release operation of the second cache space, or can be overwritten by the effect after the release operation.
[0154] In the above solution provided in this embodiment, there is no necessary causal and dependency relationship among the solution for the creation or allocation timing of the cache space provided in step 401, the usage and release solutions of the first cache space provided in steps 402 - 403, and the result data acquisition and second cache space release solutions provided in steps 404 and 405. They can be combined with the basic embodiment provided by process 200 separately to form different embodiments. This embodiment only exists as a preferred embodiment that simultaneously includes the above preferred solutions.
[0155] Different from Figure 4 the cache space creation / assignment and usage solutions provided for the first instruction and the second instruction with dependency relationships in the embodiment, this embodiment also provides a set of similar cache creation / assignment and usage solutions for independent instructions regardless of whether there are dependency relationships with other instructions. The following steps can be referred to:
[0156] First, in response to the third instruction entering the multi - stage pipeline, create or allocate a corresponding cache space for the third instruction; then, temporarily store the source operands obtained when the third instruction is executed to the data reading stage in the cache space corresponding to the third instruction; finally, in response to the successful acquisition of the source operands when the third instruction is executed to the data calculation stage, release the cache space corresponding to the third instruction.
[0157] That is, in this embodiment, the third instruction is an independent instruction that is not limited to whether there are dependency relationships with other instructions different from the first instruction and the second instruction. That is, this embodiment also provides a similar cache space creation / assignment and usage solution for such instructions, so that such instructions can also temporarily store source operands through the cache space for subsequent convenient use.
[0158] To further deepen the understanding of the form in which the cache space can exist, this embodiment also takes the number of levels of the data reading stage as N as an example, and provides a specific solution for determining the cache space through Figure 5 process 500, where N is a positive integer not less than 1, and includes the following steps:
[0159] Step 501: Determine the target number of cache spaces according to the number of source operands to be read by the second instruction through N data reading levels;
[0160] Assume that the number of source operands that the second instruction wants to read through N data read levels is X (a positive integer), and the target number of cache spaces can only be selected between 1 and X. That is, when the target number is selected as 1, all the read source operands are stored in the only cache space for temporary storage; if the target number is the number of read source operands - X, it means that the different read source operands are stored independently in different cache spaces.
[0161] Step 502: Determine the target depth of the cache space according to the number of levels of the N data read levels;
[0162] Based on Step 501, this step aims to determine the target depth of the cache space by the above-mentioned execution entity according to the number of levels of the data read levels. For example, the number of levels of the data read levels - N can be used as the number of depth levels of the target depth. Taking Figure 3 the example presented in the example that includes 3 data read levels, the cache space can also be specifically set to have a cache space with 3 levels of depth.
[0163] Step 503: Create or allocate cache spaces for the second instruction with the target number and the target depth.
[0164] Based on Step 501 and Step 502, this step aims to create or allocate cache spaces for the second instruction by the above-mentioned execution entity with the target number and the target depth. Specifically, at the data return level served by the last level among the N data read levels, cache spaces with the target number and the target depth can be created for the second instruction, that is, the cache space is set at the data return level served by the last level among the N data read levels, so as to store the data read at the corresponding level into the corresponding depth level through the depth levels equal to the number of data read levels as the number of data read levels increases.
[0165] Among them, as the solution provided in this embodiment, the cache space can be specifically represented as a first-in-first-out queue with a queue length of N (that is, the number of levels of the data read levels), and in this form of representation, the first-in-first-out queue can also be used to temporarily store the source operands obtained through the N data read levels (that is, corresponding Figure 4(the solution provided in step 402 of the embodiment). Correspondingly, in the case of specifically adopting the form of a first-in, first-out queue as the cache space, processes 202 in the above process 200 and process 406 in process 400 can be adaptively adjusted as follows: write the result data that should be written back to the corresponding write-back level into the corresponding position of the first-in, first-out queue according to the polling rule. Still taking the first-in, first-out queue with a queue length of 3 as an example, its polling rule can be specifically manifested as continuously polling in the order of 0, 1, 2. When storing the result data, it is only necessary to determine whether the current corresponding one according to the polling rule is 0, 1, or 2, so as to store the result data in the corresponding queue position.
[0166] To deepen the understanding of the specific performance of the solutions provided in the above embodiments at the hardware level, this embodiment also provides an arithmetic logic unit, including:
[0167] A multi-stage pipeline structure, including a plurality of data reading stages, a plurality of data calculation stages, and a data write-back stage arranged in sequence;
[0168] A cache controller, configured to, in response to the first instruction being executed to the data write-back stage, determine the current processing stage to which the second instruction having a dependency relationship with the first instruction is executed; in response to the current processing stage being the pre-processing stage of the target processing stage, store the result data that should be written back to the data write-back stage into the cache space corresponding to the second instruction; in response to the second instruction being executed to the target processing stage, transfer the result data stored in the cache space to the target processing stage, and the target processing stage is determined based on the instruction processing logic of the current instruction or the dependency relationship with the previous instruction in the arithmetic logic unit, and is one of the plurality of data reading stages or the plurality of data calculation stages.
[0169] Further, the arithmetic logic unit may further include:
[0170] A first data bypass, which is connected between the data write-back stage and the pre-processing stage;
[0171] The cache controller is specifically configured to: in response to the current processing stage being the pre-processing stage, store the result data that should be written back to the data write-back stage into the cache space corresponding to the second instruction in the pre-processing stage through the first data bypass.
[0172] Further, the arithmetic logic unit may further include:
[0173] A second data bypass, which is connected between the data write-back stage and the target processing stage;
[0174] The cache controller is further configured to: in response to the current processing stage being the target processing stage, directly transfer the result data that should be written back to the data write-back stage to the target processing stage through the second data bypass.
[0175] Specifically, the first data bypass and the second data bypass may include: wires and data transmission devices on the wires.
[0176] Still taking Figure 3 the arithmetic logic unit with the shown multi-stage pipeline structure as an example, Figure 6 In Figure 3 this case, a specific implementation manner of the first data bypass and the second data bypass is shown, that is, Figure 6 it is based on the target processing stage being the first data calculation stage (i.e., calculation 1). At this time, the first data bypass is connected between the data write-back stage and the data return stage, and the second data bypass is connected between the data write-back stage and calculation 1.
[0177] To further deepen the understanding of the details of the solution provided by this application, a complete set of exemplary implementation manners will be given below in combination with a specific example:
[0178] Assume that the multi-stage pipeline mechanism of the ALU has a total of N + M levels (i.e., N data read levels and M data calculation levels). The first level is the time point for initiating a data read request, the Nth level is the data return time point, and the N + Mth level is the result write-back time point. That is, an instruction entering the ALU multi-stage pipeline reaches the data return time point after N - 1 cycles after initiating the data read request, and then reaches the data write-back time point after M cycles. When the instruction enters the pipeline, it needs to carry necessary information such as the result write address and the source operand address, and the result write address needs to follow the pipeline from the first level to the last level until it is discarded after the result is written. The cycle of the entire ALU pipeline is as Figure 7-1 shown.
[0179] In this embodiment, a storage space that can cache X data is set at the data return level (i.e., the Nth level), and 2 bypasses are set. One bypass is connected from the data write-back stage to the data return stage (the Nth level), and the other bypass is connected from the data write-back stage to the next stage of the data return (the N + 1th level). And a cache control logic is set to control the read and write operations of the cache. When the instruction enters the first level of the pipeline, the cache control logic will allocate an address for the instruction. When the instruction or the data of the instruction enters the Nth level, the data will be stored in the storage space allocated for the instruction. When the instruction enters the N + 1th level, the cache control logic will take out the data corresponding to the address of the instruction for calculation.
[0180] For two dependent instructions, assuming they are Instruction A and Instruction B in sequence, the interval for entering the pipeline must be greater than or equal to M - 1 cycles. If Instruction B enters the pipeline in the (M - 1)-th cycle after Instruction A enters the pipeline, it triggers the behavior of Instruction A passing data to Instruction B from the (N + M)-th stage to the (N + 1)-th stage through bypass. The specific detection method is as follows: When Instruction B enters the first pipeline stage (if Instruction B stays in the first pipeline stage, it is based on the moment of leaving the first pipeline stage), detect the relationship between Instruction B and the instructions on the M-th pipeline stage. If there is a valid instruction (Instruction A) on the M-th pipeline stage and the result address of this instruction is the same as the source operand address of Instruction B, then set a label for Instruction B to mark that Instruction B needs to use the result of Instruction A through bypass at the (N + 1)-th stage. After Instruction B enters the (N + 1)-th pipeline stage, detect whether this label is valid. If it is valid, trigger the bypass mechanism of the (N + 1)-th pipeline stage. As Figure 7-2 shown: Assume N = 3 and M = 5. Then Instruction B is allowed to enter the pipeline in the 4th cycle after Instruction A enters the pipeline. The path working on the processor hardware is as Figure 7-3 shown.
[0181] If the moment when Instruction B enters the pipeline is after the (M - 1)-th cycle after Instruction A enters the pipeline and before the (N + M)-th cycle after A enters the pipeline, it triggers the behavior of Instruction A passing data to the storage space of the N-th stage of the pipeline from the (N + M)-th stage through bypass (i.e., the bypass presented by the thick black line in Figure 7-3 ). During the period when Instruction B is in the 1st to N-th stages of the pipeline, the cache control logic needs to detect whether the result to be written at the (N + M)-th stage is the same as the source operand address of Instruction B. If they are the same, write the calculation result of the (N + M)-th stage into the storage space allocated to Instruction B. When Instruction B executes to the (N + 1)-th stage, the cache control logic fetches the data in the space corresponding to Instruction B for calculation.
[0182] The specific approach is as follows: When any instruction enters the (N + M)-th pipeline stage, detect whether the result address of this instruction is the same as the source operand addresses of the instructions on the 1st, 2nd... N-th stages of the pipeline. If the result address of the instruction on the (N + M)-th stage is the same as the address of one or more stages of the instructions on the 1st, 2nd... N-th stages of the pipeline (assuming it is Instruction B), trigger the bypass mechanism and write the calculation result on the (N + M)-th pipeline stage into the data cache space allocated for Instruction B to use through bypass.
[0183] Data cache space allocation method: When an instruction enters the first stage of the pipeline, a data cache space is allocated for this instruction. When the instruction leaves the Nth stage of the pipeline, the data cache space of this instruction is released. As long as there is no conflict with the data cache space that is being occupied when allocating a new data cache space, it is okay. A simple allocation method is to set the data cache space to N addresses (the addresses are 0, 1, 2, ……, n - 1). When the first instruction enters the pipeline, address 0 is allocated to the first instruction. When the second instruction enters the pipeline, address 1 is allocated to the second instruction. Each time an allocation is made, the address is incremented by 1. When the address is incremented to N - 1, the next allocation loops back to address 0 to start the allocation. The working path in hardware is as Figure 7-4 shown (that is, Figure 7-4 the bypass presented by the thick black line in
[0184] If an instruction has no dependency relationship with the instructions in the pipeline when it enters the pipeline, the cache control logic will still allocate a data storage address for it. When the data read by this instruction returns to the pipeline, it will be temporarily stored in the storage space of this instruction and wait to be used by the calculation logic at the (N + 1)th stage in the next cycle.
[0185] The following introduces an ALU pipeline case implemented according to the above scheme.
[0186] Assume that the ALU execution can only perform addition and multiplication operations and requires 2 source operands, namely src0 and src1. There is an instruction scheduling module outside the ALU pipeline, which is responsible for sending calculation instructions to the ALU. The ALU pipeline sends requests to read src0 and src1 to the storage space at the first stage. If the arbitration for reading data from the storage space is successfully passed, the instruction gets the data returned by the storage space at the third stage. Otherwise, the instruction will stay at the first stage and repeatedly send read requests to the storage space. After receiving the data at the third stage, the instruction continues to be passed backward, completes the calculation through the fourth to seventh stages, and writes the result to the storage space at the eighth stage. Place 2 data cache spaces with a depth of 3 at the third stage of this ALU pipeline, which cache the data of src0 and src1 respectively. Each cache space can cache 3 pieces of data, and place a cache control logic. When an instruction enters the first stage of the ALU pipeline, a data cache address is allocated for this instruction.
[0187] Assume that instruction A and instruction B are two instructions with a dependency relationship. The cycle relationship between instruction B and instruction A can be divided into 3 types:
[0188] The first case is that the scheduling module sends instruction B to the first stage of the ALU pipeline in the fourth cycle after instruction A leaves the first stage of the ALU pipeline, and instruction B successfully passes the arbitration of the storage space when it is in the first stage, completing the reading of the source operands. In this case, when instruction B enters the fourth stage, it obtains the calculation result of instruction A in advance through the bypass from the eighth stage to the fourth stage and replaces the corresponding source operands, as Figure 7-2 shown, and the hardware working path is as Figure 7-5 shown. At this time, specifically through Figure 7-5 the currently used bypass is presented by the thick black line in
[0189] The second case is that the time when instruction B leaves the first-stage pipeline of the ALU is after the fourth cycle and before the eighth cycle when instruction A leaves the first-stage pipeline. In this case, when instruction A is at the eighth stage, instruction B may be at the first stage, the second stage, or the third stage (please refer to the schematic diagram shown in Figure 7-6 ). When instruction A is at the eighth stage, the cache control logic checks whether the result of instruction A has a dependency relationship with one or more of the instructions at the first stage, the second stage, or the third stage. If there is a dependency relationship, the calculation result of instruction A is obtained in advance through the bypass from the eighth stage to the third stage, and the calculation result of instruction A is stored in the data cache space in advance according to the address allocated by the cache control logic for each stage of the instruction (please refer to Figure 7-7 , and the currently used bypass is presented by the thick black line in the figure). If the result of instruction A is written into the data cache space of the first-stage or second-stage instruction, the cache control logic needs to record the flag indicating that the data stored at this address is valid. When the corresponding instruction enters the third stage, the data at this address does not need to be rewritten with the data returned from the storage space. When this instruction leaves the third stage, the cache control logic needs to clear the flag indicating that the stored data is valid;
[0190] The third case is that the time when instruction B leaves the first-stage pipeline of the ALU is at the eighth cycle or later when instruction A leaves the first-stage pipeline. At this time, the result of instruction A has completed the write-back operation to the data storage space, and the required data can be directly read from the storage space without using the bypass.
[0191] It can be seen from the above solutions provided by this embodiment that through the bypass design provided by the above embodiment, the execution interval between two instructions with a dependency relationship is the necessary minimum cycle of the calculation result. At any cycle after the necessary interval cycle, the second instruction can enter the pipeline for execution without waiting due to missing the bypass cycle, thereby improving the instruction processing efficiency.
[0192] For further reference, please refer to Figure 8, as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an instruction processing apparatus, which corresponds to the method embodiment shown in Figure 2 and can be specifically applied to various electronic devices.
[0193] As shown in Figure 8 , the instruction processing apparatus 800 of this embodiment may include: a current execution stage determination module 801, a result data caching module 802, and a result data transfer module 803. Among them, the current processing stage determination module 801 is configured to determine the current processing stage at which a second instruction that has a dependency relationship with the first instruction is executed in the multi-stage pipeline in response to the first instruction being executed to the data write-back stage in the arithmetic logic unit's multi-stage pipeline. The multiple data processing stages that make up the multi-stage pipeline include a plurality of data reading stages, a plurality of data calculation stages, and a data write-back stage arranged in sequence; the result data caching module 802 is configured to store the result data that should be written back at the data write-back stage into the cache space corresponding to the second instruction in response to the previous processing stage of the current processing stage being the target processing stage. The target processing stage is determined based on the instruction processing logic or dependency relationship of the second instruction and is one of the plurality of data reading stages or the plurality of data calculation stages; the result data transfer module 803 is configured to transfer the result data stored in the cache space to the target processing stage in response to the second instruction being executed to the target processing stage in the multi-stage pipeline.
[0194] In this embodiment, in the instruction processing apparatus 800: the specific processing of the current execution stage determination module 801, the result data caching module 802, and the result data transfer module 803 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201-203 in the corresponding embodiment, which will not be elaborated here.
[0195] In some other implementation manners of this embodiment, the instruction processing apparatus 800 may further include:
[0196] a first dependency relationship determination module, configured to determine that the second instruction has a dependency relationship with the first instruction according to the instruction association information passed in together with the second instruction; or
[0197] a second dependency relationship determination module, configured to perform instruction parsing on the second instruction at the instruction parsing stage of the multi-stage pipeline and determine that the second instruction has a dependency relationship with the first instruction according to the obtained instruction parsing result. The multi-stage pipeline further includes an instruction parsing stage located before the first data reading stage.
[0198] In some other implementation manners of this embodiment, the result data caching module 803 may be further configured to:
[0199] The result data is stored in the cache space corresponding to the second instruction in a preset data processing stage of the multi-stage pipeline through a first data bypass. The first data bypass exists as a data transmission path connecting the data write-back stage and the preset data processing stage.
[0200] In some other implementation manners of this embodiment, the dependency relationship is read-after-write dependency, the target processing stage is the first data calculation stage, and the preset data processing stage is the last data reading stage or the stage preceding the first data calculation stage.
[0201] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0202] A first cache space processing module, configured to create or allocate a corresponding cache space for the second instruction in response to the second instruction entering the multi-stage pipeline;
[0203] A first release module, configured to release the cache space in response to the result data being successfully acquired by the target processing stage.
[0204] In some other implementation manners of this embodiment, the number of levels of the data reading stage is N, where N is a positive integer not less than 1. The first cache space processing module is further configured to:
[0205] Determine the target number of cache spaces according to the number of source operands to be read by the second instruction through N data reading stages;
[0206] Determine the target depth of the cache space according to the number of levels of N data reading stages;
[0207] Create or allocate cache spaces for the second instruction with the target number and the target depth.
[0208] In some other implementation manners of this embodiment, the target number is the number of source operands to be read by the second instruction through N data reading stages; wherein, different source operands are independently stored in different cache spaces.
[0209] In some other implementation manners of this embodiment, the target depth is N.
[0210] In some other implementation manners of this embodiment, the cache space includes a first-in-first-out queue with a queue length of N.
[0211] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0212] A first source operand temporary storage module, configured to temporarily store the source operands acquired through N data reading stages by using the first-in-first-out queue.
[0213] In some other implementation manners of this embodiment, the result data caching module 803 may be further configured to:
[0214] Store the result data that should be written back by the data write-back stage into the corresponding position of the first selected queue according to the polling rule.
[0215] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0216] A marking addition module, configured to add a mark storing valid data to the cache space in response to that the result data that should be written back by the data write-back stage has been stored in the cache space corresponding to the second instruction.
[0217] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0218] A marking clearing module, configured to clear the mark added to the cache space in response to that the result data is successfully acquired by the target processing stage.
[0219] In some other implementation manners of this embodiment, the dependency is read-after-write related, the number of levels of the data calculation stage is M, M is a positive integer not less than 1, and the second instruction enters the multi-stage pipeline after the Mth execution cycle when the first instruction enters the multi-stage pipeline.
[0220] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0221] A second cache space processing module, configured to create or allocate a corresponding cache space for the first instruction in response to that the first instruction enters the multi-stage pipeline;
[0222] A second source operand temporary storage module, configured to temporarily store the source operand obtained when the first instruction is executed to the data reading stage into the cache space corresponding to the first instruction;
[0223] A second release module, configured to release the cache space corresponding to the first instruction in response to that the source operand is successfully acquired when the first instruction is executed to the data calculation stage.
[0224] In some other implementation manners of this embodiment, the instruction processing device 800 may further include:
[0225] A direct transfer module, configured to directly transfer the result data that should be written back by the data write-back stage to the target processing stage in response to that the current processing stage is the target processing stage.
[0226] In some other implementation manners of this embodiment, the direct transfer module is further configured to:
[0227] The result data to be written back to the write-back stage is directly passed to the target processing stage through the second data bypass. The second data bypass exists as a data transmission path connecting the write-back stage and the target processing stage.
[0228] This embodiment exists as a device embodiment corresponding to the above method embodiment. The instruction processing device provided in this embodiment, when the first instruction has been executed to the final write-back stage, stores the result data to be written back in the cache space pre-created or allocated for the second instruction, so that the second instruction can obtain it from the cache space at any time when it is executed to the target processing stage that needs to use the result data, without strictly adhering to the requirement that the target processing stage and the write-back stage must be in the same execution cycle to transfer the result data. This is equivalent to relaxing the timing for the second instruction to enter the multi-stage pipeline and start execution, optimizing, expanding, and improving the existing data bypass technology, and further enhancing the execution efficiency and accuracy of the processor for multiple instructions with dependency relationships.
[0229] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can implement the instruction processing method described in any of the above embodiments.
[0230] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions, and when the computer instructions are executed, the computer can implement the instruction processing method described in any of the above embodiments.
[0231] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which can implement the instruction processing method described in any of the above embodiments when executed by a processor.
[0232] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0233] As Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 902 or computer programs loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of device 900 can also be stored. The computing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0234] Multiple components in device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, optical disc, etc.; and a communication unit 909, such as a network card, modem, wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0235] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the instruction processing method. For example, in some embodiments, the instruction processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the instruction processing method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the instruction processing method in any other appropriate way (e.g., by means of firmware).
[0236] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0237] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0238] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0239] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0240] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0241] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.
[0242] According to the technical solution of the embodiment of the present disclosure, when the first instruction has been executed to the last data write-back stage, by storing the result data to be written back in the cache space pre-created or allocated for the second instruction, the second instruction can obtain it from the cache space at any time when it is executed to the target processing stage that needs to use the result data, without strictly following the requirement that the target processing stage must be in the same execution cycle as the data write-back stage to transfer the result data. This is equivalent to relaxing the timing for the second instruction to enter the multi-stage pipeline and start execution, thus optimizing, expanding, and improving the existing data bypass technology, and further enhancing the execution efficiency and accuracy of the processor for multiple instructions with dependency relationships.
[0243] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.
[0244] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A command processing method, characterized in that: include: In response to a first instruction being executed to a data write-back stage in a multi-stage pipeline of an arithmetic logic unit, determining a current processing stage to which a second instruction having a dependency relationship with the first instruction is executed in the multi-stage pipeline, wherein the plurality of data processing stages constituting the multi-stage pipeline include a plurality of data reading stages, a plurality of data calculation stages and one data write-back stage arranged in sequence; In response to the current processing stage being a pre-processing stage of a target processing stage, storing the result data to be written back by the data write-back stage into a cache space corresponding to the second instruction, the target processing stage is determined based on the instruction processing logic of the second instruction or the dependency relationship and is one of the plurality of data reading stages or the plurality of data computing stages, the instruction processing logic is an execution step of performing actual operations for the corresponding instruction, and the pre-processing stages are all data processing stages located before the target processing stage; In response to the second instruction being executed to the target processing stage in the multi-stage pipeline, the result data stored in the cache space is transferred to the target processing stage.
2. The method according to claim 1, characterized in that: The method further comprises: Determining, according to instruction association information transmitted along with the second instruction, that the second instruction has the dependency relationship with the first instruction; or The second instruction is parsed at the instruction parsing stage of the multi-stage pipeline, and based on the instruction parsing result obtained, it is determined that the second instruction has the dependency relationship with the first instruction. The multi-stage pipeline also includes an instruction parsing stage located before the first data reading stage.
3. The method according to claim 1, characterized in that The step of storing the result data to be written back by the data writing back stage into a cache space corresponding to the second instruction includes: The result data is stored in a cache space corresponding to the second instruction in a preset data processing stage of the multi-stage pipeline through a first data bypass, and the first data bypass exists as a data transmission path connecting the data write-back stage and the preset data processing stage.
4. The method according to claim 3, characterized in that The dependency relationship is a read-after-write dependency, the target processing stage is the first data calculation stage, and the preset data processing stage is the last data reading stage or the stage before the first data calculation stage.
5. The method according to claim 1, characterized in that The method further comprises: In response to the second instruction entering the multi-stage pipeline, creating or allocating a corresponding cache space for the second instruction; In response to the result data being successfully acquired by the target processing stage, the cache space is released.
6. The method according to claim 5, characterized in that The number of the data reading stages is N, where N is a positive integer not less than 1, and the step of creating or allocating a corresponding cache space for the second instruction includes: Determining a target number of cache spaces according to the number of source operands to be read by the second instruction through the N data reading stages; Determining a target depth of the cache space according to the number of the N data reading stages; A cache space having the target number and the target depth is created or allocated for the second instruction.
7. The method according to claim 6, characterized in that The target number is the number of source operands to be read by the second instruction through the N data reading stages; wherein different source operands are independently stored in different cache spaces.
8. The method according to claim 6, characterized in that The target depth is N.
9. The method according to claim 8, characterized in that The cache space includes a first-in-first-out queue with a queue length of N.
10. The method according to claim 9, characterized in that The method further comprises: The first-in first-out queue is used to temporarily store source operands obtained through the N data reading stages.
11. The method according to claim 9, characterized in that The step of storing the result data to be written back by the data writing back stage into a cache space corresponding to the second instruction includes: The result data to be written back by the data write-back stage is stored in the corresponding position of the first-in-out queue according to the polling rule.
12. The method according to claim 1, characterized in that The method further comprises: In response to the result data to be written back by the data write-back stage having been stored in the cache space corresponding to the second instruction, a mark in which valid data is stored is added to the cache space.
13. The method according to claim 12, characterized in that The method further comprises: In response to the result data being successfully acquired by the target processing stage, the tag attached to the cache space is cleared.
14. The method according to claim 1, characterized in that The dependency is read-after-write, the number of data calculation levels is M, M is a positive integer not less than 1, and the second instruction enters the multi-stage pipeline after the first instruction enters the Mth execution cycle of the multi-stage pipeline.
15. The method according to claim 1, characterized in that The method further comprises: In response to the first instruction entering the multi-stage pipeline, creating or allocating a corresponding cache space for the first instruction; temporarily storing the source operand obtained when the first instruction is executed to the data reading stage in a cache space corresponding to the first instruction; In response to the source operand being successfully acquired when the first instruction is executed to the data calculation stage, a cache space corresponding to the first instruction is released.
16. The method according to claim 1, characterized in that The method further comprises: In response to a third instruction entering the multi-stage pipeline, creating or allocating a corresponding cache space for the third instruction; temporarily storing the source operand obtained when the third instruction is executed to the data reading stage in a cache space corresponding to the third instruction; In response to the source operand being successfully acquired when the third instruction is executed to the data calculation stage, a cache space corresponding to the third instruction is released.
17. The method according to any one of claims 1 to 16, characterized in that: The method further comprises: In response to the current processing stage being the target processing stage, the result data to be written back by the data write-back stage is directly transferred to the target processing stage.
18. The method according to claim 17, characterized in that The step of directly transferring the result data to be written back by the data writing back stage to the target processing stage comprises: The result data to be written back by the data write-back stage is directly transferred to the target processing stage through a second data bypass, and the second data bypass exists as a data transmission path connecting the data write-back stage and the target processing stage.
19. An arithmetic logic unit, characterized in that: include: A multi-stage pipeline structure, including a plurality of data reading stages, a plurality of data calculation stages and a data write-back stage arranged in sequence; The cache controller is configured to, in response to a first instruction being executed to the data write-back stage, determine a current processing stage to which a second instruction having a dependency relationship with the first instruction is executed; in response to the current processing stage being a pre-processing stage of a target processing stage, store result data to be written back by the data write-back stage into a cache space corresponding to the second instruction; In response to the second instruction being executed to the target processing stage, the result data stored in the cache space is passed to the target processing stage, the target processing stage is determined based on the instruction processing logic of the current instruction or the dependency relationship with the previous instruction in the arithmetic logic unit, and is one of the multiple data reading stages or the multiple data calculation stages, the instruction processing logic is an execution step for the corresponding instruction to perform actual operations, and the pre-processing stage is all data processing stages located before the target processing stage.
20. The arithmetic logic unit according to claim 19, characterized in that: Also includes: a first data bypass connected between the data write-back stage and the pre-processing stage; The cache controller is specifically used for: in response to the current processing stage being the pre-processing stage, storing the result data to be written back by the data write-back stage into the cache space corresponding to the second instruction in the pre-processing stage through the first data bypass.
21. The arithmetic logic unit according to claim 20, characterized in that: The first data bypass includes: a first wire and a data transmission device on the first wire.
22. The arithmetic logic unit according to claim 19, characterized in that: Also includes: a second data bypass connected between the data write-back stage and the target processing stage; The cache controller is further configured to: in response to the current processing stage being the target processing stage, directly transfer the result data to be written back by the data write-back stage to the target processing stage through the second data bypass.
23. The arithmetic logic unit according to claim 22, characterized in that: The second data bypass includes: a second wire and a data transmission device on the second wire.
24. An instruction processing device, characterized in that: include: a current processing stage determination module, configured to determine, in response to a first instruction being executed to a data write-back stage in a multi-stage pipeline of an arithmetic logic unit, a current processing stage to which a second instruction having a dependency relationship with the first instruction is executed in the multi-stage pipeline, wherein the plurality of data processing stages constituting the multi-stage pipeline include a plurality of data reading stages, a plurality of data calculation stages and one data write-back stage arranged in sequence; A result data cache module is configured to store the result data to be written back by the data write-back stage into a cache space corresponding to the second instruction in response to the current processing stage being a pre-processing stage of a target processing stage, the target processing stage being determined based on the instruction processing logic of the second instruction or the dependency relationship and being one of the plurality of data reading stages or the plurality of data computing stages, the instruction processing logic being an execution step of performing actual operations for the corresponding instruction, and the pre-processing stage being all data processing stages located before the target processing stage; The result data transfer module is configured to transfer the result data stored in the cache space to the target processing stage in response to the second instruction being executed to the target processing stage in the multi-stage pipeline.
25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the instruction processing method according to any one of claims 1 to 18.
26. A non-volatile computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the instruction processing method according to any one of claims 1 to 18.
27. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the instruction processing method according to any one of claims 1 to 18 are implemented.
Citation Information
Patent Citations
High performance architecture for a writeback stage
US20070005941A1
Methods and systems for network flow tracing within a packet processing pipeline
US20230068914A1