Instruction execution device, chip, electronic equipment and method

By performing dependency detection in parallel during the pipeline execution of instructions, the problem of delayed loading instruction transmission time in the prior art is solved, and the instruction execution efficiency is improved.

CN120010924APending Publication Date: 2025-05-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311520394.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the loading instruction performs address dependency detection before transmission, resulting in a delay in the transmission time of the loading instruction, affecting the efficiency of instruction execution.

Method used

During the pipeline execution of instructions, the instructions are subject to dependency detection, so that the execution process of the instructions in the pipeline and the dependency detection process are executed in parallel, realizing the post-dependence detection of the instructions.

Benefits of technology

Reduces the number of clock cycles that need to wait for completion of dependency detection before sending the instruction, and improves the execution efficiency of the instruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010924A_ABST
    Figure CN120010924A_ABST
Patent Text Reader

Abstract

The invention discloses an instruction execution device, a chip, electronic equipment and a method, and belongs to the technical field of chips. The instruction execution device comprises; the instruction dispatching unit is used for sending a first instruction to the first assembly line and sending a second instruction to the second assembly line; the first instruction is one of a storage instruction and a loading instruction, and the second instruction is the other one of the storage instruction and the loading instruction; the dependency detection unit is used for determining at least one target second instruction from the at least one second instruction after the first assembly line receives the first instruction; the target second instruction has the same instruction address as the first instruction and violates an instruction execution time sequence, and the instruction execution time sequence refers to an execution sequence between the first instruction and the target second instruction; and the time sequence processing unit is used for controlling the execution process of the first instruction and the target second instruction to meet the instruction execution time sequence. The instruction execution device is helpful for improving the instruction execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of chip technology, and in particular to an instruction execution device, a chip, an electronic device and a method. Background Art

[0002] In order to improve the execution efficiency of instruction pairs, a method of out-of-order instruction dispatch is used in high-performance processors to dispatch multiple instructions in one clock cycle. Since the out-of-order process may cause errors in the execution timing of instructions with dependencies, it is necessary to perform dependency detection on instructions.

[0003] In the related art, instructions are dispatched to corresponding instruction issuing units according to the instruction types. For store instructions, the instruction dispatch unit directly dispatches the store instruction to the corresponding instruction issuing unit, and the store instruction is issued to the corresponding pipeline through the instruction issuing unit. As for load instructions, before the instruction dispatch unit dispatches the load instruction, it is necessary to perform an address dependency check on the load instruction (such as load mem violation check0 / 1). If the load instruction passes the address dependency check, the instruction dispatch unit dispatches the load instruction to its corresponding instruction issuing unit; the load instruction is issued to the data loading pipeline through the instruction issuing unit.

[0004] However, in the related art, address dependency detection is performed before a load instruction is issued, which causes a delay in the issuance time of the load instruction and affects the efficiency of instruction execution. Summary of the invention

[0005] The present application provides an instruction execution device, a chip, an electronic device and a method. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, an instruction execution device is provided, the instruction execution device comprising: an instruction dispatch unit, a first pipeline, a second pipeline, a dependency detection unit and a timing processing unit;

[0007] The instruction dispatch unit is used to send a first instruction to the first pipeline and send a second instruction to the second pipeline; wherein the first instruction is one of a store instruction and a load instruction, and the second instruction is the other of the store instruction and the load instruction;

[0008] The dependency detection unit is configured to determine at least one target second instruction from at least one second instruction after the first pipeline receives the first instruction; wherein the target second instruction has the same instruction address as the first instruction and violates an instruction execution sequence, wherein the instruction execution sequence refers to an execution order between the first instruction and the target second instruction;

[0009] The timing processing unit is used to control the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

[0010] According to one aspect of an embodiment of the present application, a chip is provided, and the chip includes: the instruction execution device as described above.

[0011] According to one aspect of an embodiment of the present application, an electronic device is provided, which includes the instruction execution device as described above.

[0012] According to one aspect of an embodiment of the present application, there is provided an instruction processing method applied to an instruction execution device, the instruction execution device comprising: an instruction dispatch unit, a first pipeline, a second pipeline, a dependency detection unit and a timing processing unit;

[0013] The instruction dispatch unit sends a first instruction to the first pipeline, and sends a second instruction to the second pipeline; wherein the first instruction is one of a store instruction and a load instruction, and the second instruction is the other of the store instruction and the load instruction;

[0014] After the first pipeline receives the first instruction, the dependency detection unit determines at least one target second instruction from at least one of the second instructions; wherein the target second instruction has the same instruction address as the first instruction and violates an instruction execution sequence, wherein the instruction execution sequence refers to an execution order between the first instruction and the target second instruction;

[0015] The timing processing unit controls the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

[0016] The technical solution provided in the present application provides an instruction execution device, which performs dependency detection on instructions during the process of executing instructions in a pipeline, so that the execution process of the instructions in the pipeline and the dependency detection process are executed in parallel, thereby realizing post-dependency detection of the instructions.

[0017] Compared with the related art, the dependency check is performed on the instruction before the instruction dispatch unit distributes the instruction, which reduces the number of clock cycles that need to wait for the dependency check to be completed before issuing the instruction, and helps to improve the execution efficiency of the instruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the instruction flow during the execution of program instructions;

[0019] Figure 2 It is a schematic diagram of dependency detection in related art;

[0020] Figure 3 It is a schematic diagram of an inventive concept provided by an exemplary embodiment of the present application;

[0021] Figure 4 is a schematic diagram of an instruction execution device provided by an exemplary embodiment of the present application;

[0022] Figure 5 is a detection schematic diagram of a first detection unit provided by an exemplary embodiment of the present application;

[0023] Figure 6 is a schematic diagram of maintaining a mirror queue provided by an exemplary embodiment of the present application;

[0024] Figure 7 is a schematic diagram of detecting the local oldest selection unit provided by an exemplary embodiment of the present application;

[0025] Figure 8 is a detection schematic diagram of a second detection unit provided by an exemplary embodiment of the present application;

[0026] Fig. 9 is a schematic diagram of a post-dependency detection process provided by an exemplary embodiment of the present application;

[0027] Fig.10 is a schematic diagram of determining a start re-execution instruction provided by an exemplary embodiment of the present application;

[0028] Fig.11 It is a flowchart of an instruction processing method applied to an instruction execution device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0030] Before introducing the present application, the terms involved in the present application are explained.

[0031] The data loading pipeline stage 0 (Load stage 0, Load s0) refers to an execution stage of a load instruction in the data loading pipeline. The data loading pipeline includes multiple execution stages. Exemplarily, if a load instruction completes each execution stage in the data loading pipeline in sequence, the load instruction is executed and the load instruction leaves the data loading pipeline.

[0032] The first stage of the data loading pipeline (load stage 1, Load s1) refers to another execution stage in which the load instruction is executed in the data loading pipeline. Exemplarily, Load s0 and Load s1 are two consecutive execution stages in the data loading pipeline; after entering the data loading pipeline, the load instruction first enters Load s0, and after Load s0 is executed, the load instruction enters Load s1.

[0033] In-pipeline check (In-pipe check) is used to perform checks during the execution of the pipeline.

[0034] The dependency violation between store and load instructions (store to load violation, Stand violation) refers to the unreasonable actual execution sequence between the store instruction and the load instruction. For the store instruction and the load address with address dependency, it is necessary to ensure that the store instruction is executed first and the load instruction is executed after to avoid the load instruction reading wrong data or being unable to read data. However, during the actual execution of the instructions, the load instruction is executed first and the store instruction is executed after. That is, a violation occurs between the load instruction and the store instruction.

[0035] The compare selector (compare select, Cmp sel) is used to compare two input data and select one data from the two input data as output according to the comparison result.

[0036] Figure 1 It is a diagram of the instruction flow during the execution of program instructions.

[0037] When executing program instructions, the processor must first ensure semantic correctness. Since program instructions are executed sequentially, there may be dependencies between the instructions executed later and the instructions executed earlier. That is, the instructions executed later need to observe the execution results of the instructions executed earlier. Figure 1 As shown, instruction 1 (INSTR1) is a storage instruction, and the address it writes is A. In the subsequent instruction stream, there is an instruction n (INSTRn) which is a load instruction, used to load data from address A. In order for instruction n to load the correct data from address A, the processor needs to execute instruction 1 first and then instruction n.

[0038] In the process of executing program instructions, a high-performance processor can issue instructions out of order in order to maximize the performance of executing program instructions. Although issuing instructions out of order enables a high-performance processor to execute instructions in parallel, for two instructions that are dependent on each other, it may cause timing violations in the actual execution of the two instructions. For example, if the issuance order of instruction n is earlier than instruction 1, the load instruction will read the old value in address A, that is, instruction n reads the wrong data. This affects the correctness of the program instruction execution process. In order to avoid this problem, a dependency detection mechanism is provided in the high-performance processor, and the dependency detection is used to determine whether a dependency detection violation occurs during the execution of instructions.

[0039] Figure 2 It is a schematic diagram of dependency detection in related technologies.

[0040] As above Figure 2 As shown, the instruction dispatch unit (dispatch) dispatches the store instruction and the load instruction to different issuing units according to the instruction type. For the store instruction, after the instruction data and instruction address of the store instruction are determined, the instruction dispatch unit dispatches the store instruction to the instruction issuing unit (storeissue) for issuing the store instruction. The instruction issuing unit of the store instruction sends the store instruction to the data storage pipeline. After the multi-stage execution of the data storage pipeline, the storage operation corresponding to the store instruction is completed.

[0041] For a load instruction, after the instruction dispatch unit obtains the instruction address of the load instruction, it first performs an address dependency check on the load instruction. Figure 2 The dependency detection logic shown requires 2 clock cycles (load memviolation check0 / 1) to complete.

[0042] When the load instruction is determined to have no dependencies, the instruction dispatch unit dispatches the load instruction to the instruction issue unit for issuing the load instruction. The instruction issue unit for the load instruction sends the load instruction to the data load pipeline. Similarly, the load instruction needs to be executed by multiple stages of the data load pipeline to complete the storage operation corresponding to the load instruction.

[0043] It can be seen that compared with the issuance process of storage instructions, the instruction issuance unit used to issue load instructions is relatively delayed in issuing load instructions. The load instruction needs to wait for several clock cycles (including the clock cycle used for timing detection and the clock cycle used to calculate the instruction address of the preceding storage instruction of the load instruction) during issuance, resulting in low execution efficiency of the load instruction.

[0044] In the dependency detection method provided by the related art, the dependency detection of the load instruction is performed before the load instruction is issued. On the one hand, this requires waiting for the instruction addresses of all storage instructions before the load instruction to be calculated. In an out-of-order processor, the instruction address calculation of the storage instruction may be related to the calculation result of the preceding instruction. If a serious blockage occurs during the execution of the preceding instruction of the storage instruction, resulting in the instruction address of the storage instruction having to wait for multiple clock cycles to be obtained, the dependency detection of the load instruction will be greatly delayed, greatly affecting the normal issuance of the load instruction.

[0045] On the other hand, after all the instruction addresses of the preceding storage instructions of the load instruction are calculated, it takes several clock cycles to complete the dependency detection logic according to the storage instruction queue, which postpones the issuance time of the load instruction again. It can be seen that the preceding dependency detection of the load instruction in the related art will lead to low execution efficiency of the load instruction.

[0046] Figure 3 It is a schematic diagram of the inventive concept provided by an exemplary embodiment of the present application.

[0047] Compared with the related art that performs address dependency detection logic before issuing a load instruction, an instruction execution device is proposed in an embodiment of the present application, which can perform address dependency detection after the instruction is issued. Specifically, during the process of executing the instruction by the pipeline, the instruction execution device performs address dependency detection on the instruction.

[0048] Figure 3 This is an example case, in which the inventive concept of the present application is introduced and explained by performing dependency detection violation on a storage instruction. Figure 3 As shown, after receiving the load instruction, the instruction dispatch unit (dispatch) distributes the load instruction to the instruction issue unit (load issue) of the load instruction. Then, the instruction issue unit corresponding to the load instruction issues the load instruction. After receiving the load instruction, the data load pipeline executes the load instruction.

[0049] After receiving the storage instruction, the instruction dispatch unit distributes the storage instruction to the instruction issuing unit (storeissue) of the storage instruction. Subsequently, the instruction issuing unit corresponding to the storage instruction issues the storage instruction. After receiving the storage instruction, the data storage pipeline starts to execute the storage instruction in multiple stages. At the same time, the dependency detection unit detects whether there is a load instruction that depends on the data that needs to be stored by the storage instruction, and the execution progress of the load instruction takes precedence over the execution progress of the storage instruction. If such a load instruction exists, the load instruction needs to be executed again.

[0050] The post-dependency detection logic is used to perform dependency detection on instructions, which reduces the waiting time before the instruction issuance stage and helps improve the execution efficiency of instructions.

[0051] The instruction execution device provided in the embodiment of the present application can be used as a complete high-performance processor, or a part of the high-performance processor can be encapsulated in a chip. The chip can be an AI (Artificial Intelligence) chip for model training, a chip for image processing, a chip for video processing, etc. The present application does not limit the type of chip used by the instruction execution device.

[0052] The chip corresponding to the instruction execution device provided in the embodiment of the present application can also be applied to the intelligent transportation system. Intelligent Traffic System (ITS), also known as Intelligent Transportation System (ITS), is an effective and comprehensive application of advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) to transportation, service control and vehicle manufacturing, strengthening the connection between vehicles, roads and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment and saves energy. For example, embodiments of the present invention can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, assisted driving, etc., to improve the execution efficiency of instructions by the instruction execution device in the intelligent transportation system.

[0053] Figure 4 4 is a schematic diagram of an instruction execution device provided by an exemplary embodiment of the present application. The instruction execution device 400 includes: an instruction dispatch unit 410 , a first pipeline 420 , a second pipeline 430 , a dependency detection unit 440 and a timing processing unit 450 .

[0054] In some embodiments, the instruction execution device 400 is implemented as a high-performance processor, or a partial structure in a high-performance processor. The high-performance processor refers to a processor that sends multiple instructions out of order in one clock cycle. The high-performance processor may be a superscalar processor.

[0055] Some storage instructions and load instructions need to follow the Write After Read (WAR) principle during execution. That is, there is a dependency relationship between the storage instructions and load instructions corresponding to the same instruction address. The storage instruction needs to be executed first, and then the load instruction, so that after the storage instruction stores the data in a certain instruction address, the load instruction can read the correct data from the instruction address. Optionally, the instruction address can be an address in the cache or an address in the memory.

[0056] For the storage instruction a and the load instruction b that have a dependency relationship, due to the existence of the mechanism of out-of-order instruction issuance, the load instruction b may be executed first and the storage instruction a may be executed later, that is, a dependency violation occurs between the storage instruction a and the load instruction b. During the instruction execution process, it is necessary to perform dependency detection on the instructions in order to find the instructions with timing violations in time and correct them to avoid errors in instruction execution.

[0057] The instruction execution device 400 provided in the present application supports performing dependency detection after sending the instruction to the pipeline, that is, performing post-dependency detection on the instruction, so that the instruction execution process and the dependency detection process are performed in parallel, which helps to avoid delaying the execution timing of the instruction due to dependency detection.

[0058] Next, each component of the instruction execution device 400 is introduced and described.

[0059] In some embodiments, the instruction dispatch unit 410 is used to send a first instruction to the first pipeline 420 and a second instruction to the second pipeline 430; wherein the first instruction is one of a storage instruction and a load instruction, and the second instruction is the other of the storage instruction and the load instruction.

[0060] Optionally, the instruction dispatch unit 410 is used to determine the instruction issuing unit corresponding to the instruction according to the instruction type, and the instruction issuing unit sends the instruction to the pipeline. Exemplarily, the instruction type of the instruction includes but is not limited to at least one of the following: a storage instruction and a load instruction. The storage instruction is used to store the data to be stored into the instruction address of the storage instruction; the load instruction is used to read data from the instruction address of the load instruction.

[0061] Different from the related art in which a load instruction is subjected to dependency detection before the instruction dispatch unit distributes the load instruction, the instruction execution device 400 provided in the present application performs dependency detection on the instruction during the pipeline execution of the instruction. In some embodiments, the first instruction is a storage instruction and the second instruction is a load instruction. In other embodiments, the first instruction is a load instruction and the second instruction is a storage instruction. In other words, the instruction processing device 400 can perform dependency detection on the storage instruction during the execution of the storage instruction, and can also perform dependency detection on the load instruction during the execution of the load instruction.

[0062] In some embodiments, the first pipeline 420 is used to execute a first instruction. Optionally, the first pipeline 420 includes at least one execution stage, and at least one execution stage is connected in series. In the same clock cycle, different first instructions are executed in each execution stage in the first pipeline 420. That is, in one clock cycle, the first pipeline 420 includes multiple first instructions, and the multiple first instructions are respectively in different execution stages of the first pipeline 420. After a first instruction completes all execution stages of the first pipeline 420, the execution of the first instruction is completed, that is, the first instruction leaves the first pipeline 420.

[0063] Optionally, the instruction execution device 400 includes a plurality of first pipelines 420 , and the plurality of first pipelines 420 can execute different first instructions in parallel, and the plurality of first pipelines 420 do not interfere with each other.

[0064] In some embodiments, the second pipeline 430 is used to execute the second instruction. Optionally, the second pipeline 430 includes at least one execution stage, and at least one execution stage is connected in series. Optionally, the instruction execution device 400 includes multiple second pipelines 430, and the multiple second pipelines 430 can execute different second instructions in parallel.

[0065] It should be noted that the number of execution stages included in the first pipeline 410 and the second pipeline 420, as well as the operations corresponding to each execution stage, need to be determined according to the actual type of the high-performance processor, and this application does not limit this.

[0066] In one example, after the instruction dispatch unit 410 receives the first instruction, the instruction dispatch unit 410 distributes the first instruction to the instruction issuing unit corresponding to the first instruction according to the instruction type of the first instruction. The instruction issuing unit corresponding to the first instruction issues the first instruction; and the first pipeline executes the first instruction after receiving the first instruction. After the instruction dispatch unit 410 receives the second instruction, the instruction dispatch unit 410 distributes the second instruction to the instruction issuing unit corresponding to the second instruction according to the instruction type of the second instruction, and the instruction issuing unit corresponding to the second instruction issues the second instruction; and the second pipeline executes the second instruction after receiving the second instruction.

[0067] Exemplarily, the instruction issuing unit may issue an instruction as follows: the instruction issuing unit issues a broadcast in the processor; after receiving the broadcast, the pipeline executes the instruction based on a control signal related to the instruction in the broadcast.

[0068] In some embodiments, the dependency detection unit 440 is used to determine at least one target second instruction from at least one second instruction after the first pipeline 420 receives the first instruction; wherein the target second instruction has the same instruction address as the first instruction and violates the instruction execution sequence, and the instruction execution sequence refers to the execution order between the first instruction and the target second instruction.

[0069] In some embodiments, the dependency detection unit 440 is used to perform dependency detection on the first instruction during the execution of the first instruction, so as to determine a second instruction that has a dependency violation with the first instruction.

[0070] Based on von Neumann's computer architecture, program instructions follow the principle of sequential execution to ensure that the execution results of program instructions are correct. When executing program instructions, the program instructions need to be converted into multiple instructions that the machine can recognize, mainly including: storage instructions, loading instructions, and operation instructions. There is an execution order between multiple instructions, and this execution order is the instruction execution sequence. Simply put, the instruction execution sequence is used to characterize the execution order of each instruction when it is executed in sequence.

[0071] Optionally, there may be at least two instructions with correlation among the multiple instructions. Correlation is also called dependency. If the execution process of a certain instruction requires the use of the execution result of another instruction, then the two instructions are correlated. Taking instruction a and instruction b as an example, if the execution process of instruction a requires the use of the execution result of instruction b, then instruction a and instruction b are correlated. In other words, there is a dependency between instruction a and instruction b.

[0072] In this example, the instruction execution sequence indicates that instruction b is executed before instruction a, so that instruction a does not need to wait when being executed and can directly obtain the execution result of instruction b.

[0073] Optionally, the instruction execution sequence is represented by the instruction execution number carried by each instruction, and instruction a and instruction b each have a different instruction execution number, and the instruction execution number of instruction b is smaller than the instruction execution number of instruction a. Optionally, the instruction execution sequence is represented by an execution timing table, and the dependency detection unit 440 stores the execution timing table. The dependency detection unit 440 searches the execution timing table for the position of the instruction identifier of a certain instruction to determine the instruction execution number of the instruction. It should be noted that the instruction execution timing is determined by other functional units in the instruction execution device, and the dependency detection unit 440 can directly obtain the instruction execution timing. The present application does not limit the representation form of the instruction execution timing.

[0074] For high-performance processors, in order to improve the efficiency of instruction execution, multiple instructions need to be issued in the same clock cycle (for example, x storage instructions and y load instructions occur in the same clock cycle, where x and y are positive integers greater than 1). This instruction generation method is called out-of-order issuance, which causes the execution order of two related instructions to not meet the instruction execution timing during the actual execution process, that is, a dependency violation occurs.

[0075] In the embodiments of the present application, dependency violation is also called read-before-write violation. Dependency violation means that the actual execution order of storage instructions and load instructions with correlation does not conform to the instruction execution timing. Simply put, for a certain clock cycle of program execution, for storage instructions and load instructions with the same instruction address, if the load instruction takes precedence over the storage instruction, a dependency violation occurs between the storage instruction and the load instruction.

[0076] The dependency detection unit 440 in the embodiment of the present application is used to detect storage instructions and load instructions with dependency violations during instruction execution, so as to adjust the actual execution order of the storage instructions and load instructions with dependency violations through the subsequent timing processing unit.

[0077] For a certain storage instruction and a certain load instruction, if the storage instruction and the load instruction have the same instruction address, and in the instruction stream, the execution sequence of the storage instruction takes precedence over the execution sequence of the load instruction, it means that there is a dependency relationship between the storage instruction and the load instruction, and the storage instruction needs to be executed before the load instruction. If during the instruction execution process, the load instruction takes precedence over the storage instruction, it means that a dependency violation occurs between the storage instruction and the load instruction.

[0078] Optionally, there is a dependency violation between the first instruction and the target second instruction. The dependency detection unit 400 is used to find a second instruction that has a dependency violation with the first instruction from at least one second instruction during the execution of the first instruction in the first pipeline.

[0079] Optionally, whether there is a dependency relationship between the first instruction and the second instruction is determined based on the instruction address of the first instruction and the instruction address of the second instruction. If the instruction addresses of the first instruction and the second instruction are the same, it means that there is a dependency relationship between the first instruction and the second instruction. In this case, it is necessary to ensure that the first instruction and the second instruction that are storage instructions are executed first. The dependency relationship can also be called read-before-write.

[0080] In some embodiments, the instruction execution sequence is used to characterize the execution order of multiple instructions. Optionally, when the first instruction is a storage instruction and the second instruction is a load instruction, the instruction execution sequence indicates that the first instruction is executed before the target second instruction. Exemplarily, the instruction execution sequence refers to the execution order of each instruction in the instruction stream. Assuming that the first instruction is a storage instruction, the instruction execution sequence indicates that the first instruction is executed before the second instruction; assuming that the first instruction is a load instruction, the instruction execution sequence indicates that the second instruction is executed before the first instruction.

[0081] Optionally, if there is no dependency relationship between the first instruction and the second instruction, the first instruction 1 may be executed before the second instruction 2 , and the second instruction 2 may also be executed before the first instruction 1 .

[0082] In some embodiments, the target second instruction refers to an instruction in at least one second instruction that has a dependency violation with the first instruction.

[0083] Optionally, when the first instruction is a storage instruction and the second instruction is a load instruction, at least one second instruction is a load instruction being executed in the instruction processing device 400, or a load instruction that has been executed. Optionally, when the first instruction is a load instruction and the second instruction is a storage instruction, at least one second instruction is a storage instruction that has not been executed by the instruction processing device 400.

[0084] In one example, after the first pipeline 420 receives the first instruction, the first pipeline 420 issues a dependency detection request to the dependency detection unit 440 when the first instruction is in the first execution stage in the first pipeline; the dependency detection unit 440 starts executing the step of determining at least one target second instruction from at least one second instruction according to the dependency detection request.

[0085] Optionally, the dependency detection request includes an instruction address of the first instruction and timing information of the first instruction. The instruction address of the first instruction is used to represent the storage address of the processing data of the first instruction, and the timing information of the first instruction is used to represent the age of the first instruction, that is, the execution order of the first instruction in the instruction stream.

[0086] In some embodiments, the instruction execution device includes multiple first pipelines. In order to ensure that each first instruction is timely subjected to dependency detection, each first pipeline 420 has its own dependency detection unit 440. That is, the number of dependency detection units 440 included in the instruction execution device 400 is greater than or equal to the number of first pipelines 420 included in the instruction execution device 400. In this case, the composition structures of the multiple dependency detection units 440 are similar and the principles are the same. For the specific content of the dependency detection unit 440, please refer to the following introduction.

[0087] In some embodiments, the timing processing unit 450 is used to control the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

[0088] Optionally, the timing processing unit 450 is used for the second pipeline to re-execute the target second instruction in the second pipeline. For example, when the first instruction is a storage instruction and the second instruction is a load instruction, the timing processing unit 450 is used to re-execute the target load instruction (i.e., the target second instruction mentioned above) in the second pipeline to ensure that after the storage instruction stores data in the instruction address, the target load instruction is executed to read the data from the instruction address. In this case, there is a routing connection between the first pipeline 420 and the dependency detection unit 440; there is a routing connection between the dependency detection unit 440 and the second pipeline 430; there is a routing connection between the dependency detection unit 440 and the timing processing unit 450; and there is a routing connection between the timing processing unit 450 and the second pipeline 430.

[0089] Optionally, the timing processing unit 450 acts on the first pipeline to delay the execution of the first instruction. For example, when the first instruction is a load instruction and the second instruction is a storage instruction, the timing processing unit 450 is used to control the first pipeline to start executing the first instruction after k clock cycles, so that the target storage instruction (that is, the target second instruction mentioned above) can be executed in k clock cycles, and the data is written to the instruction address of the storage instruction to ensure that the target storage instruction is executed before the load instruction. In this case, there is a routing connection between the first pipeline 420 and the dependency detection unit 440; there is a routing connection between the dependency detection unit 440 and the timing processing unit 450; there is a routing connection between the dependency detection unit 440 and the second pipeline 430; there is a routing connection between the timing processing unit 450 and the first pipeline 420.

[0090] In some embodiments, the instruction dispatch unit 410 (a functional unit for implementing the emission of instructions), the first pipeline 420, and the second pipeline 430 are functional units provided in the existing instruction processing device. By adding the dependency detection unit 440 and the timing processing unit 450 to the existing instruction processing device, the instruction processing device in the present application can be obtained. Therefore, the instruction execution device 400 provided in the present application has little impact on the component layout of the existing instruction execution device, and is helpful for modification based on the layout of the existing instruction execution device. For the purpose of each functional unit in the instruction execution device 400, and the connection relationship between them, please refer to the following embodiment.

[0091] In summary, the instruction execution device performs dependency detection on the instruction during the pipeline execution process, so that the execution process of the instruction in the pipeline is executed in parallel with the dependency detection process, and the post-dependency detection of the instruction is realized. Compared with the related art, the dependency detection of the instruction is performed before the instruction dispatch unit distributes the instruction, which reduces the number of clock cycles that need to wait for the completion of the dependency detection before issuing the instruction, which helps to improve the execution efficiency of the instruction.

[0092] As can be seen from the above description, the instruction execution device provided by the present application can perform dependency detection on storage instructions during execution, and can also perform dependency detection on load instructions during execution. Since the logic of dependency detection of storage instructions and load instructions is opposite, the following introduces the detection logic of dependency detection of storage instructions and the logic of dependency detection of load instructions by the instruction execution device provided by the present application through several embodiments.

[0093] Example 1: The first instruction is a storage instruction, the second instruction is a load instruction, the first pipeline is a data storage pipeline, the second pipeline is a data loading pipeline, and the target second instruction refers to a second instruction with the same instruction address as the first instruction, and before the first instruction reaches the first pipeline, the second instruction starts to be executed in the second pipeline.

[0094] The following describes the logic of the instruction execution device performing dependency detection on the first instruction and the instruction being executed through several embodiments.

[0095] In some embodiments, the first instruction is a storage instruction, and the second instruction is a load instruction; the dependency detection unit includes a first detection unit; the first detection unit is used to determine the target second instruction from at least one second instruction being executed; wherein the second instruction being executed refers to a second instruction in any execution stage (Stage) of the second pipeline.

[0096] Optionally, the first detection unit acts on the second pipeline. In this example, the second pipeline is a data loading pipeline. It can be seen from the above that the second pipeline can include multiple execution stages, and the second instruction being executed is any second instruction being executed in the second pipeline. Exemplarily, the second instruction being executed refers to an execution stage in the second pipeline in which the instruction address of the second instruction can be determined.

[0097] For example, the second pipeline includes execution stage 0 (Stage 0), execution stage 1 (Stage 1) and execution stage 2 (Stage 2); wherein, execution stage 1 and execution stage 2 are stages where the instruction address of the second instruction is known, then the second instruction being executed includes: the instruction in execution stage 1 of the second pipeline and the second instruction in execution stage 2.

[0098] In some embodiments, a wiring connection is provided between the first pipeline and the first detection unit, and after receiving the first instruction, the first pipeline generates a dependency detection request of the first instruction. The dependency detection request of the first instruction is transmitted to the first detection unit through the wiring between the first pipeline and the first detection unit; after receiving the dependency detection request of the first instruction, the first detection unit performs the step of determining a target second instruction from at least one second instruction being executed.

[0099] After the data storage pipeline receives the storage instruction, the dependency detection unit performs dependency detection on the storage instruction, so that the dependency detection on the storage instruction and the execution process of the storage instruction in the data storage pipeline can be executed in parallel. Without affecting the execution efficiency of the storage instruction, the target load instruction is screened out through dependency detection, which helps to improve the correctness of the program instruction execution result in the scenario of out-of-order instruction issuance.

[0100] In some embodiments, the first detection unit includes at least one comparator; the comparator is used to determine the candidate second instruction as the target second instruction when the instruction address of the candidate second instruction is consistent with the instruction address of the first instruction; wherein the candidate second instruction refers to the second instruction in the execution stage of the second pipeline corresponding to the comparator.

[0101] Optionally, the first pipeline is connected to the at least one comparator by a wire. Exemplarily, the first pipeline sends a dependency detection request of the first instruction to the at least one comparator through the wire, and the dependency detection request includes an instruction address of the first instruction.

[0102] Optionally, the instruction address of the first instruction refers to the storage address of the operation data (referred to as data) corresponding to the first instruction in the storage space. Exemplarily, the instruction address of the first instruction includes but is not limited to: an address in a cache space (cache) and a main memory address.

[0103] In some embodiments, the comparator requires two input signals and can generate an output signal. Optionally, when performing dependency detection, the two input signals are the instruction address of the first instruction and the instruction address of the candidate second instruction, and the output signal of the comparator is used to indicate whether the instruction address of the candidate second instruction is the target second instruction.

[0104] Optionally, the comparator obtains the instruction address of the candidate second instruction from the second pipeline, obtains the instruction address of the first instruction from the first pipeline, and compares whether the two instruction addresses are the same. Exemplarily, the comparator determines the instruction address of the first instruction by relying on the detection request, and determines the execution address of the candidate second instruction from the execution stage corresponding to the comparator in the second pipeline. For example, the execution stage corresponding to the comparator in the second pipeline is Stage 2, then the comparator reads the instruction address of the candidate second instruction from Stage 2 of the second pipeline.

[0105] Exemplarily, if the instruction address of the first instruction and the instruction address of the candidate second instruction are the same, the comparator determines that the candidate second instruction is the target second instruction; if the instruction address of the first instruction and the instruction address of the candidate second instruction are different, the comparator determines that the candidate second instruction is not the target second instruction.

[0106] Exemplarily, if the instruction address of the first instruction is the same as the instruction address of the candidate second instruction, the comparator compares the timing information of the candidate second instruction with the timing information of the first instruction; if it is determined according to the timing information that the first instruction should be executed before the second instruction, the comparator determines that the candidate second instruction is the target second instruction. Optionally, the timing information of the instruction is used to characterize the execution order of the instruction in the instruction stream, and the timing information of the instruction can be determined according to the instruction execution timing.

[0107] In some embodiments, the number of comparators included in the first detection unit is related to the number of second pipelines included in the instruction execution device. Optionally, the number of comparators included in the first detection unit is greater than or equal to the number of second pipelines included in the instruction execution device.

[0108] Exemplarily, the number of comparators included in the first detection unit is equal to the number of second pipelines included in the instruction execution device. For example, if the instruction execution device includes three second pipelines, then each of the three comparators included in the first detection unit corresponds to one pipeline. For a certain comparator, each execution stage included in the pipeline corresponding to the comparator is the execution stage corresponding to the comparator in the pipeline.

[0109] Exemplarily, the number of comparators included in the first detection unit is greater than the number of second pipelines included in the instruction execution device. Assume that each pipeline includes n execution stages, and the second instructions in m of the n execution stages need to be candidate second instructions, that is, in the same clock cycle, the second instructions respectively executed in the m execution stages belong to the second instruction being executed; then each of the m execution stages corresponds to a comparator; n is a positive integer, and m is a positive integer less than or equal to n. If the instruction execution device includes p second pipelines, the first detection unit includes m*p comparators.

[0110] Since the second instructions being executed in the second pipeline are in different execution stages in the same clock cycle, the execution stage of the second instructions being executed may change after the next clock cycle starts. In order to avoid missing any second instruction being executed during the instruction address detection process, at least one comparator in the first detection unit completes the step of comparing the instruction address of the candidate second instruction with the instruction address of the first instruction in the same clock cycle.

[0111] On the one hand, since the instruction address of the instruction being executed is determined, the comparator can directly obtain the instruction addresses of the storage instructions and the load instructions during the dependency detection process, reducing the time spent waiting for the instruction address to be generated during the dependency detection process, which helps to improve the execution efficiency of the instructions and the efficiency of dependency detection on the instructions.

[0112] On the other hand, by setting a plurality of comparators, the steps of detecting a plurality of load instructions being executed can be completed synchronously, which helps to further improve the efficiency of dependency detection.

[0113] In some embodiments, the first detection unit is used to determine the target second instruction from the second instructions being executed in at least one second pipeline when the first instruction is in the first execution stage; wherein the first execution stage refers to: an execution stage in which the instruction address of the first instruction has been determined in at least one execution stage included in the first pipeline.

[0114] Optionally, the first execution stage refers to the earliest execution stage that the first instruction passes through in at least one execution stage in the first pipeline in which the instruction address of the first instruction is known. For example, the first pipeline includes Stage 0, Stage 1, and Stage 2; when the first instruction is in Stage 1 or Stage 2, the instruction address of the first instruction is determined, and the first instruction first passes through Stage 1 and then Stage 2, then the first execution stage is Stage 1.

[0115] In one example, after receiving the first instruction, the first pipeline generates a dependency detection request for the first instruction when the first instruction enters the first execution stage. The first pipeline sends the dependency detection request to the dependency detection unit, and the dependency detection unit performs a dependency detection on at least one second instruction being executed in the second pipeline when the first instruction is in the first execution stage to determine the target second instruction.

[0116] In some embodiments, the timing processing unit includes an instruction reissue unit; the instruction reissue unit is used to generate at least one clock bubble in the second pipeline, and the clock bubble is used to roll back the target second instruction to the first execution stage of the second pipeline.

[0117] In some embodiments, after the target second instruction is determined from the second instructions being executed, the target second instruction is restarted to be executed in the second pipeline through a reissue mechanism.

[0118] Optionally, the reissue mechanism for the target second instruction is implemented by an instruction reissue unit. The instruction reissue unit acts on the second pipeline. There is a wiring connection between the first detection unit and the instruction reissue unit. After the first detection unit determines the target second instruction; the first detection unit sends the target second instruction to the instruction reissue unit, and the instruction reissue unit controls the second pipeline to re-execute the target second unit.

[0119] Exemplarily, when the comparator in the first detection unit determines that the instruction address of the candidate second instruction is consistent with the instruction address of the first instruction, the comparator notifies the instruction reissuing unit to reissue the target second instruction (in-pipeline).

[0120] Optionally, the instruction reissuing unit is used to generate at least one clock bubble according to the execution stage of the target second instruction in the second pipeline; the instruction reissuing unit is used to insert the above-mentioned at least one clock bubble in the second pipeline, so that the target second instruction is re-executed from the first execution stage in the second pipeline. The specific structure of the instruction reissuing unit to implement instruction reissuance can refer to the relevant technology, and this application will not go into details here. Exemplarily, the first execution stage in the second pipeline refers to Stage 0 in the second pipeline, that is, the first execution stage of the target second instruction in the second pipeline.

[0121] Since the first instruction has completed at least two execution stages in the first pipeline when the target second instruction is re-executed from the first execution stage, this setting allows the first instruction to be executed before the second target instruction, thereby resolving the dependency violation between the first instruction and the second target instruction.

[0122] Figure 5It is a detection schematic diagram of a first detection unit provided by an exemplary embodiment of the present application.

[0123] like Figure 5 The instruction execution device shown includes two second pipelines, namely second pipeline 0 (loadpipeline 0) and second pipeline 1 (load pipeline 1). The second instructions being executed include: the second instruction in the second pipeline is in the first execution stage (load s1), and the second instruction in the second execution stage (load s2).

[0124] When the first instruction is in the first execution stage (store1 s1), the first pipeline generates and sends a dependency check request (store_query) for the first instruction. The dependency check request includes the instruction address of the first instruction. The first detection unit includes at least one selector, and the selector in the first detection unit obtains the instruction address of the first instruction from the dependency check request and obtains the instruction address of the second instruction being executed from the first execution stage and the second execution stage of the second pipeline. Figure 5 As shown, the first detection unit includes 4 selectors, which are respectively used to detect whether the second instruction in the first execution stage of the second pipeline 0, the second instruction in the second execution stage of the second pipeline 0, the second instruction in the first execution stage of the second pipeline 1, and the second instruction in the second execution stage of the second pipeline 1 are the target second instructions. If the instruction address of a second instruction being executed is the same as that of the first instruction, the second instruction being executed is determined to be the target second instruction. In this case, the instruction reissuing unit in the timing processing unit reissues the target second instruction, so that the target second instruction is executed from the first execution stage (such as loads0) in its corresponding second pipeline.

[0125] Through this timing processing logic, the execution progress of the first instruction in the first pipeline is ahead of the execution progress of the target second instruction in the second pipeline, thereby resolving the dependency violation between the first instruction and the target second instruction, and helping to ensure the correctness of the program instruction execution results.

[0126] The following describes and illustrates the execution logic of performing dependency detection on a first instruction and a second instruction that has been executed in an instruction processing device through several embodiments.

[0127] In some embodiments, the first instruction is a storage instruction, and the second instruction is a load instruction; the dependency detection unit includes a second detection unit; the second detection unit is used to determine the target second instruction from at least one executed second instruction; wherein the executed second instruction refers to the second instruction that has left the second pipeline.

[0128] The second instruction that has been executed is a second instruction that has completed all execution stages in the second pipeline, that is, the second instruction that has been executed is a second instruction that has left the second pipeline but has not been committed.

[0129] Optionally, the second instruction queue records at least one second instruction that has been executed, and the second instruction queue includes the instruction address and timing information of the second instruction that has been executed. Exemplarily, the second instruction queue obtains the second instruction through a broadcast transmitted by the instruction transmitting unit of the second instruction.

[0130] In one example, the second detection unit can determine the second instruction that has been executed from the second instruction queue. Since the second instruction queue is closely related to the second pipeline in the instruction execution device, the second instruction queue is arranged around the second pipeline, and the routing distance between the second instruction queue and the first pipeline is relatively far.

[0131] In another example, in order to facilitate routing arrangement, a mirror queue is set around the first pipeline, and the mirror queue is used to record at least one second instruction that has been executed; the second detection unit obtains the second instruction that has been executed from the mirror queue. For the specific content of this embodiment, please refer to the embodiment below.

[0132] In some embodiments, the second detection unit includes a selector; the selector is used to determine the target second instruction from at least one executed second instruction based on the instruction address of the first instruction and the timing information of the first instruction.

[0133] In some embodiments, for the selector in the second detection unit, the selector is used to compare the instruction address of the completed second instruction with the instruction address of the first instruction, and to compare the timing information of the first instruction with the timing information of the second instruction.

[0134] Exemplarily, when the instruction address of a second instruction that has been executed is the same as the instruction address of a first instruction, and, based on the timing information of the two instructions, it is determined that the first instruction takes precedence over the second instruction, the selector in the second detection unit determines the second instruction that has been executed as the target second instruction.

[0135] Exemplarily, when the instruction address of a second instruction that has been executed is different from the instruction address of the first instruction, the selector in the second detection unit determines that the second instruction that has been executed is not the target second instruction.

[0136] Exemplarily, when it is determined according to the timing information that a second instruction that has been executed takes precedence over the first instruction, the selector in the second detection unit determines that the second instruction that has been executed is not the target second instruction.

[0137] In some embodiments, the second detection unit is wired to the first pipeline. Optionally, the first pipeline generates a dependency detection request for the first instruction when the first instruction enters the first execution stage, and the first pipeline sends the dependency detection request to the selector in the second detection unit; the dependency detection request includes the instruction address of the first instruction and the timing information of the first instruction.

[0138] Optionally, the timing information of the instruction is determined based on the instruction execution timing. Exemplarily, the timing information of the instruction includes the instruction execution sequence number carried by the instruction in the above embodiment. In one example, the selector in the second detection unit obtains the instruction address of the first instruction and the timing information of the first instruction from the dependency detection request, and the selector in the second detection unit obtains the instruction address and timing information of the executed second instruction from the second instruction queue.

[0139] In this case, a selector in the second detection unit is wired to the second instruction queue. Optionally, the number of selectors in the second detection unit is related to the number of entries in the second instruction queue, each of which is used to store a second instruction. The cells in the second instruction queue are used to store an instruction address of a second instruction that has been executed and timing information of the second instruction.

[0140] Exemplarily, the number of selectors in the second detection unit is equal to the number of units included in the second instruction queue.

[0141] In another example, the selector in the second detection unit obtains the instruction address and timing information of the first instruction from the dependency detection request, and the selector in the second detection unit obtains the instruction address and timing information of the executed second instruction from the mirror queue.

[0142] In this case, a selector in the second detection unit is wired to the mirror queue, and the number of selectors in the second detection unit is related to the capacity of the mirror queue. Each unit is used to store a second instruction. Exemplarily, the number of selectors in the second detection unit is equal to the number of units included in the mirror queue.

[0143] In some embodiments, the instruction execution device also includes: a mirror queue, wherein the routing distance between the mirror queue and the first pipeline is smaller than the routing distance between the mirror queue and the second pipeline; the mirror queue is used to store the executed second instructions; and the second detection unit obtains at least one executed second instruction through the mirror queue so as to determine the target second instruction from the at least one executed second instruction.

[0144] In some embodiments, a mirror queue refers to a storage element capable of storing a second instruction that has been executed. Optionally, the mirror queue includes at least one unit, and different units are used to store different second instructions. The number of units included in the mirror queue is the capacity of the mirror queue. It should be noted that the mirror queue can be a storage element capable of realizing a queue structure, or a storage element for realizing other data structures. This application does not limit the type of element to which the mirror queue belongs.

[0145] Exemplarily, for a unit in the mirror queue, the data stored in the unit includes at least one of the following: an instruction identifier of a second instruction that has been executed, an instruction address of the second instruction that has been executed, timing information of the second instruction that has been executed, and a second pipeline identifier.

[0146] Among them, the instruction identifier of the executed second instruction is used to uniquely identify the second instruction, and the instruction identifier of the executed second instruction includes but is not limited to at least one of the following: the instruction serial number of the executed second instruction, and the instruction name of the executed second instruction.

[0147] The instruction address of the executed second instruction refers to the action address of the second instruction. In Example 1, the second instruction is a load instruction, and the instruction address of the executed second instruction refers to the storage address of the data that the load instruction needs to read in the storage space, which can be understood as the action address of the load instruction.

[0148] The timing information of the executed second instruction is used to indicate the execution order of the second instruction. For example, the timing information is used to indicate which first instructions need to be executed before the second instruction.

[0149] The second pipeline identifier is used to indicate a second pipeline that executes the second instruction.

[0150] Optionally, a selector in the second detection unit is wired to a corresponding unit in the mirror queue. For a selector 1 in the second detection unit, the selector 1 corresponds to unit 1 in the mirror queue, and the selector 1 is wired to unit 1, and the selector 1 obtains the instruction address and timing information of the executed second instruction stored therein from the unit 1.

[0151] In some embodiments, in the instruction execution device, the mirror queue is arranged near the first pipeline, and the wiring distance between the mirror queue and the first pipeline is less than the wiring distance between the mirror queue and the second pipeline. Of course, the mirror queue can also be arranged at any position in the instruction execution unit, and this application does not limit it here.

[0152] Since the first pipeline needs to notify the second detection unit to determine the target second instruction from the executed second instruction after receiving the first instruction, arranging the second detection unit near the first pipeline helps to shorten the wiring length between the second detection unit and the first pipeline.

[0153] On this basis, a mirror queue for recording the executed second instruction is set near the first pipeline, which not only ensures that the second detection unit can obtain the executed second instruction, but also helps to shorten the wiring length between the mirror circuit and the second detection unit. By setting up a mirror queue, it helps to simplify the wiring simplicity in the instruction execution device. It also helps to shorten the delay for the second detection unit to obtain the instruction address and timing information of the executed second instruction.

[0154] In order to ensure that the mirror queue records and maintains the executed second instructions in a timely manner, when a second instruction leaves the second pipeline, the second instruction and information related to the second instruction are written into the mirror queue in a timely manner to ensure the comprehensiveness of the executed second instructions stored in the mirror queue and avoid errors in the dependency detection process.

[0155] In some embodiments, there is a routing connection between the mirror queue and the second pipeline, and there is a routing connection between the mirror queue and the dependency detection unit; the second pipeline is used to send the executed second instructions to the mirror queue, and the maximum number of executed second instructions that can be written in parallel to the mirror queue in one clock cycle is greater than or equal to the number of second pipelines included in the instruction execution device.

[0156] Optionally, the mirror queue is connected to the last execution stage of the second pipeline. After the second instruction being executed in the second pipeline completes the last execution stage, the second instruction is converted from being executed to having been executed. Before the second instruction leaves the second pipeline, the second pipeline writes the second instruction into the mirror queue.

[0157] Exemplarily, the second pipeline transmits at least one of the following to the mirror queue through routing: an instruction identifier of the second instruction, an instruction address of the second instruction, timing information of the second instruction, and a pipeline identifier of the second pipeline.

[0158] In the case where the instruction execution device includes multiple second instructions, since each second pipeline can execute different second instructions in parallel, multiple second instructions may be executed in the same clock cycle. In order to ensure that all the executed second instructions can be written into the mirror queue, the maximum parallel write quantity of the mirror queue in one clock cycle is greater than or equal to the number of second pipelines included in the instruction execution device.

[0159] Optionally, in the case of x second pipelines included in the instruction execution device, the mirror queue includes at least x write ports, where x is a positive integer. Each second pipeline has a routing connection with a write port in the mirror queue, and different second pipelines are connected to different write ports in the mirror queue. Different write ports correspond to different units in the mirror queue, and the second pipeline writes the executed second instruction into the corresponding unit through routing. The corresponding unit refers to the unit corresponding to the write port that has a routing connection with the second pipeline.

[0160] Figure 6 It is a schematic diagram of maintaining a mirror queue provided by an exemplary embodiment of the present application.

[0161] Assume that the instruction execution device includes two second pipelines, namely, second pipeline 0 and second pipeline 1, and the second execution stage (load s2) in the two second pipelines is the last execution stage in the pipeline. Then, there is a routing connection between the mirror queue and the second pipeline 0, and there is a routing connection between the mirror queue and the second pipeline 1.

[0162] Taking the second pipeline 0 as an example, if a second instruction is executed in the second pipeline 0, the second pipeline 0 regards the second instruction as the second instruction that has been executed, and transmits the data related to the second instruction to a unit in the mirror queue through the routing. The data related to the second instruction includes at least one of the following: an instruction identifier, an instruction address, timing information, and an identifier of the second pipeline 0.

[0163] The above mechanism enables the mirror queue to record the executed second instruction, so that the first pipeline performs dependency detection on the first instruction and the executed second instruction based on the mirror queue in the first execution stage of the first instruction.

[0164] By setting up a wiring connection between the mirror queue and the pipeline, the executed instructions in the pipeline can be written into the mirror queue in time, ensuring that the mirror queue includes all the executed instructions. This avoids the situation where some instructions with dependency violations cannot be found in the dependency detection process due to the incompleteness of the executed instructions included in the mirror queue, resulting in errors in the execution of program instructions and affecting the reliability of the dependency detection logic.

[0165] Since the number of completed second instructions is relatively large, multiple target second instructions may be detected by the selector in the second detection unit, and the multiple target second instructions need to be re-executed. That is, redirection is performed for each target second instruction. By providing a starting re-execution instruction among the multiple target second instructions, the re-execution operation of the target second instructions is automatically started from the starting execution instruction, which helps to simplify the processing logic for re-executing the target second instructions.

[0166] In some embodiments, the second detection unit also includes: a local oldest selection group, which has a wiring connection with the selector; the selector is also used to pass the target second instruction to the local oldest selection group after determining the target second instruction; the local oldest selection group is used to determine the starting re-execution instruction from the target second instruction, and the starting re-execution instruction refers to the target second instruction with the earliest execution timing.

[0167] In some embodiments, the start re-execution instruction has the earliest execution order among the multiple target second instructions. That is, the start re-execution instruction refers to the target second instruction with the earliest execution order among the multiple target second instructions in the instruction stream. For example, the multiple target second instructions include second instruction 1, second instruction 2, second instruction 3, and second instruction 4, and the execution order of the target second instructions from early to late in the instruction stream is: second instruction 4, second instruction 2, second instruction 1, second instruction 3, then second instruction 4 is the start re-execution instruction.

[0168] Optionally, when the selector in the second detection unit determines one target second instruction from the executed second instructions, the local oldest selection group in the second detection unit does not perform the step of determining the starting re-execution instruction from the target second instruction. Optionally, when the selector in the second detection unit determines multiple target second instructions from the executed second instructions, the local oldest selection group in the second detection unit performs the step of determining the starting re-execution instruction from the target second instruction.

[0169] In some embodiments, the local oldest selection group is used to select a starting re-execution unit from a plurality of target selection instructions according to timing information between the plurality of target selection instructions.

[0170] Optionally, the local oldest selection group uses local comparison logic to select a starting re-execution unit from a plurality of target second instructions.

[0171] Since the starting re-execution instruction is the second target instruction with the earliest execution sequence among the multiple target second instructions, the instruction re-execution process starting from the starting re-execution instruction is more in line with the sequential nature of instruction execution. Since the starting re-execution instruction is the earliest in sequence, the execution sequence of other instructions related to the starting re-execution instruction (such as a calculation instruction that performs calculations based on data read by the starting re-execution instruction) is earlier than the execution sequence of other instructions related to other target second instructions. By re-executing the target second instruction starting from the starting re-execution instruction, it is helpful to reduce the impact of re-executing multiple target second instructions on other instructions.

[0172] In some embodiments, the local oldest selection group includes at least one level of local oldest selection unit (selectpartial oldest unit); for the i-th level local oldest selection unit in the at least one level of local oldest selection unit, if the i-th level local oldest selection unit is the lowest-level local oldest selection unit, then the i-th level local oldest selection unit is used to determine the i-th level local oldest instruction from multiple target selection instructions, where i is a natural number; if the i-th level local oldest selection unit corresponds to the i-1-th level local oldest selection unit, then the i-th level local oldest selection unit is used to determine the i-th level local oldest instruction from the i-1-th level local oldest selection instruction; if the number of the i-th level local oldest instruction is 1, then the i-th level local oldest instruction is the starting re-execution instruction.

[0173] In some embodiments, the local oldest selection group includes at least one local oldest selection unit. Optionally, the number of local oldest selection units included in the local oldest selection group is related to the capacity of the mirror queue and the data processing capability of the local oldest selection unit. The local oldest selection unit is used to select the local oldest instruction from q target second instructions, and the larger q is, the stronger the data processing capability is; q is a positive integer.

[0174] Exemplarily, the stronger the data processing capability of the local oldest selection unit, the fewer the number of local oldest selection units included in the oldest selection group; the weaker the data processing capability of the local oldest selection unit, the more the number of local oldest selection units included in the local oldest selection group.

[0175] Exemplarily, the smaller the capacity of the mirror queue, the smaller the number of local oldest selection units included in the local oldest selection group; the larger the capacity of the mirror queue, the larger the number of local oldest selection units included in the local oldest selection group.

[0176] In some embodiments, the local oldest selection unit is stacked by at least one selection comparison unit. For any selection comparison unit, the selection comparison unit selects one input data from the two input data as the output data of the selection comparison unit based on the comparison result of the two input data. In the embodiment of the present application, the two input data of the selection comparison unit are two target second instructions respectively, and the output result of the selection comparison unit is the target second instruction with an earlier execution order among the two target second instructions.

[0177] Optionally, in the local oldest selection unit, at least one selection comparison unit is stacked into a tree structure. Assuming that the local oldest selection unit includes j layers of selection comparison units, the first layer includes 2 j-1 There are selection and comparison units, and the tth layer includes 2 j-tThe selection and comparison units of the first layer obtain two target second instructions from the outside of the local oldest selection unit.

[0178] Figure 7 It is a detection schematic diagram of the local oldest selection unit provided by an exemplary embodiment of the present application.

[0179] like Figure 7 As shown, the local oldest selection unit includes 3 layers of 7 comparison selection units, wherein the first layer includes 4 selection comparison units, the second layer includes 2 selection comparison units, and the third layer includes 1 selection comparison unit. The 4 selection comparison units of the first layer are used to obtain c target second instructions from the outside of the selection comparison unit, where c is a positive integer less than or equal to 8. The selection comparison unit of the third layer of the local oldest selection unit is used to output the target second instruction with the earliest execution order among the c target second instructions as the local oldest instruction determined by the local oldest selection unit.

[0180] In some embodiments, when the local oldest selection group includes multiple local oldest selection units, that is, when the local oldest selection group includes at least two levels of local oldest selection units, a register is arranged between the i-th level local oldest selection unit and the i-1-th level local oldest selection unit, and the register is used to store the i-1-th level local oldest instruction determined by the i-1-th level local oldest selection unit.

[0181] Figure 8 It is a detection schematic diagram of a second detection unit provided by an exemplary embodiment of the present application.

[0182] like Figure 8 As shown, the second detection unit performs dependency detection on the executed second instruction and the first instruction. In this example, when the first instruction is in the first execution stage, the first pipeline generates and sends a dependency detection request for the first instruction. The dependency detection request includes the instruction address and timing information of the first instruction.

[0183] The second detection unit includes at least one comparator, and the comparator 810 in the at least one comparator obtains the instruction address and timing information of the second instruction a from the unit 820 corresponding to the mirror queue. The comparator 810 compares whether the instruction address of the first instruction is consistent with the instruction address of the second instruction a, and determines whether the execution order of the first instruction is prior to the execution order of the second instruction according to the timing information.

[0184] If the instruction address of the first instruction is consistent with the instruction address of the second instruction a, and the execution order of the first instruction is prior to the execution order of the second instruction, the second instruction a is determined to be the target second instruction. Optionally, the process of the selector determining the target second instruction from the executed second instructions is completed in the first execution stage of the first instruction in the first pipeline.

[0185] If multiple target second instructions are determined, it is necessary to determine the starting re-execution instruction from the multiple target second instructions through the local oldest selection unit in the local oldest selection group. The multiple target second instructions are re-executed starting from the starting re-execution instruction through the instruction resending unit. Assuming that the capacity of the mirror queue is 64, the local oldest selection unit can select one local oldest instruction from 8 target second instructions at a time, then two levels of local oldest selection units are required to determine the starting re-execution second instruction from the multiple target second instructions.

[0186] Fig. 9 It is a schematic diagram of a post-dependency detection process provided by an exemplary embodiment of the present application.

[0187] Depend on Fig. 9 As shown, the post-dependency check for the storage instruction includes two branches, one branch 910 is to perform dependency check on the storage instruction and the load instruction being executed in the data loading pipeline; the other branch 920 is to perform dependency check on the storage instruction and the load instruction that has been executed. For the specific execution content of the above two branches, please refer to the above embodiments, which will not be repeated here.

[0188] Fig.10 It is a schematic diagram of determining a start re-execution instruction provided by an exemplary embodiment of the present application.

[0189] Fig.10 As shown, multiple load instructions depend on the same store instruction. Specifically, the instruction addresses of the three load instructions INSTR6 / 7 / 10 are the same as the instruction address of the store instruction INSTR1, and these three load instructions are executed before the store instruction. If during the execution of the program instruction, INSTR6 / 7 / 10 is issued first, and the execution is completed and enters the mirror queue, and INSTR1 is issued later, when INSTR1 passes the store s1 stage, a dependency check is performed based on the mirror queue, and it is found that three load instructions have dependency violations on the current store instruction. Then, the local oldest selection group logic is used to find the starting re-execution instruction. In this example, the starting re-execution instruction is INSTR6, and the subsequent re-execution starts from the starting re-execution instruction and flushes the entire load pipeline.

[0190] The post-dependency detection logic is implemented by the above instruction execution device, so that the load instruction is issued as early as possible. The post-dependency detection logic redirects the data load pipeline after a dependency violation occurs in the data pipeline, or reissues the load instruction that has a timing sequence with the storage instruction. This helps to improve the performance of the instruction processing device in post-processing instructions.

[0191] Example 2: In some embodiments, the first instruction is a load instruction, and the second instruction is a store instruction; the dependency detection unit includes a third detection unit; the third detection unit is used to determine the target second instruction from at least one unexecuted second instruction; wherein the unexecuted second instruction refers to a second instruction that has not been issued to the second pipeline.

[0192] In one example, the unexecuted second instruction refers to a second instruction that has not reached the first execution stage of the second pipeline. The third detection unit is used to determine the second instructions with the same instruction address as the first instruction from the unexecuted second instructions, and determine these instructions as target second instructions.

[0193] Optionally, after the third detection unit waits for the instruction address of the second instruction whose execution sequence is prior to the first instruction to be calculated, it determines whether the unexecuted second instruction is the target second instruction by comparing whether the instruction address of the first instruction is the same as the instruction address of the unexecuted second instruction.

[0194] If the instruction address of the first instruction is the same as the instruction address of the unexecuted second instruction, the third detection unit determines that the unexecuted second instruction is the target second instruction; if the instruction address of the first instruction is different from the instruction address of the unexecuted second instruction, the third detection unit determines that the unexecuted second instruction is not the target second instruction.

[0195] That is, in the solution provided by the embodiment of the present application, the instruction dispatch unit directly distributes the first instruction to the instruction generation unit, and the instruction emission unit emits the first instruction to the first pipeline. After the first instruction reaches the first pipeline, the third detection unit starts to perform dependency detection on the first instruction.

[0196] Optionally, the third detection unit includes at least one comparator, and the comparator is used to compare whether the instruction address of the first instruction is the same as the instruction address of the second instruction.

[0197] Compared to the related art, in which dependency checking is performed on the load instruction before the instruction dispatch unit generates the load instruction, the instruction execution device in the embodiment of the present application performs dependency checking on the load instruction after sending the load unit to the data loading pipeline. Since it takes a certain number of clock cycles for the load instruction to reach the first pipeline from the instruction dispatch unit, the instruction address of the second instruction that is not executed during these clock cycles may be calculated. Therefore, this setting helps to reduce the clock cycles that need to be waited for in the process of dependency checking on the load instruction, and helps to improve the execution efficiency of the load instruction by the instruction execution device.

[0198] In some other embodiments, the dependency detection unit includes a fourth detection unit; the fourth detection unit is used to determine the execution result of the dependent second instruction from at least one second instruction that is being executed or has been executed.

[0199] The dependent second instruction refers to the second instruction on which the first instruction depends, that is, the storage instruction on which the load instruction depends. Optionally, if the execution result of the dependent second instruction indicates that the dependent second instruction has been completed or has completed the target execution stage in the second pipeline, it means that the execution of the first instruction will not cause a dependency violation, and the first pipeline continues to execute the first instruction. The target execution stage refers to the stage in the second pipeline where the dependent second instruction writes data to its instruction address.

[0200] Optionally, if the execution result of the dependent second instruction indicates that the dependent second instruction is not executed, it means that there will be a timing error in executing the first instruction, and the first pipeline delays the execution of the first instruction. Exemplarily, the execution of the first instruction is delayed for k clock cycles.

[0201] In some embodiments, the timing processing unit includes a delay processing unit; the delay processing unit is used to control the first instruction to be temporarily executed in the first pipeline until the target second instruction is completely executed in the second pipeline.

[0202] The delay processing unit is used to generate at least one clock bubble and insert the clock bubble into the first pipeline, so that the first instruction waits for the target second instruction to store the operation data at the instruction address in the second pipeline, and then the first pipeline executes the first instruction.

[0203] This post-dependency detection logic helps avoid waiting for a long number of clock cycles before executing a load instruction, helps improve the execution efficiency of the load instruction, and improves the performance of the instruction execution device.

[0204] The following is an embodiment of the method of the present application. For details not disclosed in the embodiment of the method of the present application, please refer to the embodiment of the instruction execution device of the present application. Fig.11It is a flowchart of an instruction processing method applied to an instruction execution device provided by an exemplary embodiment of the present application.

[0205] The instruction execution device includes: an instruction dispatch unit, a first pipeline, a second pipeline, a dependency detection unit and a timing processing unit; the instruction processing method applied to the instruction execution device may include steps (1110-1130).

[0206] In step 1110, the instruction dispatch unit sends a first instruction to the first pipeline and sends a second instruction to the second pipeline; wherein the first instruction is one of a store instruction and a load instruction, and the second instruction is the other of the store instruction and the load instruction.

[0207] Step 1120, after the dependency detection unit receives the first instruction in the first pipeline, it determines at least one target second instruction from at least one second instruction; wherein the target second instruction has the same instruction address as the first instruction and violates the instruction execution sequence, and the instruction execution sequence refers to the execution order between the first instruction and the target second instruction.

[0208] Step 1130 , the timing processing unit controls the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

[0209] In some implementations, the first instruction is a storage instruction, and the second instruction is a load instruction; the dependency detection unit includes a first detection unit; after the first pipeline receives the first instruction, the dependency detection unit determines at least one target second instruction from at least one second instruction, including: the first detection unit determines the target second instruction from at least one second instruction being executed; wherein the second instruction being executed refers to a second instruction in any execution stage of the second pipeline.

[0210] In some embodiments, the first detection unit includes at least one comparator; the first detection unit determines the target second instruction from at least one second instruction being executed, including: the comparator determines the candidate second instruction as the target second instruction when the instruction address of the candidate second instruction is consistent with the instruction address of the first instruction; wherein the candidate second instruction refers to the second instruction in the execution stage of the second pipeline corresponding to the comparator.

[0211] In some embodiments, when the first instruction is in the first execution stage, the first detection unit determines the target second instruction from the second instructions being executed in at least one second pipeline; wherein the first execution stage refers to: an execution stage in which the instruction address of the first instruction has been determined in at least one execution stage included in the first pipeline.

[0212] In some embodiments, the timing processing unit includes an instruction reissuing unit; the timing processing unit controls the execution process of the first instruction and the target second instruction to meet the instruction execution timing, including: the instruction reissuing unit generates at least one clock bubble in the second pipeline, and the clock bubble is used to roll back the target second instruction to the first execution stage of the second pipeline.

[0213] In some embodiments, the first instruction is a storage instruction, and the second instruction is a load instruction; the dependency detection unit includes a second detection unit; after the first pipeline receives the first instruction, the dependency detection unit determines at least one target second instruction from at least one second instruction, including: the second detection unit determines the target second instruction from at least one second instruction that has been executed; wherein the second instruction that has been executed refers to the second instruction that has left the second pipeline.

[0214] In some embodiments, the second detection unit includes a selector; the second detection unit determines the target second instruction from at least one second instruction that has been executed, including: the selector determines the target second instruction from at least one second instruction that has been executed based on the instruction address of the first instruction and the timing information of the first instruction.

[0215] In some embodiments, the instruction execution device also includes: a mirror queue, the routing distance between the mirror queue and the first pipeline is smaller than the routing distance between the mirror queue and the second pipeline; the mirror queue stores the second instructions that have been executed; the second detection unit obtains at least one second instruction that has been executed through the mirror queue, so as to determine the target second instruction from the at least one second instruction that has been executed.

[0216] In some embodiments, there is a routing connection between the mirror queue and the second pipeline, and there is a routing connection between the mirror queue and the dependency detection unit; the instruction execution method also includes: the second pipeline sends the executed second instruction to the mirror queue, and the maximum number of executed second instructions that can be written in parallel to the mirror queue in one clock cycle is greater than or equal to the number of second pipelines included in the instruction execution device.

[0217] In some embodiments, the second detection unit also includes: a local oldest selection group, which has a wiring connection with the selector; the instruction execution method also includes: after the selector determines the target second instruction, the selector passes the target second instruction to the local oldest selection group; the local oldest selection group determines the starting re-execution instruction from the target second instruction, and the starting re-execution instruction refers to the target second instruction with the earliest execution timing.

[0218] In some embodiments, the local oldest selection group includes at least one level of local oldest selection unit; for the i-th level local oldest selection unit in the at least one level of local oldest selection unit, if the i-th level local oldest selection unit is the lowest-level local oldest selection unit, then the i-th level local oldest selection unit determines the i-th level local oldest instruction from multiple target selection instructions, where i is a natural number; if the i-th level local oldest selection unit corresponds to the i-1-th level local oldest selection unit, then the i-th level local oldest selection unit determines the i-th level local oldest instruction from the i-1-th level local oldest selection instruction; if the number of the i-th level local oldest instruction is 1, then the i-th level local oldest instruction is the starting re-execution instruction.

[0219] In some embodiments, the selector determines at least one target second instruction within one execution stage of the first instruction, and the local oldest selection group determines a starting re-execution instruction from at least one target second instruction within at most log2N execution stages of the first instruction, where N is the instruction capacity of the mirror queue.

[0220] In some embodiments, the timing processing unit includes a redirection unit; the timing processing unit controls the execution process of the first instruction and the target second instruction to meet the instruction execution timing, including: the redirection unit clears the second instruction in each execution stage of the second pipeline where the target second instruction is located, and controls the target second instruction to be re-executed from the instruction issuance stage.

[0221] The redirection unit includes at least one comparator, which is used to characterize the timing information of each second instruction of the second pipeline and the timing information of the starting re-execution instruction, determine at least one instruction to be re-executed, and control the at least one instruction to be re-executed to be re-executed from the instruction issuance stage. Optionally, the instruction issuance clock cycle of at least one instruction to be re-executed is later than the clock cycle of the starting re-execution instruction.

[0222] Optionally, if the timing information of a second instruction in the second pipeline is earlier than the timing information of the starting re-execution instruction, it is determined that the second instruction is not an instruction to be redirected; if the timing information of a second instruction in the second pipeline is later than the timing information of the starting re-execution instruction, it is determined that the second instruction is an instruction to be re-executed.

[0223] Optionally, the redirection unit sorts at least one target second instruction determined from at least one instruction to be redirected and the second instruction that has been executed according to the timing information to obtain a redirection instruction sequence; wherein, instructions with earlier timing information are arranged at the front of the redirection instruction sequence, instructions with later timing information are arranged at the back of the redirection instruction sequence, and the start re-execution instruction is arranged at the front of the redirection instruction sequence. The redirection unit sequentially controls the re-execution of each instruction included in the redirection instruction sequence from the instruction emission stage. The redirection unit preferentially controls the start re-execution instruction to be re-executed from the instruction emission stage.

[0224] Exemplarily, the slave instruction issue stage is a stage where an instruction issue unit generates instructions to a pipeline.

[0225] Optionally, each instruction included in the redirection instruction sequence can be respectively transmitted to different second pipelines, so as to improve the time required to complete the execution of each instruction included in the redirection instruction sequence and improve the execution efficiency of the instruction execution device on the instructions.

[0226] The specific mechanism of the redirection unit controlling the target second instruction to be re-executed starting from the instruction issuance stage can be referred to in the related art and is not limited here.

[0227] In some embodiments, the first instruction is a load instruction, and the second instruction is a store instruction; the dependency detection unit includes a third detection unit; after the first pipeline receives the first instruction, the dependency detection unit determines at least one target second instruction from at least one second instruction, including: the third detection unit determines the target second instruction from at least one unexecuted second instruction; wherein the unexecuted second instruction refers to the second instruction that has not been issued to the second pipeline.

[0228] In some embodiments, the timing processing unit includes a delay processing unit; the delay processing unit controls the first instruction to be suspended in the first pipeline until the target second instruction is completed in the second pipeline.

[0229] For the specific content of this embodiment, please refer to the above embodiment, which will not be repeated here.

[0230] The embodiment of the present application further provides a chip, the electronic chip comprising the instruction execution device as described above. Optionally, the chip comprises the instruction execution device and a memory.

[0231] The embodiment of the present application also provides an electronic device, which includes the instruction execution device as described above. Optionally, the electronic device includes a chip, which includes a memory and an instruction execution device. The instruction execution device is used to execute instructions using the solution provided in the above embodiment and complete the dependency detection process on the instructions.

[0232] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0233] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent switching, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An instruction execution device, characterized in that: The instruction execution device comprises: an instruction dispatch unit, a first pipeline, a second pipeline, a dependency detection unit and a timing processing unit; The instruction dispatch unit is used to send a first instruction to the first pipeline and send a second instruction to the second pipeline; wherein the first instruction is one of a store instruction and a load instruction, and the second instruction is the other of the store instruction and the load instruction; The dependency detection unit is configured to determine at least one target second instruction from at least one second instruction after the first pipeline receives the first instruction; wherein the target second instruction has the same instruction address as the first instruction and violates an instruction execution sequence, wherein the instruction execution sequence refers to an execution order between the first instruction and the target second instruction; The timing processing unit is used to control the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

2. The instruction execution device according to claim 1, characterized in that: The first instruction is the store instruction, and the second instruction is the load instruction; the dependency detection unit includes a first detection unit; The first detection unit is used to determine the target second instruction from at least one second instruction being executed; The second instruction being executed refers to a second instruction in any execution stage of the second pipeline.

3. The instruction execution device according to claim 2, characterized in that: The first detection unit includes at least one comparator; The comparator is configured to determine the candidate second instruction as the target second instruction when the instruction address of the candidate second instruction is consistent with the instruction address of the first instruction; The candidate second instruction refers to a second instruction in an execution stage of the second pipeline corresponding to the comparator.

4. The instruction execution device according to claim 2 or 3, characterized in that: The first detection unit is configured to determine the target second instruction from at least one second instruction being executed by the second pipeline when the first instruction is in the first execution stage; The first execution stage refers to: an execution stage in which the instruction address of the first instruction has been determined in at least one execution stage included in the first pipeline.

5. The instruction execution device according to claim 2, characterized in that: The timing processing unit includes an instruction reissuing unit; The instruction reissuing unit is used to generate at least one clock bubble in the second pipeline, and the clock bubble is used to roll back the target second instruction to the first execution stage of the second pipeline.

6. The instruction execution device according to claim 1, characterized in that: The first instruction is the store instruction, and the second instruction is the load instruction; the dependency detection unit includes a second detection unit; The second detection unit is used to determine the target second instruction from at least one second instruction that has been executed; The second instruction that has been executed is the second instruction that has left the second pipeline.

7. The instruction execution device according to claim 6, characterized in that: The second detection unit includes a selector; The selector is used to determine the target second instruction from the at least one executed second instruction based on the instruction address of the first instruction and the timing information of the first instruction.

8. The instruction execution device according to claim 7, characterized in that: The second detection unit further includes: a local oldest selection group, the local oldest selection group having a wiring connection with the selector; The selector is further configured to, after determining the target second instruction, pass the target second instruction to the local oldest selection group; The local oldest selection group is used to determine a starting re-execution instruction from the target second instruction, and the starting re-execution instruction refers to the target second instruction with the earliest execution timing.

9. The instruction execution device according to claim 6, characterized in that: The instruction execution device further includes: a mirror queue, wherein a routing distance between the mirror queue and the first pipeline is smaller than a routing distance between the mirror queue and the second pipeline; The mirror queue is used to store the executed second instructions; the second detection unit obtains the at least one executed second instruction through the mirror queue, so as to determine the target second instruction from the at least one executed second instruction.

10. The instruction execution device according to claim 9, characterized in that: A wiring connection is provided between the mirror queue and the second pipeline, and a wiring connection is provided between the mirror queue and the dependency detection unit; The second pipeline is used to send the executed second instruction to the mirror queue, and the maximum number of the executed second instructions that can be written in parallel by the mirror queue in one clock cycle is greater than or equal to the number of the second pipelines included in the instruction execution device.

11. The instruction execution device according to claim 10, characterized in that: The local oldest selection group includes at least one level of local oldest selection unit; For the i-th level local oldest selection unit in the at least one level local oldest selection unit, if the i-th level local oldest selection unit is the local oldest selection unit with the lowest level, the i-th level local oldest selection unit is used to determine the i-th level local oldest instruction from the multiple target selection instructions, where i is a natural number; If the i-th level local oldest selection unit corresponds to the i-1-th level local oldest selection unit, the i-th level local oldest selection unit is used to determine the i-th level local oldest instruction from the i-1-th level local oldest selection instruction; If the number of the oldest local instruction at the i-th level is 1, the oldest local instruction at the i-th level is the starting re-execution instruction.

12. The instruction execution device according to claim 10, characterized in that: The selector is used to determine at least one of the target second instructions within an execution stage of the first instruction, and the local oldest selection group is used to determine the starting re-execution instruction from the at least one target second instruction in at most log2N execution stages of the first instruction, where N is the instruction capacity of the mirror queue.

13. The instruction execution device according to claim 6, characterized in that: The timing processing unit includes a redirection unit; The redirection unit is used to clear the second instructions in each execution stage of the second pipeline where the target second instruction is located, and control the target second instruction to be re-executed starting from the instruction issuance stage.

14. The instruction execution device according to claim 1, characterized in that: The first instruction is the load instruction, and the second instruction is the store instruction; the dependency detection unit includes a third detection unit; The third detection unit is used to determine the target second instruction from at least one unexecuted second instruction; The unexecuted second instruction refers to a second instruction that has not been issued to the second pipeline.

15. The instruction execution device according to claim 14, characterized in that: The timing processing unit includes a delay processing unit; The delay processing unit is used to control the first instruction to be temporarily executed in the first pipeline until the target second instruction is completely executed in the second pipeline.

16. A chip, characterized in that: The chip comprises a memory and an instruction execution device as claimed in any one of claims 1 to 15.

17. An electronic device, characterized in that: The electronic device comprises the instruction execution device according to any one of claims 1 to 15.

18. An instruction execution method applied to an instruction execution device, characterized in that: The instruction execution device comprises: an instruction dispatch unit, a first pipeline, a second pipeline, a dependency detection unit and a timing processing unit; The instruction dispatch unit sends a first instruction to the first pipeline, and sends a second instruction to the second pipeline; wherein the first instruction is one of a store instruction and a load instruction, and the second instruction is the other of the store instruction and the load instruction; After the first pipeline receives the first instruction, the dependency detection unit determines at least one target second instruction from at least one of the second instructions; wherein the target second instruction has the same instruction address as the first instruction and violates an instruction execution sequence, wherein the instruction execution sequence refers to an execution order between the first instruction and the target second instruction; The timing processing unit controls the execution process of the first instruction and the target second instruction to meet the instruction execution timing.

Citation Information

Cited By

  • Instruction processing method, processor, electronic equipment and storage medium

    CN120578426A