Processor core, processor, computing device and instruction processing method
By writing multiple instructions to the transmit queue in the processor, and writing only the first instruction that does not meet the preset conditions to the ROB, while the second instruction that meets the conditions does not need to write the ROB, the problem of limited ROB capacity is solved and the performance of the processor core is improved.
Patent Information
- Application Number
- CN202411808836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-13
AI Technical Summary
In the superscalar processor for out-of-order execution, the capacity of the reorder cache (ROB) is limited, resulting in the upper limit of out-of-order execution being limited, and the bandwidth of the ROB submits instructions is limited, affecting the performance of the processor core.
By writing multiple instructions to the transmit queue and writing the first instruction to the ROB according to the original order, the second instruction that meets the preset conditions does not need to be written to the ROB, the instructions in the transmit queue are sent to the instruction execution unit for execution in an out-of-order manner, and when the oldest first instruction in the ROB is completed, it is submitted.
This improves the utilization rate of ROB and reduces the limitations of ROB on the processor core performance, thereby improving the performance of the processor core.
Smart Images

Figure CN119987858A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of chip technology, and in particular, to a processor core, a processor, a computing device, and an instruction processing method. Background Art
[0002] In an out-of-order superscalar processor, instructions are executed out of order in the processor, which can reduce pauses while maintaining data flow, thereby improving processor performance. Programs are executed sequentially. In order to ensure the correctness of program execution results, a reorder buffer (ROB) is set in the processor to maintain the order of instructions through the ROB.
[0003] At present, instructions are recorded in the ROB after being decoded by the Instruction Decoder Unit (IDU). When an instruction is executed and is the oldest instruction in the ROB, the instruction can be committed and leave the ROB in order, thus ensuring that the instruction execution process is out of order, and the commit is made in the order of the instructions in the program, thus ensuring the correctness of the program execution result.
[0004] However, due to the constraints of cost and chip area, the capacity of ROB is limited. ROB records the order and status of all instructions, resulting in the upper limit of out-of-order execution being limited by the capacity of ROB. In addition, the bandwidth for ROB to submit instructions is limited, and instructions cannot leave ROB in time, which will cause ROB to limit the performance of the processor core. Summary of the invention
[0005] In view of this, embodiments of the present disclosure provide a processor core, a processor, a computing device, and an instruction processing method to at least partially solve the above problems.
[0006] According to a first aspect of an embodiment of the present disclosure, a processor core is provided, comprising: an instruction decoding unit, for writing a plurality of instructions into a transmission queue, and writing a first instruction among the plurality of instructions into a reordering cache according to an original order of the plurality of instructions, the first instruction being an instruction among the plurality of instructions other than a second instruction that satisfies a preset condition; an instruction transmission unit, for sending the instructions in the transmission queue to an instruction execution unit for execution in a disorderly manner; and an instruction retirement unit, for submitting the first instruction that needs to be submitted first in the reordering cache after the first instruction is executed by the instruction execution unit, so as to delete the first instruction from the reordering cache.
[0007] According to a second aspect of an embodiment of the present disclosure, a processor is provided, comprising: at least one processor core as described in the first aspect above.
[0008] According to a third aspect of an embodiment of the present disclosure, a computing device is provided, comprising: at least one processor as described in the second aspect above, and a memory, the memory being coupled to the processor and used to store instructions to be executed.
[0009] According to a fourth aspect of an embodiment of the present disclosure, there is provided an instruction processing method, comprising: writing a plurality of instructions into an emission queue; writing a first instruction among the plurality of instructions into a reordering cache according to an original order of the plurality of instructions, the first instruction being an instruction among the plurality of instructions other than a second instruction that satisfies a preset condition; sending the instructions in the emission queue to an instruction execution unit for execution in a disorderly manner; and submitting the first instruction, which needs to be submitted first in the reordering cache, after the first instruction is executed by the instruction execution unit, so as to delete the first instruction from the reordering cache.
[0010] It can be seen from the above technical solution that multiple instructions are written into the emission queue, and the first instruction that meets the preset conditions among the multiple instructions is written into the ROB, while the second instruction that meets the preset conditions does not need to be written into the ROB, and the instructions in the emission queue are sent to the instruction execution unit in a disordered manner for execution, and when the oldest first instruction in the reorder cache has been completed, the first instruction is submitted. Since the second instruction that meets the preset conditions does not need to enter the reorder cache, the utilization rate of the reorder cache is improved, and the limiting effect of the reorder cache on the performance of the processor core is reduced, thereby improving the performance of the processor core. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present disclosure. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0012] Figure 1 is a schematic diagram of a computing device according to an embodiment of the present disclosure;
[0013] Figure 2 is a schematic diagram of a processor according to an embodiment of the present disclosure;
[0014] Figure 3 is a schematic diagram of a processor core according to an embodiment of the present disclosure;
[0015] Figure 4 is a schematic diagram of a processor core according to another embodiment of the present disclosure;
[0016] Figure 5 The figure is a flowchart of an instruction processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The present disclosure is described below based on embodiments, but the present disclosure is not limited to these embodiments. In the detailed description of the present disclosure below, some specific details are described in detail. It is possible for a person skilled in the art to fully understand the present disclosure without the description of these details. In order to avoid confusing the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. In addition, the drawings are not necessarily drawn to scale.
[0018] First, some nouns or terms that appear in the process of describing the embodiments of the present disclosure are subject to the following explanations.
[0019] Out-of-order execution: Out-of-order execution refers to the processor's dynamic scheduling of the instruction stream in the expectation that an instruction can be executed immediately when its operands and instruction execution units are available, without waiting for the execution of the instruction's predecessor.
[0020] Instruction issuance: Instruction issuance refers to the process of sending instructions to the instruction execution unit to start executing the instructions, which is usually implemented by the instruction issuance unit in the processor core.
[0021] Instruction commit: Instruction commit indicates that the instruction is allowed to modify the processor state to produce execution effects visible to the programming model, usually implemented by an instruction retirement unit in the processor core.
[0022] Instruction completion: The instruction execution unit returns the execution result (instruction completion information) except the destination register value. Instruction completion means that the instruction execution unit determines that the instruction can be executed. Instruction completion is usually implemented by the instruction retirement unit in the processor core.
[0023] Reorder Buffer: Reorder Buffer (ROB) is a mechanism used in computer architecture to implement out-of-order execution and sequential submission of instructions. The core idea of ROB is to record the order of instructions in the program. An instruction will not be submitted immediately after execution (submission here refers to modifying the processor state, such as modifying the logical register stack). Instead, it will wait in the cache (buffer) until all previous instructions have been submitted before modifying the processor state (such as submitting the result to the logical register stack).
[0024] Issue Queue: The Issue Queue is also called the Reservation Station, which is used to store instructions that have been renamed by registers but have not yet been sent to the Functional Unit (FU) for execution. The Issue Queue selects instructions whose source operands are ready according to certain rules, and sends these instructions to the FU for execution. This process is called issuance. The Issue Queue is the dividing line from sequential execution to out-of-order execution. Before the issue stage, instructions flow in the order specified in the program, and after the issue stage, instructions are executed out of order until the commit stage, when the instructions are pulled back to the original order specified in the program through the ROB.
[0025] Computing equipment
[0026] Figure 1 A schematic block diagram of a computing device 10 according to an embodiment of the present disclosure is shown. The computing device 10 can be constructed based on various types of processors and driven by any operating system such as WINDOWS operating system, UNIX operating system, Linux operating system, etc. In addition, the computing device 10 can be implemented in hardware and / or software such as a personal computer (PC), a desktop, a notebook, a server, and a mobile communication device.
[0027] like Figure 1 As shown, the computing device 10 may include one or more processors 12, and a memory 14. The memory 14 in the computing device 10 may be used as a main memory (abbreviated as main memory or internal memory) for storing instruction information and / or data information represented by data signals. For example, the memory 14 may store data provided by the processor 12, and may also be used to implement data exchange between the processor 12 and an external storage device 16 (or referred to as an auxiliary memory or external memory).
[0028] In some cases, the processor 12 needs to access the memory 14 through the bus 11 to obtain data in the memory 14 or modify the data in the memory 14. Since the access speed of the memory 14 is slow, in order to alleviate the speed gap between the processor 12 and the memory 14, the computing device 10 also includes a cache memory 18 connected to the bus 11 for caching some program data or message data in the memory 14 that may be repeatedly called. The cache memory 18 can be implemented by a storage device such as a static random access memory (SRAM). The cache memory 18 can be a multi-level structure, such as a three-level cache structure with a first-level cache (L1Cache), a second-level cache (L2Cache) and a third-level cache (L3Cache). The cache memory 18 can also be a cache structure of more than three levels or other types of cache structures. In some embodiments, a portion of the cache memory 18 (such as a first-level cache, or both a first-level cache and a second-level cache) can be integrated inside the processor 12 or integrated with the processor 12 in the same system on a chip.
[0029] The information exchange between the memory 14 and the cache memory 18 can be organized according to data blocks. In some embodiments, the cache memory 18 and the memory 14 can be divided into data blocks according to the same spatial size, and the data block can be used as the minimum unit of data exchange between the cache memory 18 and the memory 14 (including one or more data of a preset length). For the sake of simplicity and clarity, the data blocks in the cache memory 18 are referred to as cache blocks (or cache lines or cache lines), and different cache blocks have different cache block addresses. The data blocks in the memory 14 are referred to as memory blocks, and different memory blocks have different memory block addresses. The cache block address and / or the memory block address may include a physical address tag for locating the data block.
[0030] In addition, the computing device 10 may also include an external storage device 16, a display device, an audio device, a mouse / keyboard and other input / output devices. The external storage device 16 may be a hard disk, an optical disk, a flash memory and other devices for information access coupled to the bus 11 through a corresponding interface. The display device may be coupled to the bus 11 through a corresponding graphics card, and is used to display according to the display signal provided by the bus 11.
[0031] The computing device 10 may also include a communication device 17, so that the computing device 10 can communicate with a network or other devices in various ways. The communication device 17 may include one or more communication modules, and the communication device 17 may include a wireless communication module suitable for a specific wireless communication protocol. For example, the communication device 17 may include a WLAN module for implementing WiFi communication that complies with the 802.11 standard established by the Institute of Electrical and Electronics Engineers (IEEE). The communication device 17 may include a WWAN module for implementing wireless wide area communication that complies with cellular or other wireless wide area protocols. The communication device 17 may also include a communication module using other protocols such as a Bluetooth module, or a communication module of other custom types. The communication device 17 may also be a port for serial transmission of data.
[0032] It should be noted that the structure of the computing device 10 may vary according to the motherboard, operating system and instruction set architecture of different computing devices 10. For example, many computing devices are currently provided with an input / output control center connected between the bus 11 and various input / output devices, and the input / output control center may be integrated into the processor 12 or independent of the processor 12.
[0033] processor
[0034] Figure 2 FIG. 1 is a schematic block diagram of a processor 12 according to an embodiment of the present disclosure. Figure 2 As shown, each processor 12 may include one or more processor cores 120 for processing instructions, and the processing and execution of instructions may be controlled by the user (e.g., through an application) and / or the system platform. Each processor core 120 may be used to process a specific instruction set, and the instruction set may support complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW)-based computing. It should be particularly noted that the processor core 120 is suitable for processing the RISC-V instruction set. Different processor cores 120 may each process different or the same instruction sets. The processor 12 may also include other processing modules, such as a digital signal processor (DSP), etc. As an example, Figure 2 , processor core 1 to processor core m are shown, where m is a positive integer.
[0035] Figure 1The illustrated cache memory 18 may be fully or partially integrated into the processor 12. Depending on the architecture, the cache memory 18 may be a single or multiple levels of internal cache memory located within and / or outside each processor core 120 (e.g., Figure 2 The 3-level cache memory L1 to L3 shown, Figure 2 The processor 12 may include an instruction-oriented cache and an instruction cache and a data-oriented cache. The components in the processor 12 may share at least a portion of the cache memory. For example, the processor cores 1 to m may share the third-level cache memory L3. The processor 12 may also include an external cache (not shown). Other cache structures may also serve as the external cache of the processor 12.
[0036] like Figure 2 As shown, the processor 12 may include a register file 121. The register file 121 may include multiple registers for storing different types of data and / or instructions. These registers may be of different types. For example, the register file 121 may include integer registers, floating point registers, status registers, instruction registers, and pointer registers. The registers in the register file 121 may be implemented by general registers, or may be implemented by a specific design according to the actual needs of the processor 12.
[0037] The processor 12 may include a memory management unit (MMU) 122 for implementing the translation of virtual addresses to physical addresses. The memory management unit 122 caches a portion of the table entries in the page table, and the memory management unit 122 may also obtain uncached table entries from the memory. One or more memory management units 122 may be set in each processor core 120, and the memory management units 122 in different processor cores 120 may be synchronized with the memory management units 122 in other processors or processor cores, so that each processor or processor core can share a unified virtual storage system.
[0038] Processor core
[0039] like Figure 2 As shown, the processor 12 is used to execute an instruction sequence (i.e., a program). The process of the processor 12 executing each instruction includes: fetching the instruction from the memory storing the instruction, decoding the fetched instruction, executing the decoded instruction, and saving the instruction execution result, etc., and this cycle is repeated until all instructions in the instruction sequence are executed or a shutdown instruction is encountered.
[0040] In order to achieve the above process, Figure 3The schematic block diagram of the processor core 120 shown in FIG. 1 may include an instruction fetch unit 123 , an instruction decoding unit 124 , an instruction issuing unit 125 , an instruction executing unit 126 , an instruction retiring unit 127 , and the like.
[0041] The instruction fetch unit 124 acts as a startup engine for the processor core 120 and is used to transfer instructions from the memory 14 to the instruction register (which may be Figure 2 A register in the register stack 121 shown is used to store instructions) and receives the next instruction fetch address or calculates the next instruction fetch address according to an instruction fetch algorithm. The instruction fetch algorithm can be an incrementing address or a decrementing address according to the instruction length.
[0042] After fetching the instruction, the processor core 120 enters the instruction decoding stage, and the instruction decoding unit 124 decodes the fetched instruction according to the predetermined instruction format to obtain the operand acquisition information required by the fetched instruction, thereby preparing for the operation of the instruction execution unit 126. The operand acquisition information may include pointing to an immediate number, a register, or other software / hardware that can provide a source operand.
[0043] The instruction issuing unit 125 is usually present in the high-performance processor core 120, and is located between the instruction decoding unit 124 and the instruction execution unit 126, and is used for scheduling and controlling instructions, so as to efficiently allocate each instruction to different instruction execution units 126, so that the parallel operation of multiple instructions becomes possible. After the instruction is fetched, decoded and scheduled to the corresponding instruction execution unit 126, the corresponding instruction execution unit 126 starts to execute the instruction, that is, to execute the operation indicated by the instruction and realize the corresponding function.
[0044] The instruction retirement unit 127 (also called the instruction retirement unit or the instruction write-back unit) is mainly used to write the execution results generated by the instruction execution unit 126 back to the corresponding storage location (for example, the register inside the processor core 120) so that subsequent instructions can quickly obtain the corresponding execution results from the storage location.
[0045] For different types of instructions, different instruction execution units 121 may be provided in the processor core 120 accordingly. The instruction execution unit 126 may be a computing unit (e.g., including an arithmetic logic unit, a shaping processing unit, a vector computing unit, etc., for performing operations according to operands and outputting operation results), a memory execution unit (e.g., for accessing a memory according to an instruction to read data in the memory or write specified data to the memory, etc.), a coprocessor, etc. In the processor core 120, each instruction execution unit 126 may run in parallel and output corresponding execution results.
[0046] When executing certain types of instructions (e.g., memory access instructions), the instruction execution unit 126 needs to access the memory 14 to obtain information stored in the memory 14 or provide data to be written into the memory 14. It should be noted that the instruction execution unit 126 for executing memory access instructions may also be referred to as a memory execution unit, which may be a load store unit (LSU) and / or other units for memory access.
[0047] After the memory access instruction is fetched by the instruction fetch unit 123, the instruction decoding unit 124 can decode the memory access instruction so that the source operand of the memory access instruction can be obtained. The decoded memory access instruction is provided to the corresponding instruction execution unit 126, and the instruction execution unit 126 can perform corresponding operations on the source operand of the memory access instruction (for example, the arithmetic logic unit performs operations on the source operand stored in the register) to obtain the address information corresponding to the memory access instruction, and initiate corresponding requests according to the address information, such as address translation requests, write access requests, etc.
[0048] The source operand of the memory access instruction usually includes an address operand, and the instruction execution unit 126 operates on the address operand to obtain the virtual address or physical address corresponding to the memory access instruction. When the memory management unit 122 is disabled, the instruction execution unit 126 can directly obtain the physical address of the memory access instruction through logical operations. When the memory management unit 122 is enabled, the corresponding instruction execution unit 126 initiates an address translation request according to the virtual address corresponding to the memory access instruction, and the address translation request includes the virtual address corresponding to the address operand of the memory access instruction; the memory management unit 122 responds to the address translation request, and converts the virtual address in the address translation request into a physical address according to the table entry matching the virtual address, so that the instruction execution unit 126 can access the cache memory 18 and / or the memory 14 according to the translated physical address.
[0049] Depending on the function, memory access instructions may include load instructions and store instructions. The execution process of the load instruction usually does not require modification of the information in the memory 14 or the cache memory 18. The instruction execution unit 126 only needs to read the data stored in the memory 14, the cache memory 18 or the external storage device according to the address operand of the load instruction. Different from the load instruction, the source operand of the store instruction includes not only the address operand, but also the data information. The execution process of the store instruction usually requires modification of the information in the memory 14 and / or the cache memory 18. The data information of the store instruction can point to the write data, and the source of the write data can be the execution result of the operation instruction, the load instruction and other instructions, or it can be the data provided by the register or other storage unit in the processor 12, or it can be an immediate number.
[0050] For the processor 12 of the out-of-order execution architecture, the processor core 120 may further include an instruction distribution unit, which may convert the instruction from the in-order state to the out-of-order state and assign a reordering number (instruction ID) to the instruction, so that the instruction execution unit 126 and the instruction retirement unit 127 may determine the new and old relationship of the instruction according to the reordering number. After the instruction distribution unit converts the instruction to the out-of-order state, the instruction emission unit 125 may emit the instruction in the out-of-order state to the instruction execution unit 126, and the instruction execution unit 126 then executes the instruction out-of-order.
[0051] The disclosed embodiments mainly focus on the process of writing instructions by the instruction decoding unit 124, the instruction issuing process by the instruction issuing unit 125, and the instruction submitting process by the instruction retiring unit 127. The instruction issuing and submitting processes will be described in detail later.
[0052] like Figure 4 In the schematic block diagram of the processor core 120 shown in FIG. 1 , the instruction decoding unit 124 can write multiple instructions into the emission queue, and write the first instruction of the multiple instructions into the reorder buffer (ROB) according to the original order of the multiple instructions. The multiple instructions include a first instruction and a second instruction, the second instruction is an instruction in the multiple instructions that meets a preset condition, and the instructions in the multiple instructions other than the second instruction are the first instructions. The instruction emission unit 125 can send the multiple instructions in the emission queue to the instruction execution unit 126 for execution in a disorderly manner. The instruction retirement unit 127 can commit the first instruction that needs to be committed first in the ROB after the instruction execution unit 126 completes execution of the first instruction, so as to delete the first instruction from the ROB.
[0053] Since the instruction emission unit 125 can send the instructions in the emission queue to the instruction execution unit 126 for execution in a disorderly manner, the actual execution order of the instructions may be different from the original order in the program, but when the instructions are submitted, they need to be submitted in the order in the program to ensure the accuracy of the program execution result. However, the submission order of the second instruction that meets the preset condition does not affect the execution result of the program, and the submission order of the first instruction that does not meet the preset condition will affect the execution result of the program, so only the first instruction can be written into the ROB, so that the instruction retirement unit 127 submits the first instruction in sequence.
[0054] The ROB records the status information of the first instruction, and the status information indicates whether the corresponding first instruction has been executed by the instruction execution unit 126. The storage location of the first instruction in the ROB can indicate the relative execution order between the first instructions in the program. After the instruction execution unit 126 executes the first instruction, it will generate completion information (complete) for the first instruction, and the status information of the first instruction in the ROB will be updated to completion based on the completion information. The instruction retirement unit 127 can commit the first instruction with a completed status and the oldest instruction in the ROB according to the status information recorded in the ROB and the storage location of the first instruction. The committed first instruction leaves the ROB, and the first instruction to be committed is deleted from the ROB. After the first instruction is deleted from the ROB, the status information of the first instruction is also deleted from the ROB.
[0055] Since the second instruction is not stored in the ROB, the second instruction does not involve a commit process. After the second instruction is executed by the instruction execution unit 126, a write-back process may be performed to release resources except the physical register number (Physical Tag, ptag).
[0056] In the embodiment of the present disclosure, the instruction decoding unit 124 writes multiple instructions into the emission queue, and writes the first instruction that does not meet the preset condition among the multiple instructions into the ROB, while the second instruction that meets the preset condition among the multiple instructions does not need to be written into the ROB, and the instruction emission unit 125 sends the multiple instructions in the emission queue to the instruction execution unit 126 for execution in a disordered manner, and when the oldest first instruction in the ROB has been completed, the instruction retirement unit 127 submits the first instruction. Since the second instruction that meets the preset condition does not need to enter the ROB, the utilization rate of the ROB is improved, and the limiting effect of the ROB on the performance of the processor core 120 is reduced, thereby improving the performance of the processor core 120.
[0057] In a possible implementation, the preset conditions include the following five conditions:
[0058] Condition 1: No error will be reported during instruction execution;
[0059] Condition 2: The instruction will not cause other instructions to execute incorrectly;
[0060] Condition 3: The source operand of the instruction comes from an immediate value or a physical register, and the instruction execution result is recorded in the physical register;
[0061] Condition 4: Based on the original order of instructions in the program, the current instruction and the previous instruction are both executed and committed or neither is committed;
[0062] Condition 5: The instruction is not the last instruction in the instruction block.
[0063] For an instruction in the issue queue, if the instruction satisfies the above five conditions, the instruction is the second instruction, and the instruction decoding unit 124 will not write the instruction into the ROB. For an instruction in the issue queue, if the instruction does not satisfy any one or more of the above conditions, the instruction is the first instruction, and the instruction decoding unit 124 will write the instruction into the ROB.
[0064] For the above condition 1, no error will be reported during the execution of the instruction, that is, the instruction will not fault. For the above condition 2, the instruction will not cause other instructions to execute errors, that is, the instruction will not issue a flush. For the above condition 3, the source operand of the instruction comes from the immediate value or the physical register indicated by the ptag, and the execution results of the instruction are all recorded in the physical register indicated by the ptag. For the above condition 4, the current instruction and the previous instruction will be executed and submitted by the processor, or neither will be submitted, that is, the previous instruction will not flush next. If the previous instruction is submitted and the current instruction is not submitted, the current instruction does not meet the above condition 4. If the previous instruction and the current instruction are both executed by the processor but neither is submitted, the current instruction does not meet the above condition 4. For the above condition 5, the instruction is not the last instruction in the program block, that is, the instruction is not EOB (End of Block).
[0065] The instruction decoding unit 124 can classify the instructions in the issue queue according to the above five conditions to determine whether the instruction belongs to the first instruction or the second instruction. The instruction decoding and issuing unit 124 will not write the second instruction into the ROB, but will write the first instruction into the ROB.
[0066] In the disclosed embodiment, the submission order of the second instruction that meets the above five conditions will not affect the execution result of the program, that is, the submission order of the second instruction is different from the original order specified in the program, and will not affect the correctness of the program execution result, so after the second instruction is executed, there is no need to pull it back to the original order specified in the program, and the instruction is written to the ROB to pull the out-of-order executed instructions back to the original order specified in the program, so there is no need to write the second instruction to the ROB. The second instruction is screened out by the above five conditions to ensure that the out-of-order submission of the screened second instruction will not affect the execution result of the program, while improving the utilization rate of the ROB and ensuring the correctness of the program execution result.
[0067] In a possible implementation, the instructions satisfying the above five conditions include arithmetic logic operation instructions and floating point operation instructions, the arithmetic logic operation instructions are executed by an arithmetic logic unit (ALU), and the floating point operation instructions are executed by a floating point unit (FPU). In an example, the second instruction may be an arithmetic logic operation instruction satisfying the above five conditions.
[0068] Among the instructions executed by the processor core 120, a large proportion of the instructions are arithmetic logic operation instructions. For example, in the standardized benchmark test suite SPECint2017, about 50% of the instructions are arithmetic logic operation instructions. By taking the instructions in the arithmetic logic operation instructions that meet the above five conditions as the second instructions, a larger number of second instructions can be determined, so that a larger number of instructions do not need to enter the ROB, thereby improving the utilization rate of the ROB.
[0069] In the disclosed embodiment, since the program includes a large number of arithmetic and logical operation instructions, the instructions in the arithmetic and logical operation instructions that meet the above five conditions are determined as the second instructions. This can simplify the logic of instruction classification, reduce the resource consumption of instruction classification, and determine a large number of second instructions, thereby greatly improving the utilization rate of ROB.
[0070] In a possible implementation, when the instruction decoding unit 124 writes the first instruction into the ROB, it may write the first instruction into the ROB according to the original order corresponding to the first instruction, so that the storage position of the first instruction in the ROB can indicate the relative execution order of the first instruction in the program.
[0071] The instruction decoding unit 124 can write the first instruction to the corresponding position in the ROB according to the original order of the first instruction in the program. Different write positions in the ROB correspond to different reordering numbers (ROB ID, Rid), so that different first instructions correspond to different Rids, and the relative execution order of the first instructions in the program can be determined according to the corresponding Rids. For example, the first instruction executed first in the program corresponds to a smaller Rid, and the first instruction executed later in the program corresponds to a smaller Rid.
[0072] In one example, the instruction decoding unit 124 can obtain the Rid corresponding to the first instruction when writing the first instruction into the ROB, and write the corresponding Rid into the emission queue when writing the first instruction into the emission queue, so that the instruction execution unit 126 and the instruction retirement unit 127 can determine the new and old relationship of the first instruction according to Rid. Since different first instructions are stored in different storage locations in the ROB, different first instructions correspond to different Rids, and the size relationship of Rid can indicate the relative execution order of the corresponding first instructions in the program. In one example, Rid is an increasing integer, and the first instruction corresponding to the smaller Rid in the program is executed before the first instruction corresponding to the larger Rid. For example, in the original order specified by the program, the first instruction 1 is before the first instruction 2, then the Rid of the first instruction 1 is less than the Rid of the first instruction 2.
[0073] In the disclosed embodiment, the instruction decoding unit 124 writes the first instruction into the corresponding storage location in the ROB according to the original order of the first instruction in the program, and different storage locations in the ROB correspond to different Rids, so that different first instructions correspond to different Rids, and thus the Rid can indicate the relative execution order of the corresponding first instructions in the program. Therefore, the instruction retirement unit 127 can determine the oldest first instruction in the ROB according to the storage position of the first instruction in the ROB, that is, according to the Rid corresponding to the first instruction, and then after the oldest first instruction in the ROB is completed, the first instruction can be submitted to ensure that the first instructions are submitted in the original order specified by the program, thereby ensuring the correctness of the program execution.
[0074] In one possible implementation, when the instruction decoding unit 124 writes an instruction into the transmit queue, the Rid corresponding to the instruction is also recorded in the transmit queue. Since both the first instruction and the second instruction will be sent to the transmit queue, the transmit queue records the Rid of each first instruction and second instruction in the transmit queue.
[0075] The instruction decoding unit 124 determines the Rid of the first instruction according to the storage position of the first instruction in the ROB, and then when the instruction cryptographic unit 124 writes the first instruction into the emission queue, the Rid of the first instruction can be recorded into the emission queue. The instruction decoding unit 124 assigns the same Rid as the previous first instruction to the second instruction, that is, the second instruction has the same Rid as the previous first instruction of the second instruction, and then when the instruction decoding unit 124 writes the second instruction into the emission queue, the Rid of the second instruction can be recorded into the emission queue. For example, the emission queue includes instructions 0, 1, 2, 3, and 4 from front to back in the corresponding original order, instructions 0 and 3 are the first instructions, instructions 1, 2, and 4 are the second instructions, the Rid of instruction 0 is 0, the Rid of instruction 3 is 1, then the Rid of instruction 1 and instruction 2 is 0, and the Rid of instruction 4 is 1.
[0076] In one example, as shown in Table 1 below, the program includes 4 instructions in the order from front to back in the original order in the program, the first instruction is a load instruction (ld) / store instruction (st) / other instructions, the second instruction is an arithmetic logic operation instruction (alu), the third instruction is an arithmetic logic operation instruction (alu), and the fourth instruction is a branch jump instruction (br). Since the first instruction does not flush or can only flush self, the second instruction satisfies the above condition 4, since the second instruction is not the last instruction in the program block, the second instruction satisfies the above condition 5, and since the second instruction is an arithmetic logic operation instruction, the second instruction satisfies the above conditions 1, 2 and 3, so it can be determined that the second instruction is the second instruction, and similarly, the third instruction can be determined to be the second instruction. Since both the first instruction and the fourth instruction do not meet the above conditions 1-3, it is determined that the first instruction and the fourth instruction are both the first instruction. The Rid of the first instruction is 0. Since the second and third instructions are the second instructions, the Rid of the second and third instructions is the same as the Rid of the first instruction. Rid increases with a step size of 1, and the Rid of the fourth instruction is 1. The second and third instructions will not be written into the ROB, and do not occupy the space of the ROB. The ROB does not need to collect the completion information of these two instructions.
[0077] Table 1
[0078] instruction Rid illustrate ld / st / others 0 No flush occurs or only self is flushed alu 0 This instruction does not enter the ROB alu 0 This instruction does not enter the ROB br 1
[0079] In another example, as shown in Table 2 below, the program includes 4 instructions in the order from front to back in the original order in the program, and these 4 instructions are all arithmetic logic operation instructions (alu). The Rid of the first instruction is 0. Since the second instruction meets the above 5 conditions, the Rid of the second instruction is the same as the Rid of the first instruction. Since the third instruction is the last instruction in the program block, that is, the third instruction is EOB, the Rid of the third instruction is 1. Since the fourth instruction meets the above 5 conditions, the Rid of the fourth instruction is the same as the Rid of the third instruction. The third instruction is EOB, this instruction will be written into the ROB and a new Rid will be assigned to it, but there is no need to collect the completion information (complete) of the instruction. When this instruction is the oldest instruction in the ROB, it can leave the ROB directly. Due to the existence of EOB, some processing of some dependent instructions of the processor core 120 will not fail to respond.
[0080] Table 2
[0081]
[0082] After the instruction issuing unit 125 sends the instruction in the issuing queue to the instruction executing unit 126 , the instruction and the Rid of the instruction are deleted from the issuing queue.
[0083] In the disclosed embodiment, since the instruction decoding unit 124 will not write the second instruction into the ROB, the second instruction will not be submitted in the original order of the instructions in the program, but the resources of the first instruction and the second instruction (the ptag of the physical register used to store the instruction execution result) need to be released in the corresponding original order in the program. Therefore, according to the original order of the instructions in the program, the same Rid is assigned to the second instruction and the previous first instruction of the second instruction, and then according to the Rid of the first instruction in the ROB and the Rid of the instruction in the emission queue, the oldest instruction in the pipeline can be determined, and then the ptag can be released based on the oldest instruction in the pipeline, so that the ptag resources are correctly recovered, thereby ensuring the correctness of program execution.
[0084] In one possible implementation, the instruction retirement unit 127 can determine the first Rid of the first instruction that corresponds to the original order and is at the first position in the ROB, that is, determine the Rid of the oldest first instruction in the ROB as the first Rid, and can determine the second Rid of the second instruction that corresponds to the original order and is at the first position in the emission queue, that is, determine the Rid of the oldest second instruction in the emission queue as the second Rid.
[0085] The instruction retirement unit 127 can determine the Rid with the previous corresponding instruction in the first Rid and the second Rid as the third Rid according to the corresponding original order, that is, determine the Rid with the older corresponding instruction in the first Rid and the second Rid as the third Rid. It should be noted that the first Rid and the second Rid can be the same, in which case the first Rid, the second Rid and the third Rid are the same.
[0086] After the instruction retirement unit 127 determines the third Rid, the corresponding instruction is determined as the fourth Rid according to the corresponding original order. For example, according to the corresponding original order from front to back, instructions 0, 1, 2, 3 and 4, instructions 0 and 3 are the first instructions, instructions 1, 2 and 4 are the second instructions, and the Rid of instruction 3 is determined as the third Rid, then the Rid of instruction 2 is determined as the fourth Rid. Since the Rids of instructions 0, 1 and 2 are the same, the Rids of instructions 0, 1 and 2 are all the fourth Rid.
[0087] In one example, according to the corresponding original order from front to back, the Rid of the first instruction is an increasing integer sequence with a step length of 1. If the third Rid is 5, the fourth Rid is 4.
[0088] After the instruction retirement unit 127 determines the fourth Rid, the ptag occupied by the instruction using the same logical register number (logic tag, ltag) as the instruction indicated by the fourth Rid is released. In one example, the program includes instructions 1 to 100 in the corresponding original order from front to back. When the Rid of instruction 60 is determined to be the fourth Rid, it is determined that instruction 6 and instruction 60 use the same ltag, then the ptag occupied by instruction 6 is released, and then the instruction after instruction 60 can use the ptag previously occupied by instruction 6, such as assigning the ptag previously occupied by instruction 6 to instruction 61 for storing the execution result of instruction 61.
[0089] In one example, the instruction retirement unit 127 may broadcast the fourth Rid to release the ptag occupied by the instruction that uses the same ltag as the instruction indicated by the fourth Rid.
[0090] It should be noted that, since there are the first and second instructions corresponding to the same Rid, the instruction indicated by the fourth Rid may be one or more instructions. If the fourth Rid indicates multiple instructions, the ptags occupied by the instructions using the same ltag as each of the multiple instructions will be released respectively.
[0091] In the embodiment of the present disclosure, after the third Rid is determined based on the first Rid and the second Rid, the instruction indicated by the third Rid will leave the pipeline within the expected limited time. The instruction that is older than the instruction indicated by the third Rid and indicated by the fourth Rid has already left the pipeline, so the ptag occupied by the older instruction that uses the same ltag as the instruction indicated by the fourth Rid can be released, ensuring that the ptag resources can be recovered in time, and the ptag used by the instruction is returned after the instruction leaves the pipeline, thereby ensuring the correctness of program execution.
[0092] In one possible implementation, an age matrix is provided in the transmission queue, and the new and old relationship of any two instructions in the transmission queue can be compared according to the age matrix, so that the instruction retirement unit 127 can determine the second oldest instruction in the transmission queue as the target instruction according to the age matrix, that is, determine the second instruction at the first position in the corresponding original order in the transmission queue as the target instruction, and then determine the Rid of the target instruction as the second Rid.
[0093] In the embodiment of the present disclosure, an age matrix is added to the transmission queue, and the age matrix can record the new and old relationship between any two instructions in the transmission queue, so that the instruction retirement unit 127 can determine the oldest second instruction in the transmission queue through the age matrix, and then determine the Rid of the oldest second instruction in the transmission queue as the second Rid, thereby ensuring the accuracy of the determined second Rid.
[0094] In a possible implementation, after submitting the first instruction, the instruction retirement unit 127 may release other resources occupied by the first instruction except for the ptag.
[0095] After the instruction retirement unit 127 submits the first instruction, the resources such as the instruction counter cache (Program Counter Buffer, PCB) occupied by the first instruction can be released. The instruction counter cache is used to save the instruction counter (Program Counter, PC) and is used to indicate that the ptag of the physical register storing the execution result of the first instruction is not released temporarily. Subsequently, the ptag occupied by the first instruction is released by determining the fourth Rid in accordance with the above embodiment.
[0096] It should be noted that the second instruction releases the resources occupied by it except the ptag after leaving the pipeline. In various embodiments of the present disclosure, unless otherwise specified, the ptag occupied by an instruction refers to the ptag used to store the execution result of the instruction.
[0097] In the disclosed embodiment, after the instruction retirement unit 127 submits the first instruction, the resources occupied by the first instruction except the ptag have been used up, and the resources occupied by the first instruction except the ptag are released in time, which can improve the performance of the processor core 120. The ptag resources occupied by the submitted first instruction are released by the fourth Rid in the above embodiment, which can ensure the correctness of the program execution result.
[0098] In the above embodiments, it is not necessary to write the second instruction into the ROB, it is not necessary to collect the completion information (complete) of the second instruction, and the processor core 120 does not need to perform these operations, which can save the power consumption of the processor core 120. The second instruction does not enter the ROB, nor does it occupy the bandwidth discharged by the ROB, which can increase the range of the Point Of No Return (PNR) protection instruction, thereby improving the performance of the processor core 120. Since the second instruction does not enter the ROB, the use requirements of the ROB can be met by a ROB with a smaller storage space, and a ROB with a smaller storage space can reduce the area of the ROB on the chip.
[0099] Instruction processing method
[0100] Figure 5 FIG. 1 is a flowchart of an instruction processing method according to an embodiment of the present disclosure, and the instruction processing method may be executed by the processor core 120 in the above embodiment. Figure 5 As shown, the instruction processing method includes the following steps:
[0101] Step 501, write multiple instructions into the transmission queue;
[0102] Step 502: write a first instruction among the multiple instructions into a reordering cache according to the original order of the multiple instructions, wherein the first instruction is an instruction among the multiple instructions except a second instruction that meets a preset condition;
[0103] Step 503: Send the instructions in the transmit queue to the instruction execution unit in a disorderly manner for execution;
[0104] Step 504: After the first instruction that needs to be committed first in the reorder cache is executed by the instruction execution unit, the first instruction is committed to delete the first instruction from the reorder cache.
[0105] In the embodiment of the present disclosure, multiple instructions are written into the emission queue, and the first instruction that meets the preset condition among the multiple instructions is written into the ROB, while the second instruction that meets the preset condition does not need to be written into the ROB, and the instructions in the emission queue are sent to the instruction execution unit in a disordered manner for execution, and when the oldest first instruction in the ROB has been completed, the first instruction is submitted. Since the second instruction that meets the preset condition does not need to enter the ROB, the utilization rate of the ROB is improved, and the limiting effect of the ROB on the performance of the processor core 120 is reduced, thereby improving the performance of the processor core 120.
[0106] In one possible implementation, the preset conditions include: no error is reported during the execution of the instruction; the instruction will not cause execution errors of other instructions; the source operand of the instruction comes from an immediate value or a physical register, and the execution result of the instruction is recorded in the physical register; according to the original order of the instructions in the program, the current instruction and the previous instruction are both executed and submitted or neither is submitted; the instruction is not the last instruction in the program block.
[0107] In the disclosed embodiment, the submission order of the second instruction that meets the above five conditions will not affect the execution result of the program, that is, the submission order of the second instruction is different from the original order specified in the program, and will not affect the correctness of the program execution result, so after the second instruction is executed, there is no need to pull it back to the original order specified in the program, and the instruction is written to the ROB to pull the out-of-order executed instructions back to the original order specified in the program, so there is no need to write the second instruction to the ROB. The second instruction is screened out by the above five conditions to ensure that the out-of-order submission of the screened second instruction will not affect the execution result of the program, while improving the utilization rate of the ROB and ensuring the correctness of the program execution result.
[0108] In a possible implementation, when writing the first instruction into the ROB, the first instruction may be written into the ROB according to the original order of the instructions, so that the storage position of the first instruction in the ROB may indicate the relative execution order of the first instruction in the program. After the instruction in the emission queue is sent to the instruction execution unit, the instruction and the Rid of the instruction may be deleted from the emission queue, wherein, according to the corresponding original order, the reordering number of the second instruction is the same as the previous first instruction of the second instruction, and the Rid of the first instruction is used to indicate the storage position of the first instruction in the ROB.
[0109] In the disclosed embodiment, the first instructions are written into corresponding storage locations in the ROB according to their original order in the program, and different storage locations in the ROB correspond to different Rids, so that different first instructions correspond to different Rids, and thus the Rid can indicate the relative execution order of the corresponding first instructions in the program. Therefore, according to the storage position of the first instruction in the ROB, that is, according to the Rid corresponding to the first instruction, the oldest first instruction in the ROB can be determined, and then after the oldest first instruction in the ROB is completed, the first instruction can be submitted, ensuring that the first instructions are submitted in the original order specified by the program, thereby ensuring the correctness of program execution.
[0110] In one possible implementation, the ptag occupied by the instruction can be released by the following method:
[0111] S1, determine the first Rid of the first instruction in the ROB corresponding to the first instruction in the original order;
[0112] S2, determining a second Rid corresponding to a second instruction at the first position in the original order in the transmit queue;
[0113] S3, according to the corresponding original order, determine the Rid with the corresponding instruction preceding the first Rid and the second Rid as the third Rid;
[0114] S4, determining the Rid of the previous instruction of the first instruction indicated by the third Rid as the fourth Rid;
[0115] S5. Release the physical register number occupied by the instruction that uses the same logical register number as the instruction indicated by the fourth Rid.
[0116] In the disclosed embodiment, after the third Rid is determined based on the first Rid and the second Rid, the instruction indicated by the third Rid will leave the pipeline within the expected limited time. For example, the instruction indicated by the third Rid is older and the instruction indicated by the fourth Rid has already left the pipeline, so the ptag occupied by the older instruction using the same ltag as the instruction indicated by the fourth Rid can be released to ensure that the ptag resources can be recovered in time and the ptag used by the instruction is returned after the instruction leaves the pipeline to ensure the correctness of program execution.
[0117] It should be noted that, since the details of the instruction processing method have been described in detail in the above processor core embodiment in conjunction with the structural diagram, the specific process can be found in the description of the above processor core embodiment, and will not be repeated here.
[0118] It should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0119] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the method embodiment, since it is basically similar to the method described in the device and system embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of other embodiments.
[0120] It should be understood that the above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] It should be understood that an element described in singular form herein or shown in the drawings as only one does not limit the number of the element to one. In addition, modules or elements described or shown as separate herein may be combined into a single module or element, and modules or elements described or shown as single herein may be split into multiple modules or elements.
[0122] It should also be understood that the terms and expressions used herein are for description only, and one or more embodiments of this specification should not be limited to these terms and expressions. The use of these terms and expressions does not mean to exclude any equivalent features of the illustrations and descriptions (or parts thereof), and it should be recognized that various modifications that may exist should also be included in the scope of the claims. Other modifications, changes and substitutions may also exist. Accordingly, the claims should be deemed to cover all such equivalents.
Claims
1. A processor core, comprising: an instruction decoding unit, configured to write a plurality of instructions into an issue queue, and write a first instruction of the plurality of instructions into a reordering cache according to an original order of the plurality of instructions, wherein the first instruction is an instruction of the plurality of instructions other than a second instruction that meets a preset condition; An instruction issuing unit, used to send the instructions in the issuing queue to the instruction executing unit for execution in a disorderly manner; The instruction retirement unit is used to submit the first instruction that needs to be submitted first in the reorder cache after the first instruction is executed by the instruction execution unit, so as to delete the first instruction from the reorder cache.
2. The processor core according to claim 1, wherein: The preset conditions include: No errors will be reported during the execution of the instruction; The instruction will not cause other instructions to be executed incorrectly; The source operand of the instruction comes from an immediate value or a physical register, and the execution result of the instruction is recorded in the physical register; According to the original order, the current instruction and the previous instruction are both executed and submitted or neither is submitted; The instruction is not the last instruction in the block.
3. The processor core according to claim 2, wherein: The second instruction includes an arithmetic logic operation instruction that meets the preset condition.
4. The processor core according to claim 1, wherein: The instruction decoding unit is used to write the first instruction into the reorder cache according to the corresponding original order, so that the storage position of the first instruction in the reorder cache indicates the relative execution order of the first instruction in the program.
5. The processor core according to claim 4, wherein: The instruction issuing unit is used to delete the instruction and the reordering number of the instruction from the issuing queue after the instruction in the issuing queue is sent to the instruction execution unit, wherein, according to the corresponding original order, the second instruction has the same reordering number as the previous first instruction of the second instruction, and the reordering number of the first instruction is used to indicate the storage position of the first instruction in the reordering cache.
6. The processor core according to claim 5, wherein: The instruction retirement unit is used to determine the first reordering number of the first instruction corresponding to the original order in the reordering cache, and determine the second reordering number of the second instruction corresponding to the original order in the transmission queue; according to the corresponding original order, the reordering number of the instruction preceding the first reordering number and the second reordering number is determined as the third reordering number, and the reordering number of the previous instruction corresponding to the first instruction indicated by the third reordering number is determined as the fourth reordering number, and the physical register number occupied by the instruction using the same logical register number as the instruction indicated by the fourth reordering number is released.
7. The processor core according to claim 6, wherein: The instruction retirement unit is used to determine the second instruction at the first position in the corresponding original order in the transmission queue as the target instruction according to the age matrix set in the transmission queue, and determine the reordering number of the target instruction as the second reordering number.
8. The processor core according to any one of claims 1 to 7, wherein: The instruction retirement unit is used to release the resources occupied by the first instruction except the physical register number after the first instruction is submitted.
9. A processor, comprising: At least one processor core according to any one of claims 1-8.
10. A computing device comprising: at least one processor according to claim 9; The memory is coupled to the processor and stores instructions to be executed.
11. A method for processing an instruction, comprising: Write multiple instructions into the issue queue; Writing a first instruction among the plurality of instructions into a reordering cache according to an original order of the plurality of instructions, the first instruction being an instruction among the plurality of instructions except a second instruction that meets a preset condition; Sending the instructions in the transmit queue to the instruction execution unit for execution in a disorderly manner; After the first instruction that needs to be committed first in the reorder cache is executed by the instruction execution unit, the first instruction is committed to delete the first instruction from the reorder cache.
12. The method according to claim 11, wherein: The preset conditions include: No errors will be reported during the execution of the instruction; The instruction will not cause other instructions to be executed incorrectly; The source operand of the instruction comes from an immediate value or a physical register, and the execution result of the instruction is recorded in the physical register; According to the original order, the current instruction and the previous instruction are both executed and submitted or neither is submitted; The instruction is not the last instruction in the block.
13. The method according to claim 11, wherein: Writing the first instruction of the plurality of instructions into the reorder cache according to the original order of the plurality of instructions comprises: writing the first instruction into the reorder cache according to the corresponding original order so that the position of the first instruction in the reorder cache indicates the relative execution order of the first instruction in the program; The method also includes: after the instruction in the transmission queue is sent to the instruction execution unit, deleting the instruction and the reordering number of the instruction from the transmission queue, wherein, according to the corresponding original order, the second instruction has the same reordering number as the previous first instruction of the second instruction, and the reordering number of the first instruction is used to indicate the storage position of the first instruction in the reordering cache.
14. The method according to claim 13, further comprising: Determine a first reordering number of the first instruction in the reordering cache corresponding to the first instruction in the original order; Determine a second reordering number of the second instruction that is first in the original order in the issue queue; According to the corresponding original order, the reordering number of the first reordering number and the second reordering number corresponding to the previous instruction is determined as a third reordering number; Determine the reordering number of the previous instruction whose corresponding instruction is the first instruction indicated by the third reordering number as the fourth reordering number; The physical register number occupied by the instruction using the same logical register number as the instruction indicated by the fourth reordering number is released.