Scheduling queue, processor, electronic device and instruction scheduling method
By dividing the scheduling queue into static and dynamic subqueues, the mapping relationship is used to reduce data movement, which solves the problem of high dynamic power consumption in the scheduling queue and improves the computing performance of the processor.
Patent Information
- Application Number
- CN202510396760.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The dynamic power consumption caused by data movement in existing scheduling queues is high, affecting processor performance.
The dispatch queue is divided into static subqueues and dynamic subqueues. The instruction data in the static subqueue remains fixed. The dynamic subqueue is responsible for storing and moving the item index of the instruction data in sequence, reducing the amount of data movement by establishing a mapping relationship.
Reduces dynamic power consumption caused by data movement and improves processor computing performance.
Smart Images

Figure CN120335956A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a scheduling queue, a processor, an electronic device, and an instruction scheduling method. Background Art
[0002] To improve the instruction set parallelism, the CPU adopts a pipeline design. Instructions are decoded into micro-operations (uops). During the execution of micro-operations, after register renaming, the micro-operations enter a scheduling queue (SQ). After entering the scheduling queue, the micro-operations wait for the source operands to be ready. If all the source operands of a micro-operation are ready, the micro-operation can be picked out from the scheduling queue and enter the corresponding execution unit for operation; otherwise, if the source operands of the micro-operation are not ready, the micro-operation will continue to wait in the scheduling queue. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a scheduling queue, including a first sub-queue, a second sub-queue, and a queue control module. The first sub-queue includes N first data items, and each first data item is configured to store instruction data. The second sub-queue includes N second data items, and each second data item is configured to store the item index corresponding to the first data item. The queue control module is configured to determine an idle first data item in the first sub-queue for storing the instruction data input to the first sub-queue, and map and store the item index of the non-empty first data item in the first sub-queue to the second data item in the second sub-queue to obtain the corresponding non-empty second data item, and maintain the correspondence between the non-empty first data item in the first sub-queue and the non-empty second data item in the second queue. The second queue is further configured to continuously store the non-empty second data items among the N second data items in a predetermined order, where N is an integer greater than 1.
[0004] For example, in the scheduling queue provided by at least one embodiment of the present disclosure, the queue control module includes: an item index allocation module, configured to, in response to the target instruction data input to the first sub-queue, obtain the item index of the target first data item that is currently in an idle state in the first sub-queue, and allocate the target first data item for storing the target instruction data.
[0005] For example, in the scheduling queue provided by at least one embodiment of the present disclosure, the item index allocation module is further configured to allocate the non-empty target first data item that has been allocated in the first sub-queue to the target second data item in the second sub-queue to obtain the target second data item storing the item index of the target first data item.
[0006] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the queue control module further includes: an item index mapping module configured to map the ready information of the target instruction data corresponding to the target first data item into the target second data item of the second sub-queue, so as to store the ready information of the target instruction data in the target second data item as well.
[0007] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the second sub-queue is further configured to output the item index of the target second data item corresponding to the ready information according to the ready information of the instruction data stored in the target first data item obtained from the first sub-queue, so as to be used to select the instruction data to be issued corresponding to the target second data item in the first sub-queue.
[0008] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the item index mapping module includes: a mapping sub-module configured to output the item index of the target second data item according to the item index of the target first data item stored in the input target second data item.
[0009] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the first sub-queue is further configured to keep the instruction data stored in each non-empty first data item in the first sub-queue fixed and wait for the ready information corresponding to the stored instruction data.
[0010] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the first sub-queue is further configured to, in response to the source operands of the instruction data in the non-empty target first data item in the first sub-queue being ready and being selected to be issued out of the scheduling queue, indicate that the valid bit information of the target first data item is invalid; the queue control module is further configured to, in response to the valid bit information of the target first data item indicating invalid, indicate that the valid bit information of the target second data item corresponding to the target first data item is invalid.
[0011] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the second sub-queue is further configured to, in response to the valid bit information of the target second data item indicating invalid, move at least one other non-empty second data item except the target second data item, so as to continuously save the current non-empty second data items in the second sub-queue in a predetermined order.
[0012] For example, the scheduling queue provided by at least one embodiment of the present disclosure, wherein the second sub-queue is further configured to shift the item index of the first data item stored in the second data item bit by bit according to the new and old order in a displacement manner.
[0013] At least one embodiment of the present disclosure further provides a processor including the scheduling queue provided by any one of the above embodiments.
[0014] For example, the processor provided by at least one embodiment of the present disclosure further includes: an instruction selection module configured to select instruction data to be issued from the first sub-queue in response to receiving at least one item index of the first data item determined according to the item index of at least one second data item corresponding to the ready information provided by the second sub-queue.
[0015] For example, in the processor provided by at least one embodiment of the present disclosure, the instruction selection module is further configured to select, from the item indices of at least one first data item, the instruction data of the first data item corresponding to the oldest ready information for issuing from the first sub-queue.
[0016] For example, the processor provided by at least one embodiment of the present disclosure further includes: a renaming module configured to perform register renaming on the received instruction data and then send it to the first sub-queue for storage in the first sub-queue; and an execution unit module configured to receive the instruction data issued by the first sub-queue for execution.
[0017] At least one embodiment of the present disclosure further provides an electronic device including the processor provided by any of the above embodiments.
[0018] At least one embodiment of the present disclosure further provides an instruction scheduling method, including: determining free first data items among N first data items in a first sub-queue of a scheduling queue for storing instruction data input to the first sub-queue; mapping and storing the item indices of non-empty first data items in the first sub-queue into N second data items in a second sub-queue to obtain corresponding non-empty second data items, where there is a corresponding relationship between non-empty first data items in the first sub-queue and non-empty second data items in the second queue; continuously storing non-empty second data items among the N second data items in the second sub-queue in a predetermined order, where N is an integer greater than 1.
[0019] For example, in the instruction scheduling method provided by at least one embodiment of the present disclosure, determining free first data items among N first data items in a first sub-queue of a scheduling queue for storing instruction data input to the first sub-queue includes: in response to target instruction data input to the first sub-queue, obtaining the item index of a target first data item that is currently in an idle state in the first sub-queue, and allocating the target first data item for storing the target instruction data of the first sub-queue.
[0020] For example, in the instruction scheduling method provided by at least one embodiment of the present disclosure, mapping and storing the item indices of non-empty first data items in the first sub-queue into N second data items in the second sub-queue to obtain corresponding non-empty second data items includes: mapping an allocated non-empty target first data item in the first sub-queue into a target second data item in the second sub-queue to obtain a target second data item storing the item index of the target first data item.
[0021] For example, the instruction scheduling method provided by at least one embodiment of the present disclosure further includes: in response to the source operands of the instruction data in the non-empty target first data item in the first sub-queue being ready and being selected to be emitted from the scheduling queue, indicating that the valid bit information corresponding to the target first data item is invalid, and indicating that the valid bit information corresponding to the target second data item corresponding to the target first data item is invalid.
[0022] For example, the instruction scheduling method provided by at least one embodiment of the present disclosure further includes: in response to the valid bit information indicating invalidity of the target second data item, moving at least one other non-empty second data item except the target second data item to continuously store the current non-empty second data items in the second sub-queue in a predetermined order.
[0023] For example, the instruction scheduling method provided by at least one embodiment of the present disclosure further includes: according to the ready information of the instruction data stored in the target first data item obtained from the first sub-queue, outputting the item index of the target second data item corresponding to the ready information for selecting the instruction data to be emitted corresponding to the target second data item in the first sub-queue. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and are not a limitation to the present disclosure.
[0025] Figure 1 Shows a pipeline schematic diagram of a processor core;
[0026] Figure 2 Shows Figure 1 The structural block diagram of the instruction out-of-order execution engine in;
[0027] Figure 3 Shows a structural schematic diagram of a scheduling queue provided by at least one embodiment of the present disclosure;
[0028] Figure 4 Shows a structural schematic diagram of an example of a queue control module in a scheduling queue provided by at least one embodiment of the present disclosure;
[0029] Figure 5 Shows a structural schematic diagram of a mapping sub-module in an item index mapping module provided by at least one embodiment of the present disclosure;
[0030] Figure 6 Shows a structural schematic diagram of a processor provided by at least one embodiment of the present disclosure;
[0031] Figure 7The figure shows a schematic structural diagram of another processor provided by at least one embodiment of the present disclosure;
[0032] Figure 8 The figure shows a structural example of a scheduling queue of a processor provided by at least one embodiment of the present disclosure
[0033] Figure 9 The figure shows a schematic flowchart of an instruction scheduling method provided by at least one embodiment of the present disclosure;
[0034] Figure 10 The figure shows a schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure; and
[0035] Figure 11 The figure shows a schematic block diagram of an electronic device provided by at least another embodiment of the present disclosure. Detailed implementation manners
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0037] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "include" or "comprise" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute positions of the objects being described change, the relative positional relationships may also change accordingly.
[0038] The processor core (CPU core) of a single-core processor or a multi-core processor improves the instruction execution efficiency through pipeline technology. The pipeline technology divides the complete operation steps of the CPU core into multiple sub-steps and executes these sub-steps in the form of a pipeline to improve the efficiency.
[0039] Figure 1A pipeline schematic diagram of a processor core is shown. The interior of the processor core includes multiple pipeline stages. For example, after various sources of program counters are fed into the pipeline and the next program counter (PC) is selected through a multiplexer (Mux), the instruction corresponding to the program counter needs to go through branch prediction, instruction fetch, decode, dispatch and rename, execute, retire, etc. Waiting queues are set between each pipeline stage as needed, and these queues are usually first-in-first-out (FIFO) queues. For example, after the branch prediction unit, a branch prediction (BP) FIFO queue is set to store the branch prediction results; after the instruction fetch unit, an instruction cache (IC) FIFO is set to cache the fetched instructions; after the instruction decode unit, a decode (DE) FIFO is set to cache the decoded instructions; after the instruction dispatch and rename unit, a retire (RT) FIFO is set to cache the instructions waiting for confirmation of retirement after execution. At the same time, the pipeline of the processor core also includes a scheduling queue to cache the instructions waiting for the instruction execution unit to execute after instruction dispatch and rename. To support a high operating frequency, each pipeline stage may in turn include multiple pipeline levels (clock cycles). Although each pipeline level performs limited operations, in this way each clock can be made the shortest, and the performance of the CPU core is improved by increasing the operating frequency of the CPU. Each pipeline level can also further improve the performance of the processor core by accommodating more instructions (i.e., superscalar technology).
[0040] Within the microarchitecture, the processor core translates each architecture instruction (usually simply referred to as "instruction") into one or more micro-instructions (micro-op, uop). Each micro-instruction only performs limited operations, which can ensure that each pipeline level is very short to increase the operating frequency of the processor core. For example, a memory read instruction (load) can be translated into an address generation micro-instruction and a memory read micro-instruction. The second micro-instruction depends on the result of the first micro-instruction, so the second micro-instruction will only start to execute after the first micro-instruction has finished execution. Micro-instructions contain multiple microarchitecture-related fields to transfer relevant information between pipeline levels.
[0041] Speculative Execution is another technique to improve the performance of a processor. This technique executes the instructions following an instruction before it has completed execution. The branch prediction unit (branch predictor) at the front end of the processor core predicts the jump direction of a branch instruction, prefetches and executes the instructions in that direction; another speculative execution technique is to execute a memory read instruction before the addresses of all previous memory write instructions have been obtained. Speculative execution further improves the instruction-level parallelism, thereby significantly enhancing the performance of the processor core. When a speculative execution error occurs, such as a branch prediction error being detected, or a write instruction before the memory read instruction overwriting the same address, all the instructions in the pipeline after the faulty instruction need to be flushed (or "cleared"), and then the program jumps to the error point to be re-executed to ensure the correctness of program execution. To support speculative execution, the microarchitecture of the processor core also needs to support an architectural register recovery mechanism to ensure that the architectural registers always have the correct values during speculative execution.
[0042] Figure 2 shows Figure 1 a structural block diagram of an example of the instruction out-of-order execution engine in Figure 2 shown. The instruction out-of-order execution engine of a general-purpose CPU mainly includes a register renaming module 101, a schedule queue (SQ) 102, an instruction selection module (Picker) 103, and an execution unit module (Execute) 104, corresponding respectively to Figure 1 instruction renaming, schedule queue, and instruction execution in Figure 3The described scheduling queue 200) is specifically implemented by a reservation station (RS) during the design of the processor to achieve the micro-instruction scheduling function. The reservation station is the core hardware module for modern processors to achieve out-of-order execution and dynamic scheduling. It maximizes the utilization of execution units by tracking the operand status and bypass network in real time. Its design directly affects the instructions per cycle (IPC) and energy efficiency of the processor and is the "invisible engine" of high-performance chip architectures. Unless otherwise specified, the "reservation station" in the following text is referred to as the "scheduling queue" for description.
[0043] In the register renaming stage, the architectural register (or logical register) of the source operand in the micro-instruction is mapped to a physical register number (PRN). The number of mapped PRN numbers is the maximum number of source operands allowed in the micro-instruction. For example, if the maximum number of allowed source operands is 3, but the micro-instruction a + b = c includes two source operands, valid PRN numbers are mapped for a and b, and an additional valid PRN number is newly allocated for c. The control module in the scheduling queue 102 sets the initial value of the ready signal of the source operand according to the PRN numbers mapped by the micro-instruction. In this initial value, the number of bits representing invalidity is the same as the number of source operands included in the micro-instruction.
[0044] After the processor core obtains the source operand according to the source operand address, it broadcasts the information of the obtained source operand to each micro-instruction in the scheduling queue 102. The corresponding control module can adjust the ready signal stored in the corresponding slot in the scheduling queue module according to the received broadcast information.
[0045] As described above, the scheduling queue 102 consists of several slots, and each slot can store several data. For example, the scheduling queue has 32 slots (sequentially slot0, slot1,..., slot31), and each slot can accommodate 100 bits of data. Usually, one micro-instruction occupies one slot. This 100-bit data includes the operation code, source operand address, destination operand address, valid bits, etc. of the micro-instruction, information such as whether the operand is ready (ready signal) and on which channel to execute.
[0046] The valid microinstructions stored in the scheduling queues slot0, slot1, …, slot31 will maintain the new and old order of entry (for example, slot0 is the latest and slot31 is the oldest) and are stored continuously. For example, the microinstruction stored in slot0 is the latest one to enter, and the microinstruction stored in slot31 is the oldest one. If the microinstruction in a certain slot is selected and removed, then the microinstructions in the slots younger than this microinstruction and stored in the slots with smaller serial numbers will move in the direction of larger serial numbers (i.e., older). For example, if the source operands of the microinstructions in slots slot10, slot11, slot12, and slot13 are all ready, and these four microinstructions are selected and removed for execution, then four empty slots will be left. Accordingly, the microinstructions in slots slot0 - slot9 will shift 4 slots bit by bit. Specifically, for example, the microinstruction stored in slot6 will be moved to slot10, the microinstruction stored in slot7 will be moved to slot11, the microinstruction stored in slot8 will be moved to slot12, and the microinstruction stored in slot9 will be moved to slot13; the microinstruction stored in slot2 will be moved to slot6, the microinstruction stored in slot3 will be moved to slot7, the microinstruction stored in slot4 will be moved to slot8, and the microinstruction stored in slot5 will be moved to slot9; the microinstruction stored in slot0 will be moved to slot4, the microinstruction stored in slot1 will be moved to slot5, and so on. Finally, slots slot0, slot1, slot2, and slot3 will become empty slots to receive the 4 updated microinstructions newly entering the scheduling queue in the next clock cycle.
[0047] The above scheduling queue provides a buffer for the microinstructions in the pipeline to stay and wait, and can select out - of - order, improving the instruction parallelism. The design of the scheduling queue that adopts this sequential shift method is called a Shift Queue. Each clock cycle, new microinstructions enter the scheduling queue, and the microinstructions in other slots will move in the direction of older. However, if the number of slots is large and the number of data bits in each slot is large (taking the 32 - slot, 100 - bit SQ as an example), then the amount of data that needs to be moved each clock cycle is huge, and the dynamic power consumption of the system will increase significantly, reducing the performance of the processor.
[0048] At least one embodiment of the present disclosure provides a scheduling queue, a processor, an electronic device, and an instruction scheduling method.
[0049] The scheduling queue according to an embodiment of the present disclosure includes a first sub-queue, a second sub-queue, and a queue control module. Wherein, the first sub-queue includes N first data items, and each first data item is configured to store instruction data; the second sub-queue includes N second data items, and each second data item is configured to store the item index of the corresponding first data item; the queue control module is configured to determine an idle first data item in the first sub-queue for storing the instruction data input to the first sub-queue, and map and store the item index of the non-empty first data item in the first sub-queue into the second data item of the second sub-queue of the scheduling queue to obtain the corresponding non-empty second data item, and maintain the corresponding relationship between the non-empty first data item in the first sub-queue and the non-empty second data item in the second queue; wherein, the second queue is further configured to continuously store the non-empty second data items among the N second data items in a predetermined order, and N is an integer greater than 1.
[0050] At least one embodiment of the present disclosure further provides an instruction scheduling method corresponding to the above scheduling queue. The instruction scheduling method includes: determining an idle first data item among the N first data items in the first sub-queue of the scheduling queue for storing the instruction data input to the first sub-queue; mapping and storing the item index of the non-empty first data item in the first sub-queue into the N second data items of the second sub-queue of the scheduling queue to obtain the corresponding non-empty second data item, wherein, there is a corresponding relationship between the non-empty first data item in the first sub-queue and the non-empty second data item in the second queue; continuously storing the non-empty second data items among the N second data items in the second sub-queue in a predetermined order, wherein, N is an integer greater than 1.
[0051] In the scheduling queue and the scheduling method provided by at least one embodiment of the present disclosure, the scheduling queue is divided into a first sub-queue and a second sub-queue, and the data items of the two are established with a corresponding relationship through mapping. The instruction data entering the first sub-queue maintains a fixed storage position to wait for the ready information of the operand, and the second sub-queue is responsible for orderly moving the storage position information of the instruction data in the first sub-queue in a predetermined order (such as the order of new and old). Thus, the data items of the first sub-queue storing the instruction data do not move, and when the data items of the second sub-queue storing the storage position information are moved, the amount of data to be moved is greatly reduced, effectively reducing the dynamic power consumption generated by data movement in the entire scheduling queue and improving the computing performance of the processor.
[0052] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0053] Figure 3The figure shows a schematic structural diagram of a scheduling queue provided by at least one embodiment of the present disclosure. This scheduling queue is, for example, used in various applicable queues in the pipeline of a processor core. For example, in at least one example, it is used to cache instructions or microinstructions that have undergone register renaming, and then send these instructions or microinstructions to the corresponding execution unit after meeting the conditions. Therefore, in this example, the scheduling queue is located between the register renaming unit and the execution unit in the pipeline.
[0054] As Figure 3 shown, the scheduling queue 200 includes a first sub-queue 210, a second sub-queue 220, and a queue control module 230. The first sub-queue includes N first data items 2101 to 210N, and the second sub-queue includes N second data items 2201 to 220N. For example, N can be any value such as 16, 32, or 64, and this value can be determined based on the buffering capacity required by the scheduling queue 200. Embodiments of the present disclosure do not limit this. It should be noted that the number of first data items in the first sub-queue is the same as the number of second data items in the second sub-queue to ensure that each first data item can correspond to a second data item.
[0055] Since the instruction data stored in the first sub-queue 210 (such as the data corresponding to instructions or microinstructions) remains in a fixed position without being moved, it can be called a static queue (SQ, Static Queue). Similarly, since the content stored in the second sub-queue 220 needs to be moved while maintaining the order, it can be called a dynamic queue (DQ, Dynamic Queue). The first data item and the second data item (entry) are the slots described above. For the convenience of description, hereinafter, the first data item will be called a slot, and the second data item will be called an entry. It should be noted that here, slot and entry are two ways of expressing the same concept for distinction. Unless otherwise specified, hereinafter, the first sub-queue will also be called SQ, and the second sub-queue will be called DQ.
[0056] Each first data item is configured to store the instruction data entering the scheduling queue 200. This instruction data is, for example, an instruction or a microinstruction. These instruction data wait in the first sub-queue and are sent to the execution unit for execution after they are ready; each second data item is configured to store the item index of the corresponding first data item, such as the label or serial number of the corresponding first data item in the first sub-queue. The queue control module 230 establishes a corresponding relationship between the data items in the first sub-queue 210 and the second sub-queue 220 to reduce the overall dynamic power consumption of the scheduling queue 200.
[0057] Figure 4 The figure shows a schematic structural diagram of an example of a queue control module in a scheduling queue provided by at least one embodiment of the present disclosure. As Figure 4As shown, the queue control module 230 includes an item index allocation module 2301 and an item index mapping module 2302.
[0058] For example, in a possible implementation, the above-mentioned item index allocation module 2301 is configured to, in response to target instruction data input to the first sub-queue, obtain the item index of the target first data item that is currently in an idle state in the first sub-queue, and allocate the target first data item for storing the target instruction data. Here, "target instruction data" refers to the instruction data that is currently the object of description; similarly, "target first data item" refers to the first data item that is currently the object of description.
[0059] For example, if there are 4 pieces of target instruction data (such as 4 micro-instructions) input to the scheduling queue by the processor pipeline at the current beat, and these instruction data will be stored in SQ, the item index allocation module 2301 will find 4 empty slots in the first sub-queue for the 4 pieces of instruction data newly entered into SQ, and allocate the slot numbers (such as tags) corresponding to the empty slots to the 4 pieces of instructions for storing these four pieces of instruction data. It should be noted that 4 is only exemplary, and it can be 6, 8, etc., depending on the number of pipelines, and the embodiments of the present disclosure do not limit this.
[0060] For example, the method by which the item index allocation module 2301 searches for empty slots may include the following steps: First, start searching from the top (i.e., the queue entry) of SQ until an empty slot is found, and place the first piece of instruction data among the above 4 pieces of target instruction data in it. Then, start searching from the bottom (i.e., the queue exit) of SQ until a second empty slot is found, and place the second piece of instruction data among the above 4 pieces of target instruction data in it. Then, start searching from the top of SQ again until a third empty slot is found, and place the third piece of instruction data in it. Finally, start searching from the bottom of SQ again until a fourth empty slot is found, and place the fourth piece of instruction data in it.
[0061] Again, for example, it is also possible to sequentially find four empty slots by traversing in a single direction from the top to the bottom of SQ according to actual needs, and sequentially place the above 4 pieces of target instruction data. The specific way of finding empty slots depends on the actual situation, and the embodiments of the present disclosure do not limit this.
[0062] Exemplarily, the SQ includes 32 slots (corresponding to slot0 to slot31, where slot0 is at the top of the queue and slot31 is at the bottom of the queue). Each slot can store 100 bits of instruction data, and the stored instruction data includes information such as the operation code, operand address, and valid bit of the instruction. For example, in the item index allocation module 2301, according to the above-mentioned way of traversing alternately from the top and the bottom to find an empty slot, the first empty slot is slot0, the second empty slot is slot7, the third empty slot is slot3, and the fourth empty slot is slot5. Since the input of the instruction data received by the SQ has a sequential order, the first instruction data received is stored in slot0, the second instruction data is stored in slot7, the third instruction data is stored in slot3, and the last instruction data is stored in slot5.
[0063] For example, in a possible implementation manner, the above item index allocation module 2301 is further configured to map the non-empty target first data items already allocated in the first sub-queue to the target second data items in the second sub-queue to obtain the target second data items storing the item indices of the target first data items.
[0064] Exemplarily, the DQ also includes 32 slots (i.e., entry0 to entry31, where entry0 is at the top of the queue (i.e., the queue entry) and entry31 is at the bottom of the queue (i.e., the queue exit)). The item index mapping module 2302 sequentially allocates the slot numbers corresponding to the non-empty slots already allocated in the above SQ (i.e., the item index "0" in slot0, the item index "7" in slot7, the item index "3" in slot3, and the item index "5" in slot5) to entry3, entry2, entry1, and entry0 in the DQ for storage. For example, the DQ stores the non-empty slots in the 32 slots continuously in a predetermined order from the queue exit to the queue entry, that is, "0" in slot0 is stored in entry3, and so on for the remaining cases.
[0065] It should be noted that the slot numbers in the SQ are stored in the DQ in binary form. Here, since there are only 32 slots in the SQ, each slot in the DQ can be stored as a 5-bit binary number on the premise of ensuring brevity. Correspondingly, for example, "7" in slot7 of the SQ is stored as "00111" in entry2 of the DQ, and "5" in slot5 is stored as "00101" in entry0, etc. If there are 64 slots in the SQ, then correspondingly, each slot in the DQ is stored as at least a 6-bit binary number. The number of binary digits of each slot in the DQ is adaptively determined based on the number of slots in the SQ, and the present disclosure does not limit this.
[0066] When there is data movement in the above storage method of the DQ from the queue exit to the queue entrance (for example, when there is instruction data that meets the emission condition and the SQ is selected), it can ensure sufficient space for ordered movement and space for storing the slot numbers corresponding to the newly allocated slots from the SQ.
[0067] In another example, the DQ can continuously store non-empty slots starting from the middle of the queue (for example, entry16). At this time, the DQ is divided into two relatively independent sub-queues for respective management. The upper half, entry0 to entry15, is one sub-queue, and the lower half, entry16 to entry31, is another sub-queue. Among them, for the two sub-queues, entry0 and entry16 are the latest stored data respectively. The specific situation is based on the actual design, and the embodiments of the present disclosure do not limit this.
[0068] For example, in a possible implementation manner, the above queue control module 230 further includes an item index mapping module 2302, and the item index mapping module 2302 is configured to map the ready information of the target instruction data corresponding to the target first data item to the target second data item in the second sub-queue, so as to store the ready information of the target instruction data in the target second data item.
[0069] For example, the ready information (ready signal) of the target instruction data refers to the ready signal of the source operands of the target instruction data. For example, when the ready signal is 1, it means that all source operands are ready, and when it is 0, it means that at least one source operand is not ready. For example, the ready signal is generated by the control module of the scheduling queue (or the scheduling module of the SQ). For example, if the instruction includes at most 3 source operands, the generated ready signal means that the ready signals corresponding to the 3 source operands have all been obtained. The ready information of the operands of the instruction data stored in slot0, slot7, slot3, and slot5 of the above SQ is respectively mapped to entry3, entry2, entry1, and entry0 of the DQ for storage.
[0070] For example, in a possible implementation manner, the second sub-queue is further configured to output the item index of the target second data item corresponding to the ready information according to the ready information of the instruction data stored in the target first data item obtained from the first sub-queue, so as to select the instruction data to be emitted corresponding to the target second data item in the first sub-queue.
[0071] For example, the entry3, entry2, entry1, and entry0 of the above DQ respectively store the ready information indicating whether the instruction data operands of slot0, slot7, slot3, and slot5 of the SQ are ready. When the ready information of slot7 obtained from the SQ is 1, the item index of the data item in the DQ corresponding to the obtained ready information of slot7 (i.e., the entry number of the DQ) is obtained, which is 2. The ready information of entry2 is stored as 1, indicating that the instruction corresponding to entry2 is ready. Thus, after the instruction stored in slot7 is issued from the SQ, the data item of slot7 is cleared, and the corresponding data item in entry2 is also cleared. After that, the DQ will perform a shift operation.
[0072] The scheduling queue will issue the currently all-ready instructions in the order from the oldest to the newest. Therefore, the order of issue is determined according to the item indexes of these ready instructions in the DQ.
[0073] For example, in a possible implementation manner, the above item index mapping module 2302 includes a mapping sub-module, which is configured to output the item index of the target second data item according to the item index of the target first data item stored in the input target second data item.
[0074] Figure 5 The structural schematic diagram of a mapping sub-module in a processor provided by at least one embodiment of the present disclosure is shown.
[0075] As Figure 5 shown, the input end of the mapping sub-module 300 is the slot number (tag) of the non-empty slot in the corresponding SQ stored in the non-empty slot in the above DQ. For example, if the SQ has 32 slots, the input end can be expressed as Dslot0_Tag[4:0], …, Dslot31_Tag[4:0], where [4:0] represents that the tag value is represented by a 5-bit binary number (for example, if entry31 has a corresponding relationship with slot31, that is, the “31” in slot31 is stored in entry31, then the value of [4:0] in Dslot31_Tag[4:0] is 11111).
[0076] For example, if the SQ has 64 slots (slot0 to slot63), adaptively, a 6-bit binary number is required to represent the tag value. Furthermore, the input end can be expressed as Dslot0_Tag[5:0], …, Dslot63_Tag[5:0]. The output end of the mapping sub-module 300 is the slot number (entry number) corresponding to the non-empty slot stored in the DQ.
[0077] For example, since the number of slots in DQ is the same as that in SQ, the output terminals can correspondingly be expressed as Sslot0_Dentryn[31:0], …, Sslot0_Dentryn[31:0], where [31:0] is a one-hot encoding structure, so there are a total of 32 bits.
[0078] For example, as can be seen from the foregoing example, slot0 in SQ corresponds to entry3 in DQ, that is, the tag value stored in entry3 is 00000. When Dslot0_Tag[4:0] is input to the mapping sub-module 300, the corresponding output Sslot0_Dentryn[31:0] represents the item index number of slot0 in SQ corresponding to entry3 in DQ, that is, 3. Then the value of Sslot0_Dentryn[31:0] is 00000000000000000000000000001000, that is, the value of the 4th bit (bit3) from right to left is set to 1, and the remaining bits are 0. For another example, Sslot7_Dentryn[31:0] represents that the item index of entry2 in DQ corresponding to slot7 in SQ is 2. Then the value of Sslot7_Dentryn[31:0] is 00000000000000000000000000000100, that is, the value of the 3rd bit (bit2) from right to left is set to 1, and the remaining bits are 0.
[0079] It should be noted that the output terminals can also be represented by binary numbers correspondingly. That is, the output of the above mapping sub-module 300 is expressed as Sslotn_Dentryn[4:0]. Correspondingly, the value of Sslot0_Dentryn[4:0] is 00011; the value of Sslot7_Dentryn[4:0] is 00010.
[0080] For example, in a possible implementation manner, the first sub-queue is further configured to keep the instruction data stored in each non-empty first data item in the first sub-queue fixed and wait for the ready information corresponding to the instruction data to be stored.
[0081] For example, the storage locations of the 4 instruction data input to SQ for storage in the above embodiment are kept fixed. That is, the first instruction is always stored in slot0, the second instruction is always stored in slot7, the third instruction is always stored in slot3, and the last instruction is always stored in slot5. These 4 instructions will wait for whether their corresponding source operands are ready (ready signal). Specifically, for these 4 instructions, the processor will continuously read the relevant source operands from the register file (file), the corresponding signal line or the memory at a certain frequency according to the source operand addresses in the instruction data.
[0082] For example, in one possible implementation, the first sub-queue is further configured to, in response to the source operands of the instruction data in the target first data item in the first sub-queue being ready while the first sub-queue is non-empty, be selected to be dispatched from the scheduling queue, and the valid bit information of the target first data item is indicated as invalid.
[0083] For example, the queue control module is further configured to, in response to the valid bit information of the target first data item being indicated as invalid, indicate the valid bit information of the target second data item corresponding to the target first data item as invalid.
[0084] For example, the source operands of the instruction stored in slot0 in the SQ of the above embodiment are ready, that is, all relevant source operands of the instruction are successfully read from the register file or memory and the instruction is selected to be dispatched from the SQ. After that, the valid bit of slot0 storing the instruction data is set to invalid (i.e., changed from the binary number 1 to 0), whereby the instruction data in slot0 is cleared (or equivalent to being cleared), so that new input instruction data can be received.
[0085] For example, the data capacity stored in each slot in the SQ can be 100 bits, and the 100-bit data includes information such as the opcode of the instruction, the source operand address, the destination operand address, the valid bit of the slot, whether the operand is ready, and the pipeline execution sequence.
[0086] It should be noted that, on the one hand, the arrangement manner of the data stored in the slot can be arbitrary or random during operation, and the embodiments of the present disclosure do not limit this; on the other hand, the data capacity stored in the slot can be determined based on the performance of the actually selected hardware, and the embodiments of the present disclosure also do not limit this.
[0087] Correspondingly, after the valid bit of the above slot0 is set to invalid, the queue control module 230 sets the valid bit of entry3 storing the slot number of slot0 in the DQ to invalid, that is, changed from the binary number 1 to 0, and then the slot number stored in entry3 of the DQ is cleared to prepare for receiving new allocated slot number information. For slot7, slot3, and slot5 in the above embodiment, their execution logics are exactly the same as that of slot0, and will not be elaborated here.
[0088] For example, in one possible implementation, the second sub-queue is further configured to, in response to the valid bit information of the target second data item being indicated as invalid, move at least one other non-empty second data item except the target second data item to continuously store the current non-empty second data items in the second sub-queue in a predetermined order.
[0089] For example, as described above, the valid bit of entry3 in DQ is set to 0 and cleared. If a new instruction data is then input into SQ and stored in an empty slot (such as the cleared slot0), to ensure the new and old order of the remaining slots in DQ and allocate storage space for the slot number information of the newly input slot0, the contents stored in the remaining entry2 to entry0 need to be moved one slot position towards the older direction. It should be noted that the content stored in entry0 is the latest, and the content stored in entry31 is the oldest. That is, the content in entry2 is moved to entry3, the content in entry1 is moved to entry2, and the content in entry0 is moved to entry1, and then entry0 is cleared to store the slot number information of the newly allocated slot0.
[0090] For example, in a possible implementation manner, the second sub-queue is further configured to shift the item index of the first data item stored in the second data item bit by bit according to the new and old order in a displacement manner.
[0091] For example, the slot number information of the non-empty slots in SQ stored in DQ in order will continuously move towards the older direction along with the slot number information of the newly allocated non-empty slots.
[0092] Exemplarily, on the basis that entry3 to entry0 in the above DQ already store tag values, if SQ receives 4 newly input instruction data and stores them in new empty slots and becomes non-empty slots, the above item index mapping module 2301 will map the tag values corresponding to the 4 new non-empty slots to the empty slots (empty entry) in DQ for storage. Therefore, the tag values stored in entry3 to entry0 are moved in order towards the older direction. Since entry31 is the oldest, the content in entry3 is moved to entry7, the content in entry2 is moved to entry6, the content in entry1 is moved to entry5, and the content in entry0 is moved to entry4. In this way, entry3 to entry0 are cleared to receive the newly allocated tag values in the new and old order. It should be noted that when moving the content in the entry, the ready information of the instruction data corresponding to the stored tag value also needs to be moved synchronously.
[0093] The scheduling queue provided in the above embodiments of the present disclosure includes two sub-queues, SQ and DQ. The instruction data entering SQ remains fixed, and the tag values of the instruction data corresponding to the storage in DQ move relatively few bits of data on the basis of maintaining the order, greatly reducing the dynamic power consumption generated by moving the instruction data as a whole by establishing a mapping relationship between SQ and DQ, and significantly improving the computing performance of the processor.
[0094] At least one embodiment of the present disclosure further provides a processor. For example, the processor can be a single-core processor or a multi-core processor, and can also be a single-threaded processor or a simultaneous multi-threading (SMT) processor. Figure 6 FIG. shows a schematic structural diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 6 shown, the processor 400 includes the scheduling queue 200 provided by any of the above embodiments.
[0095] Figure 7 FIG. shows a schematic structural diagram of another processor provided by at least one embodiment of the present disclosure. As Figure 7 shown, on the basis of Figure 6 the processor 400 further includes an instruction selection module 410, a renaming module 420, and an execution unit module 430.
[0096] For example, in a possible implementation manner, the instruction selection module 410 is configured to, in response to receiving, select the instruction data to be issued in the first sub-queue according to the item indexes of at least one first data item determined by the item indexes of at least one second data item corresponding to the ready information provided by the second sub-queue.
[0097] For example, in the SQ of the above embodiment, the source operands of the instruction data stored in slot0 and slot7 are ready. Through the item indexes of slot0 and slot7, it is determined that the item indexes of them are respectively stored in entry3 and entry2 of the current DQ. The ready information of slot0 and slot7 can be mapped and stored in entry3 and entry2 of the DQ. Since entry3 and entry2 in the DQ are already ready, then according to the new and old order in the DQ, the older entry3 can be selected for processing first, and entry2 for processing later. The instruction selection module 410 selects the corresponding instruction data to be issued from the SQ according to the tag values (0 and 7 respectively) of slot0 and slot7 stored in entry3 and entry2 corresponding to the ready information provided by the DQ, that is, selects slot0 first and then slot7. However, the instructions stored in slot0 and slot7 can be issued in the same clock cycle. It should be noted that the number of instructions that can be issued simultaneously is based on the number of pipelines. For example, if there are 4 pipelines, then 4 corresponding instructions can be selected and issued at one time, that is, 1 pipeline selects and issues 1 instruction.
[0098] For example, in a possible implementation manner, the instruction selection module is further configured to select the instruction data of the first data item corresponding to the oldest ready information among the item indexes of at least one first data item for issuing from the first sub-queue.
[0099] For example, on the basis that the tag values of slot0, slot7, slot3, and slot5 already correspond to entry3 to entry0 of the above DQ, the SQ newly inputs 4 new instruction data and stores them in new empty slots in sequence (for example, slot1, slot6, slot2, and slot4). The DQ receives a request to store the tag values corresponding to the new instruction data in the SQ, and shifts the tag values in entry3 to entry0 4 units in the older direction bit by bit, that is, the tag value stored in entry3 is moved to entry7, the tag value stored in entry2 is moved to entry6, the tag value stored in entry1 is moved to entry5, and the tag value stored in entry0 is moved to entry4, while the data in entry3 to entry0 is cleared to receive the newly stored tag values. Furthermore, the tag values of slot1, slot6, slot2, and slot4 are stored in entry3 to entry0 in sequence. If all the source operands of the instruction data in slot0 to slot7 in the SQ are ready at this time, that is, entry0 to entry7 of the DQ will provide corresponding ready information.
[0100] The processor of the embodiments of the present disclosure may be a single-threaded processor or a simultaneous multi-threading (SMT) processor. For example, for the embodiments of the SMT processor, the instruction data in slot0 to slot31 in the above SQ may come from different corresponding threads, and the thread information (such as the thread number) to which the instruction belongs may also be included in this instruction data.
[0101] In the embodiments of the present disclosure, the instruction selection module includes multiple channels, and the multiple channels (pipes) of the instruction selection module correspond to multiple execution units one by one. The instruction selection module will select the corresponding channels to process the instruction data of the corresponding execution units for execution.
[0102] For example, in a possible implementation manner, the processor provided in the above embodiments further includes a renaming module and an execution unit module. The renaming module is configured to perform register renaming on the received instruction data and then send it to the first sub-queue for storage in the first sub-queue;
[0103] The execution unit module is configured to receive the instruction data emitted by the first sub-queue for execution. The execution unit module includes multiple different types of execution units, such as arithmetic logic units, multiplication units, memory access units, etc.
[0104] For example, the execution unit module 430 receives the instruction data (such as operation codes, source operands, etc.) emitted by the SQ for execution.
[0105] For example, the instruction selection module includes channels 0 to 3; the execution unit module may include execution units 0 to 3 corresponding to channels 0 to 3 (e.g., in a one-to-one correspondence). The instruction selection module will select the oldest ready instruction data corresponding to channel 0 (execution unit 0) from slot0 to slot31, and issue it from the SQ to the execution unit 0 through channel 0. The execution logic of the instruction selection module for other channels is similar to that of channel 0, which will not be elaborated here.
[0106] For program code, programmers or compilers use a set of limited logical registers. When multiple instructions attempt to read and write the same logical register, data dependencies may occur, which limits the parallel execution of instructions. However, in many cases, these dependencies are actually an illusion caused by the same logical register names used by different instructions, rather than true data dependencies. The above-mentioned renaming module 420 is used to eliminate this illusion. It dynamically allocates physical registers to replace fixed architectural registers, thereby avoiding false data dependencies (i.e., pseudo-dependencies), enabling more instructions to be executed in parallel.
[0107] The processor provided by at least one embodiment of the present disclosure includes the scheduling queue provided by any of the above embodiments. The scheduling queue has two sub-queues, SQ and DQ, which establish a mapping relationship. The instruction data stored in the SQ remains fixed, and the DQ is responsible for storing and orderly moving the item index values of the instruction data stored in the SQ. The number of bits to be moved in the DQ is greatly reduced, effectively reducing the dynamic power consumption of the entire scheduling queue due to data movement, and thus significantly improving the computing performance of the processor.
[0108] The processor core of the embodiment of the present disclosure may be based on X86 microarchitecture, ARM microarchitecture, RISC-V microarchitecture, MIPS microarchitecture, etc., and the embodiments of the present disclosure do not limit this.
[0109] Figure 8 FIG. shows a structural example of the scheduling queue of a processor provided by at least one embodiment of the present disclosure.
[0110] As Figure 8 shown, the scheduling queue of the processor includes a static sub-queue Static Q, a dynamic sub-queue DynamicQ, a queue control module, and an instruction selection module. The queue control module further includes an item index allocation module and an item index mapping module; the instruction selection module includes channels p0, p1, p2, and p3.
[0111] For ease of explanation below, the static sub-queue Static Q is referred to as SQ, which includes items s0 to s47 (only some are shown in the figure); the dynamic sub-queue Dynamic Q is called DQ, which includes items d0 to d47 (only some are shown in the figure).
[0112] It can be seen that, for example, four instruction data respectively enter the SQ after being register-renamed by a rename module (not shown) in the processor pipeline through their respective pipelines. The item index allocation module searches for and obtains empty slots s0, s7, s3, and s5 in the SQ for the four instruction data. Here, the way to search for empty slots is to first find the first empty slot s0 from the top of the SQ (the position corresponds to Top1), then find the second empty slot s7 from the bottom of the SQ (the position corresponds to Bot1), then find the third empty slot s3 from the top of the SQ (the position corresponds to Top2), and finally find the fourth empty slot s5 from the bottom of the SQ (the position corresponds to Bot2). The four instruction data are stored in s0, s7, s3, and s5 respectively in the order they enter the SQ. At the same time, the slot indexes (tag numbers) of these four empty slots are also stored in the DQ in the old-new order (i.e., the order found by the item index allocation module), that is, the slot number of s0 is put into d3, the slot number of s7 is put into d2, the slot number of s3 is put into d1, and the slot number of s5 is put into d0.
[0113] It should be noted that d0 in the DQ is the queue entry for storing the latest allocated slot number from the SQ. The instructions become older in the direction from top to bottom towards the queue exit. EntryValid[47:0] indicates that there are 48 slots (s0 to s47) in the SQ for storing instruction data and their respective validity data; since the slot capacities of the DQ and the SQ are the same, there are also 48 slots (d0 to d47) in the DQ for storing the slot numbers corresponding to the instruction data, that is, d47 in the DQ is the storage position for the slot number corresponding to the oldest instruction data.
[0114] It should be noted that the slot numbers stored in the DQ are represented in the form of binary numbers. Since there are a total of 48 slots, at least 6-bit binary numbers [5:0] are required for representation on the basis of ensuring brevity; for the situation shown in the figure, the data stored in d0 is 000101, the data stored in d1 is 000011, the data stored in d2 is 000111, and the data stored in d4 is 000000.
[0115] On the one hand, the item index mapping module maps the signals in the SQ (such as item index, ready signal, etc.) to the DQ corresponding to the slot numbers. The DQ determines the operations to be performed according to the mapped signals. For example, the item index of the newly allocated item in the mapped SQ is filled into the empty item (slot) in the DQ, or the ready signal of the item in the mapped SQ is filled into the corresponding item in the DQ, etc.
[0116] For example, in the case where the SMT processor has 4 threads, in the first cycle of the SQ, 4 instruction data of thread 0 are received, in the second cycle, 4 instruction data of thread 1 are received, in the third cycle, 4 instruction data of thread 2 are received, and in the fourth cycle, 4 instruction data of thread 3 are received. The item index mapping module maps the signals of the newly input 4 instruction data in the SQ to the DQ. In order to clear the latest slot to store the newly mapped slot number, based on the shift signal generated by the item index mapping module, the data in d3 is shifted down to d7, the data in d2 is shifted down to d6, the data in d1 is shifted down to d5, and the data in d0 is shifted down to d4, where the old and new orders of the original slot numbers remain unchanged.
[0117] On the other hand, the ready information of the instruction data in the SQ is also mapped and stored in the corresponding slot of the DQ. As Figure 8 shown, Ready0_FSP, Ready1_FSP, Ready2_FSP, and Ready3_FSP represent the ready information of the instruction data from four pipelines. Since the SQ has 48 slots, each of Ready0_FSP, Ready1_FSP, Ready2_FSP, and Ready3_FSP has 48 binary bits to represent the valid information of the ready information of each instruction data. If it is ready, the corresponding binary bit is set to 1, and if it is invalid, it is set to 0. It should be noted that since the slot numbers stored in the DQ are stored in order, the 48-bit ready signals of each pipeline, Dslot_Ready0[47:0], Dslot_Ready1[47:0], Dslot_Ready2[47:0], and Dslot_Ready3[47:0], mapped from the SQ are also stored in order.
[0118] The mapping sub-module of the queue control module (not shown in the figure, for reference, see Figure 5 ) outputs the slot number information of each corresponding DQ slot according to the slot number stored in each slot of the DQ, that is, Dslot_Tag0 - 47. The instruction selection modules p0 to p3 select the slot number corresponding to the oldest ready0 - 3 signal from Dslot_Ready0 - 3[47:0] according to the ready information (i.e., Dslot_Ready0 - 3) mapped to the DQ and the slot number output by the DQ, and send it to the SQ for selection and emission from the SQ.
[0119] For example, in channel p0, according to Dslot_Ready0[47:0] and Dslot_Tag0-47, the oldest ready0
[29] is found from Dslot_Ready0[47:0]. Then, the slot number stored in d29 in DQ is the slot number corresponding to ready0
[29] . SQ selects and emits the corresponding instruction data according to this slot number to be provided to the corresponding execution unit through channel p0 for execution. For the remaining 3 channels, the execution logic is similar to that of channel 0 and will not be elaborated here.
[0120] After receiving the instruction data emitted by SQ, the execution unit of the processor uses the obtained data of the original operands for execution.
[0121] Figure 8 The illustrated processor example divides the scheduling queue into two sub-queues, SQ and DQ. A single data movement in DQ only requires a 6-bit binary number, greatly avoiding the dynamic power consumption caused by large data volume movement, and thus significantly improving the computing performance of the processor.
[0122] Figure 9 The flowchart of an instruction scheduling method provided by at least one embodiment of the present disclosure is shown. As Figure 9 shown, the instruction scheduling method includes steps S910 to S930.
[0123] Step S910: Determine the idle first data items among the N first data items in the first sub-queue of the scheduling queue to store the instruction data input to the first sub-queue.
[0124] Step S920: Map and store the item indexes of the non-empty first data items in the first sub-queue into the N second data items in the second sub-queue of the scheduling queue to obtain the corresponding non-empty second data items, where there is a corresponding relationship between the non-empty first data items in the first sub-queue and the non-empty second data items in the second queue.
[0125] Step S930: Continuously save the non-empty second data items among the N second data items in the second sub-queue in a predetermined order, where N is an integer greater than 1.
[0126] For example, the above instruction scheduling method corresponds to the scheduling queue as Figure 3 shown, and the corresponding operation steps can correspond to Figure 4 、 Figure 5 description.
[0127] For example, in a possible implementation, determining the free first data items among the N first data items in the first sub-queue of the scheduling queue to store the instruction data input to the first sub-queue includes: in response to the target instruction data input to the first sub-queue, obtaining the item index of the target first data item that is currently in an idle state in the first sub-queue, and allocating the target first data item to store the target instruction data of the first sub-queue. For example, this step may be executed by the item index allocation module 2301 in the above Figure 4 , and the specific operation method will not be elaborated here.
[0128] For example, in a possible implementation, mapping and storing the item indices of the non-empty first data items in the first sub-queue to the N second data items in the second sub-queue to obtain the corresponding non-empty second data items includes: mapping the non-empty target first data item that has been allocated in the first sub-queue to the target second data item in the second sub-queue to obtain the target second data item storing the item index of the target first data item. For example, this step may also be executed by the item index allocation module 2301 in the above Figure 4 , and the specific operation method will not be elaborated here.
[0129] For example, in a possible implementation, the above instruction scheduling method further includes: in response to the source operands of the instruction data in the non-empty target first data item in the first sub-queue being ready and being selected to be emitted from the scheduling queue, indicating that the valid bit information of the corresponding target first data item is invalid, and indicating that the valid bit information of the target second data item corresponding to the target first data item is invalid. For example, the operation step of indicating that the valid bit information of the corresponding target first data item is invalid can be performed by Figure 2 the first sub-queue 210 in the above; the operation step of indicating that the valid bit information of the target second data item corresponding to the target first data item is invalid can be performed by Figure 2 the queue control module 230 in the above, and the specific operation method will not be elaborated here.
[0130] For example, in a possible implementation, the above instruction scheduling method further includes: in response to the valid bit information of the target second data item indicating invalid, moving at least one other non-empty second data item except the target second data item to continuously store the current non-empty second data items in the second sub-queue in a predetermined order. For example, this step may be executed by Figure 2 the second sub-queue 220 in the above, and the specific operation method will not be elaborated here.
[0131] For example, in one possible implementation, the above instruction scheduling method further includes: outputting an item index of a target second data item corresponding to the readiness information according to the readiness information of the instruction data stored in the target first data item obtained from the first sub-queue, so as to be used to select the instruction data to be issued corresponding to the target second data item in the first sub-queue. For example, this step may be executed by Figure 7 the instruction selection module 410 in
[0132] In the instruction scheduling method provided by at least one embodiment of the present disclosure, the scheduling queue is divided into a first sub-queue and a second sub-queue, and the data items of the two are established in a corresponding relationship through mapping. The instruction data entering the first sub-queue maintains a fixed storage position to wait for the readiness information of the operand, and the second sub-queue is responsible for orderly moving the storage position information of the instruction data in the first sub-queue. Thus, the data items of the first sub-queue storing the instruction data do not move, and when the data items of the second sub-queue storing the storage position information are moved, the amount of data to be moved is greatly reduced, effectively reducing the dynamic power consumption generated by data movement in the entire scheduling queue and improving the computing performance of the processor.
[0133] At least one embodiment of the present disclosure further provides an electronic device, Figure 10 showing a schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure, as Figure 10 shown, the electronic device 500 includes the processor 400 provided by any of the above embodiments.
[0134] The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0135] At least another embodiment of the present disclosure further provides an electronic device, Figure 11 showing a schematic block diagram of an electronic device provided by at least another embodiment of the present disclosure.
[0136] Figure 11 The electronic device 1000 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0137] For example, as Figure 11As shown, in some examples, the electronic device 1000 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which may include the processor of any of the above embodiments and can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the computer system are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through the bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0138] For example, the following components may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; a communication device 1009 which may also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that the computer program read from it can be installed into the storage device 1008 as needed. Although Figure 10 an electronic device 1000 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0139] For example, the electronic device 1000 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 1009 may communicate with the network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0140] For example, the electronic device 1000 may be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, a server, etc., or may be any combination of a data processing device and hardware. The embodiments of the present disclosure are not limited thereto.
[0141] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made based on the embodiments of the present disclosure, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure fall within the scope of protection required by the present disclosure.
[0142] For the present disclosure, in addition to the above exemplary description, the following points need to be noted:
[0143] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.
[0144] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of the layer or region is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.
[0145] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.
[0146] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the said claims.
Claims
1. A scheduling queue, comprising a first sub-queue, a second sub-queue and a queue control module, wherein, the first sub-queue includes N first data items, and each of the first data items is configured to store instruction data; the second sub-queue includes N second data items, and each of the second data items is configured to store the item index of the corresponding first data item; the queue control module is configured to determine an idle first data item in the first sub-queue for storing the instruction data input to the first sub-queue, and map and store the item index of the non-empty first data item in the first sub-queue to the second data item in the second sub-queue to obtain a corresponding non-empty second data item, and maintain the correspondence between the non-empty first data item in the first sub-queue and the non-empty second data item in the second queue; wherein, the second queue is further configured to continuously store the non-empty second data items among the N second data items in a predetermined order, and N is an integer greater than 1.
2. The scheduling queue according to claim 1, wherein, The queue control module includes: an item index allocation module, configured to, in response to target instruction data input to the first sub-queue, obtain the item index of the currently idle target first data item in the first sub-queue, and allocate the target first data item for storing the target instruction data.
3. The scheduling queue according to claim 2, wherein The item index allocation module is further configured to allocate the non-empty target first data item that has been allocated in the first sub-queue to the target second data item in the second sub-queue to obtain a target second data item storing the item index of the target first data item.
4. The scheduling queue according to claim 3, wherein, The queue control module further includes: an item index mapping module, configured to map the ready information of the target instruction data corresponding to the target first data item to the target second data item in the second sub-queue, so that the ready information of the target instruction data is also stored in the target second data item.
5. The scheduling queue according to claim 4, wherein, The second sub-queue is further configured to output the item index of the target second data item corresponding to the ready information according to the ready information of the instruction data stored in the target first data item obtained from the first sub-queue, for selecting the instruction data to be issued corresponding to the target second data item in the first sub-queue.
6. The scheduling queue according to any one of claims 2-4, wherein, The item index mapping module includes: a mapping sub-module, configured to output the item index of the target second data item according to the item index of the target first data item stored in the input target second data item.
7. The scheduling queue according to any one of claims 1-5, wherein, The first sub-queue is further configured to keep the instruction data stored in each non-empty first data item in the first sub-queue fixed and wait for the ready information corresponding to the stored instruction data.
8. The scheduling queue according to any one of claims 1-5, wherein, The first sub-queue is further configured to, in response to the source operands of the instruction data in the non-empty target first data item in the first sub-queue being ready and being selected to be issued out of the scheduling queue, indicate that the valid bit information of the target first data item is invalid, The queue control module is further configured to, in response to the valid bit information of the target first data item indicating invalid, indicate that the valid bit information of the target second data item corresponding to the target first data item is invalid.
9. The scheduling queue according to claim 8, wherein, The second sub-queue is further configured to, in response to the valid bit information of the target second data item indicating invalid, move at least one other non-empty second data item other than the target second data item to continuously store the current non-empty second data items in the second sub-queue in the predetermined order.
10. The scheduling queue according to any one of claims 1-5, wherein, The second sub-queue is further configured to shift the item index of the first data item stored in the second data item bit by bit according to the new and old order in a displacement manner.
11. A processor, comprising the scheduling queue according to any one of claims 1-10.
12. The processor according to claim 11, further comprising: An instruction selection module, configured to, in response to receiving, select the instruction data to be issued in the first sub-queue according to the item indexes of at least one first data item determined according to the item indexes of at least one second data item corresponding to the ready information provided by the second sub-queue.
13. The processor according to claim 12, wherein, The instruction selection module is further configured to, among the item indexes of the at least one first data item, select the instruction data of the first data item corresponding to the oldest ready information to be issued from the first sub-queue.
14. The processor according to any one of claims 11-13, further comprising: A rename module, configured to perform register renaming on the received instruction data and then send it to the first sub-queue for storage in the first sub-queue; And An execution unit module, configured to receive the instruction data issued by the first sub-queue for execution.
15. An electronic device, comprising the processor according to any one of claims 11-14.
16. An instruction scheduling method, comprising: Determining free first data items among N first data items in a first sub-queue of a scheduling queue to store instruction data input to the first sub-queue; Mapping and storing the item indexes of the non-empty first data items in the first sub-queue into N second data items of the second sub-queue of the scheduling queue to obtain corresponding non-empty second data items, wherein there is a corresponding relationship between the non-empty first data items in the first sub-queue and the non-empty second data items in the second queue; Continuously storing the non-empty second data items among the N second data items in the second sub-queue in a predetermined order, where N is an integer greater than 1.
17. The instruction scheduling method according to claim 16, wherein The determining free first data items among N first data items in a first sub-queue of a scheduling queue to store instruction data input to the first sub-queue includes: In response to target instruction data input to the first sub-queue, obtaining the item index of the target first data item that is currently in an idle state in the first sub-queue, and allocating the target first data item to store the target instruction data of the first sub-queue.
18. The instruction scheduling method according to claim 17, wherein, The mapping and storing the item indexes of the non-empty first data items in the first sub-queue into N second data items of the second sub-queue to obtain corresponding non-empty second data items includes: Map the allocated non-empty target first data items in the first sub-queue to the target second data items in the second sub-queue to obtain target second data items storing the item indices of the target first data items.
19. The instruction scheduling method according to any one of claims 16-18, further comprising: In response to the source operands of the instruction data in the non-empty target first data items in the first sub-queue being ready and being selected to be issued from the scheduling queue, indicate that the valid bit information corresponding to the target first data items is invalid, and indicate that the valid bit information of the target second data items corresponding to the target first data items is invalid.
20. The instruction scheduling method according to claim 19, further comprising: In response to the valid bit information of the target second data items indicating invalid, move at least one other non-empty second data item other than the target second data item to continuously store the current non-empty second data items in the second sub-queue in a predetermined order.