Thread Scheduling Method and Device, Processor, and Computer-Readable Storage Medium
By obtaining the resource usage and reservation of threads in the processor and selecting target threads that meet the conditions for scheduling, the problem of reducing parallelism caused by resource competition among threads is solved, and the performance and resource utilization of the processor are improved.
Patent Information
- Application Number
- CN202410322880.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-03-20
AI Technical Summary
In simultaneous multi-threaded processors, as the number of threads increases, competition for shared resources between threads becomes more intense, resulting in weaker parallelism and lowering processor performance and resource utilization.
By obtaining the multiple threads to be scheduled in the processor, their number of occupies and reservations for object resources, the target thread that meets the scheduling conditions is selected for scheduling based on this information. Specifically, the target thread must be active, and the object resource has sufficient resources available for allocation after removing the reservations from other threads.
This method effectively alleviates resource competition among threads, improves the parallelism and resource utilization of the processor, avoids the situation of starving threads, and enhances the overall design robustness of the processor.
Smart Images

Figure CN118132233B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a thread scheduling method, an apparatus, a processor, and a computer-readable storage medium. Background Art
[0002] Modern multi-issue high-performance CPUs (Central Processing Units) include at least one processor core, and each processor core includes multiple execution units to execute instructions. For example, the pipeline process of instruction execution includes five stages: instruction fetch (IF), instruction dispatch / decode (ID), execution (EX), memory access (MEM), and write back (WB, updating the result obtained after instruction execution to a register).
[0003] To improve the parallelism of instruction execution in a processor, the processor can adopt Simultaneous Multithreading (SMT) technology, so that the pipeline structure (also simply referred to as "pipeline") for instruction execution in the processor can support the simultaneous execution of two or more (hardware) threads. For example, an SMT processor can be SMT2 (supporting up to two concurrent threads), SMT4 (supporting up to four concurrent threads), or SMT8 (supporting up to eight concurrent threads), etc. When a single-threaded processor is running, due to stalls such as cache misses, multiple execution units inside the processor may be idle. For a processor adopting SMT technology, even when a certain thread pauses, other threads can continue to use these hardware resources, which improves the utilization rate of hardware resources, can avoid idling, and improves the processor throughput and performance power ratio.
[0004] However, with the increase in the number of threads in an SMT processor, the competition for shared resources among threads becomes more intense, and resource competition will lead to a decrease in the parallelism of multiple threads, reducing the performance and resource utilization rate of the processor. Therefore, how to alleviate resource competition among threads and improve parallelism has become an important research direction for simultaneous multithreading processors. Summary of the Invention
[0005] At least one embodiment of the present disclosure provides a thread scheduling method, where the thread scheduling method includes: obtaining multiple threads to be scheduled for an object resource in a processor, and obtaining the current occupation quantity and reservation quantity of the multiple threads for the object resource respectively; based on the current occupation quantity and reservation quantity of the multiple threads for the object resource respectively, selecting a target thread that meets the scheduling condition from the multiple threads for scheduling.
[0006] For example, in the thread scheduling method according to at least one embodiment, the scheduling condition includes that the target thread is currently active, and the object resource has a first quantity available for allocation after removing the quantity reserved for other active threads except the target thread.
[0007] For example, in the thread scheduling method according to at least one embodiment, selecting a target thread that meets the scheduling condition from the multiple threads for scheduling includes: performing arbitration based on a selection algorithm to select the target thread from one or more alternative threads that all meet the scheduling condition.
[0008] For example, in the thread scheduling method according to at least one embodiment, the selection algorithm includes the Least Recently Used (LRU) algorithm, the Least Frequently Used (LFU) algorithm, or the Round Robin (RR) algorithm.
[0009] For example, in the thread scheduling method according to at least one embodiment, selecting a target thread that meets the scheduling condition from the multiple threads for scheduling includes: obtaining a target instruction to be sent by the target thread and a second quantity of the target instruction, and obtaining a first quantity available for allocation of the object resource after removing the quantity reserved for other active threads except the target thread, where the first quantity is greater than 0; in response to the second quantity being less than or equal to the first quantity, sending all of the target instruction to process the instruction using the object resource, or, in response to the second quantity being greater than the first quantity, sending a portion of the first quantity of the target instruction to process the instruction using the object resource.
[0010] For example, in the thread scheduling method according to at least one embodiment, before obtaining the current occupied quantity and reserved quantity of each of the multiple threads for the object resource, it further includes: determining, for each thread in the multiple threads, the reserved quantity for the object resource.
[0011] For example, in the thread scheduling method according to at least one embodiment, determining, for each thread in the multiple threads, the reserved quantity for the object resource includes: for the current thread in the multiple threads, determining whether the current thread is active; in response to the current thread being inactive, not reserving for the current thread, or, in response to the current thread being active, and determining the reserved quantity for the object resource based on the occupied quantity of the current thread for the object resource.
[0012] For example, in the thread scheduling method according to at least one embodiment, the number of the multiple threads is greater than or equal to 4.
[0013] For example, in the thread scheduling method according to at least one embodiment, the object resources include a cache, a branch predictor storage space, an out-of-order scheduler storage space, and a data prefetcher storage space.
[0014] At least one embodiment of the present disclosure further provides a thread scheduling apparatus, where the thread scheduling apparatus includes: an obtaining module configured to obtain a plurality of threads to be scheduled for object resources in a processor, and obtain the current occupation quantity and reservation quantity of the plurality of threads for the object resources respectively; a scheduling module configured to select a target thread that meets the scheduling conditions from the plurality of threads for scheduling based on the current occupation quantity and reservation quantity of the plurality of threads for the object resources respectively.
[0015] For example, in the thread scheduling apparatus according to at least one embodiment, the scheduling module includes a thread selection module, and the thread selection module is configured to perform arbitration based on a selection algorithm to select the target thread from one or more alternative threads that all meet the scheduling conditions.
[0016] For example, in the thread scheduling apparatus according to at least one embodiment, the scheduling module further includes an instruction selection module, and the instruction selection module is configured to: obtain a target instruction to be sent by the target thread and a second quantity of the target instruction, and obtain a first quantity available for allocation of the object resource after removing the quantity reserved for other active threads except the target thread, where the first quantity is greater than 0; and in response to the second quantity being less than or equal to the first quantity, send all of the target instructions to use the object resource for instruction processing, or in response to the second quantity being greater than the first quantity, send a part of the target instructions with the first quantity to use the object resource for instruction processing.
[0017] For example, the thread scheduling apparatus according to at least one embodiment further includes a resource reservation module, and the resource reservation module is configured to determine a reservation quantity for the object resources for each of the plurality of threads.
[0018] At least one embodiment of the present disclosure further provides a thread scheduling apparatus, including: a processing unit and a memory; where the memory stores computer-readable instructions and is communicatively connected to the processing unit; the processing unit executes the computer-readable instructions stored in the memory to implement the thread scheduling method provided in any of the above embodiments.
[0019] At least one embodiment of the present disclosure further provides a processor according to the thread scheduling apparatus provided in at least one embodiment of the present disclosure.
[0020] At least one embodiment of the present disclosure further provides a computer-readable storage medium, wherein computer-readable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-readable instructions, the thread scheduling method provided in any of the above embodiments is implemented. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0022] Figure 1 A schematic diagram showing a pipeline of a processor core is shown;
[0023] Figure 2 A schematic diagram showing a processor pipeline involving out-of-order execution and register renaming is shown;
[0024] Figure 3 A schematic diagram showing a thread scheduling process of an SMT4 processor is shown;
[0025] Figure 4 A schematic flowchart showing the thread scheduling method provided by at least one embodiment of the present disclosure is shown;
[0026] Figure 5 A schematic structural diagram showing a thread scheduling device provided by at least one embodiment of the present disclosure is shown;
[0027] Figure 6 A schematic structural diagram showing another thread scheduling device provided by at least one embodiment of the present disclosure is shown;
[0028] Figure 7 A schematic structural diagram showing another thread scheduling device provided by at least one embodiment of the present disclosure is shown;
[0029] Figure 8 A schematic diagram showing a computer-readable storage medium provided by at least one embodiment of the present disclosure is shown;
[0030] Figure 9 A schematic block diagram showing an electronic device provided by at least one embodiment of the present disclosure is shown. Detailed Embodiments
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0032] Unless otherwise defined, the technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a", "an" or "the" do not denote a limitation of quantity, but rather denote the presence of at least one. Words such as "comprising" or "including" mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left" and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0033] To improve the performance of a processor core, the processor core may use a pipelined approach, that is, the entire process of an instruction from being fetched, decoded, executed, and the result written back is divided into multiple pipeline stages, and in one operation cycle, an instruction can only be in one of the pipeline stages; while the processor core can run multiple instructions in different pipeline stages.
[0034] Figure 1 A schematic diagram of a pipeline of a processor core is shown, and the dashed line with an arrow in the figure represents the redirected instruction stream.
[0035] As Figure 1As shown, the processor core (e.g., CPU core) of a single-core processor or multi-core processor improves the instruction-level parallelism through pipelining technology. The pipeline inside the processor core includes multiple pipeline stages. For example, after the program counter from various sources is fed into the pipeline and the next program counter (PC) is selected through a multiplexer (Mux), the instruction corresponding to this program counter has to go through branch prediction, instruction fetch, instruction decode, instruction dispatch and rename, instruction execution, instruction retirement, etc. Waiting queues are set between each pipeline stage as needed, and these queues are usually first-in-first-out (FIFO) queues. For example, after the branch prediction unit, there is a branch prediction (BP) FIFO queue to store the branch prediction results; after the instruction fetch unit, there is an instruction cache (IC) FIFO to cache the fetched instructions; after the instruction decode unit, there is a decode (DE) FIFO to cache the decoded instructions; after the instruction dispatch and rename unit, there is a retirement (RT) FIFO to cache the instructions waiting for confirmation of retirement after execution. At the same time, the pipeline of the processor core also includes an instruction queue to cache the instructions waiting for the instruction execution unit to execute after instruction dispatch and rename.
[0036] Figure 2 FIG. shows a schematic diagram of a processor pipeline involving out-of-order execution and register renaming. As Figure 2As shown, after each instruction (such as a branch instruction) in the processor is fetched from the instruction cache according to the program counter (PC) ("instruction fetch"), the fetched instruction is decoded ("decoding"); the decoded instruction undergoes a register renaming operation ("register renaming"), and the instruction after register renaming is sent for out-of-order execution to remove false dependencies between instructions; then the instruction enters the instruction distribution module, and the instruction distribution module decides when to distribute the instruction to which execution unit for execution ("out-of-order scheduling"). For example, in different operation cycles, the instruction will be distributed to different execution units corresponding to different ports (such as port 0, port 1, port 2, port 3, etc.) (such as arithmetic logic unit (ALU), multiplication unit (MUL), division unit (DIV), load and store unit (LSU), etc.). While performing register renaming, the instruction will enter the instruction commit unit, and the instruction commit unit records the original instruction fetch order of the instructions processed in the pipeline. The instruction commit unit is used to commit the instructions in the original instruction fetch order after the instructions are executed, and at the same time, it updates the actual execution information of the branch instruction to the branch prediction unit.
[0037] As described above, the possible pipeline conflicts of WAW (write after write) and WAR (read after write) can be solved through the register renaming technology. This technology does not increase the number of general-purpose registers. Physical registers (PR) are redefined inside the processor, and the registers defined in the instruction set are called architectural registers (AR). Physical registers actually exist in the processor, such as what is called the physical register file (PRF). The processor will dynamically map the architectural register AR to the physical register PR to solve the problems of WAW and WAR dependencies; the processor completes the process of register renaming through the architectural register mapping table and the physical register available list. When the processor performs register renaming on the architectural register used in the current instruction, it needs to process both the source register and the destination register (both are logical registers) in the instruction. For the source register, the processor will look up the architectural register mapping table to find the corresponding PR number (PRN), and for the destination register, it needs to read a PR number from the physical register available list, establish a mapping relationship between the PR number and the destination register, and write it into the architectural register mapping table. If the free list is empty, the pipeline of the processor needs to pause and wait until an instruction retires and releases a PR. The architectural register mapping table is implemented through hardware structures such as storage devices (such as caches or registers), etc.
[0038] In addition, in order to improve the parallelism of instruction execution in a processor, the processor can also adopt the Simultaneous Multi-Threading (SMT) technology. The pipeline structure (also simply referred to as "pipeline") of the processor for instruction execution can support the simultaneous execution of two or more (hardware) threads. For example, it can be SMT2 (supporting up to two concurrent threads), SMT4 (supporting up to four concurrent threads), or SMT8 (supporting up to eight concurrent threads), etc. In the pipeline of a processor that supports the SMT technology, the computing resources in the processor are shared by multiple threads. For example, each thread can have its own independent logical registers, but multiple physical registers within the processor are shared by multiple threads; in the queues of various control functions in the pipeline, some can be shared by multiple threads, such as the instruction scheduling queue, and some can be statically partitioned among multiple threads, such as the instruction reordering queue. The SMT technology can utilize the parallelism between threads to improve the utilization rate of pipeline resources.
[0039] In addition, a processor adopting the Simultaneous Multi-Threading (SMT) technology can operate in the SMT mode or in the single-thread mode. For example, in the SMT mode, when one thread encounters a waiting state, other threads can continue to execute, which can effectively improve the utilization rate of hardware resources, and thus enhance the multi-threading processing ability, overall performance, and performance power ratio of the CPU core.
[0040] In the present disclosure, the "operation cycle" can be, for example, a clock cycle or a machine cycle, or can also be other time periods for completing one beat of operation in the instruction pipeline of the processor. The execution process of an instruction in each thread includes several stages, and each stage completes a basic operation (for example, fetching an instruction, memory read, memory write, etc.). The time required to complete a basic operation is called a machine cycle, also known as a CPU cycle.
[0041] Figure 3 The figure shows a schematic diagram of the thread scheduling process of an SMT4 processor. Taking the SMT4 processor that can schedule 2 threads in each operation cycle as an example. As Figure 3 shown, for the scheduling of a queue shared by multiple threads in a certain pipeline stage, the thread numbers T0 - T3 of 4 threads are stored in a First-In-First-Out (FIFO) queue. From the head to the tail of the FIFO queue, the priorities of the 4 threads are arranged from high to low. In operation cycle cycle 0, the initial state of the FIFO queue stores the thread priority order as T0, T1, T2, and T3, and all 4 threads are in a schedulable state. When threads T0 and T1 are successfully scheduled, the state of the FIFO queue is updated to T2, T3, T0, and T1. In operation cycle cycle 1, the thread priority order is T2, T3, T0, and T1. Since thread T2 is in an unschedulable state ( Figure 2(The thread in the gray box indicates that the thread is in a non - schedulable state), then the schedulable threads are thread T3, T0, and T1. After threads T3 and T0 are successfully scheduled, the state of the FIFO queue is updated to T2, T1, T3, and T0. The subsequent thread scheduling process follows the same pattern and will not be elaborated here.
[0042] For a simultaneous multi - threaded processor (hereinafter also simply referred to as an SMT processor), when the schedulable shared resources meet the requirements of all threads to be scheduled (such as thread T0 and thread T1), the threads to be scheduled can be scheduled simultaneously; otherwise, there are resource conflicts among the threads to be scheduled, and shared resource arbitration is required to determine the target thread for scheduling.
[0043] For example, the resources shared by multiple threads in an SMT processor include but are not limited to: caches (such as level - 1 data cache, level - 1 instruction cache, level - 2 cache, etc.), branch predictor storage space, out - of - order scheduler storage space for register renaming, data prefetching storage space, and so on.
[0044] The inventors of the present disclosure noticed that in existing SMT processors, the control of shared resources for multiple threads (such as the out - of - order scheduler storage space) is designed based on a two - thread - at - a - time scenario. In a two - thread scenario, each thread only needs to ensure that it does not exhaust the shared resources when competing for shared resources to ensure that the other thread still has resources available and will not starve the other thread. However, for processors with more threads (such as four or more threads), if each thread only considers not exhausting the shared resources in the processor pipeline, since each thread does not know the resource occupancy of the other three threads, it is very likely that the three threads will exhaust the shared resources together, resulting in the remaining one thread having no resources available, leading to bugs or hangs.
[0045] At least one embodiment of the present disclosure provides a thread scheduling method, apparatus, processor, and computer - readable storage medium.
[0046] The thread scheduling method includes: obtaining multiple threads to be scheduled for an object resource in a processor, and obtaining the current occupancy quantity and reserved quantity of each of the multiple threads for the object resource; based on the current occupancy quantity and reserved quantity of each of the multiple threads for the object resource, selecting a target thread that meets the scheduling conditions from the multiple threads for scheduling.
[0047] The thread scheduling device includes an acquisition module and a scheduling module; the acquisition module is configured to acquire a plurality of threads to be scheduled for object resources in a processor, and acquire the current occupation quantity and reserved quantity of the plurality of threads for the object resources respectively; the scheduling module is configured to select a target thread that meets the scheduling condition from the plurality of threads for scheduling based on the current occupation quantity and reserved quantity of the plurality of threads for the object resources respectively.
[0048] In at least one embodiment of the present disclosure, for example, in a simultaneous multi-threading (SMT4) processor, for the shared resources ("object resources") of the scheduling object, when each selected thread that uses the shared resources determines how much resources it can consume, in addition to considering how much of the total quantity of the object resources remains, it also needs to consider the quantity of resources reserved for the other three threads. When there are remaining shared resources available after reserving sufficient resources for the other threads, thread scheduling is considered, otherwise, it is necessary to wait. Further, for example, before scheduling each thread, it is first checked whether object resources are reserved for the thread and how much quantity of resources is reserved.
[0049] The thread scheduling method and device of the above embodiments of the present disclosure can ensure fairness among multiple threads in the processor, and will not starve a certain thread, improving the overall robustness of the processor design.
[0050] The following non-limitingly describes the disclosed thread scheduling method and device in conjunction with one or more specific embodiments.
[0051] Figure 4 The flowchart of the thread scheduling method provided by at least one embodiment of the present disclosure is shown, and the thread scheduling method includes step S410 - step S420.
[0052] Step S410: Acquire a plurality of threads to be scheduled for object resources in a processor, and acquire the current occupation quantity and reserved quantity of the plurality of threads for the object resources respectively;
[0053] Step S420: Select a target thread that meets the scheduling condition from the plurality of threads for scheduling based on the current occupation quantity and reserved quantity of the plurality of threads for the object resources respectively.
[0054] For example, the processor is an SMT processor (such as SMT4, SMT8, etc.), and there are multiple instructions belonging to multiple threads running in the pipeline of the processor. For the scheduling of the resources ("object resources") shared by multiple threads in a certain pipeline stage, it is necessary to consider the current occupation quantity and reserved quantity of the object resources by the multiple threads respectively, and then select a target thread that meets the scheduling condition from the multiple threads for scheduling based on the occupation quantity and reserved quantity.
[0055] Here, for the object resources at a certain pipeline stage, the "occupied quantity" is used to refer to the quantity of object resources occupied by a certain thread and not yet released during the current operation cycle of the thread; the "reserved quantity" is used to refer to the quantity of object resources reserved and idle (and thus available for allocation) by a certain thread for the current operation cycle.
[0056] For example, in at least one example, before obtaining the occupied quantities and reserved quantities of multiple threads for object resources respectively, the thread scheduling method according to the embodiments of the present disclosure further includes: determining the reserved quantity of object resources for each thread among the multiple threads.
[0057] For example, in at least one example, determining the reserved quantity of object resources for each thread among the multiple threads includes: for the current thread among the multiple threads, determining whether the current thread is active; in response to the current thread being inactive, not reserving for the current thread, or, in response to the current thread being active, and based on the occupied quantity of the current thread for the object resources, determining the reserved quantity of the current thread for the object resources.
[0058] For example, taking an SMT4 processor and an out-of-order scheduler storage space as an example, where the 4 threads to be processed are threads T0 to T3 respectively, before scheduling in the current operation cycle, determining the reserved quantity of object resources for each of the threads T0 to T3. For example, the reservation rules are as follows:
[0059] · If a certain thread is inactive, no resources are reserved for this thread;
[0060] · If a certain thread is active and the resources already occupied in the out-of-order scheduler storage space are 0, then 2 storage resources are reserved for this thread;
[0061] · If a certain thread is active and the resources already occupied in the out-of-order scheduler storage space are 1, then 1 storage resource is reserved for this thread;
[0062] · If a certain thread is active and the resources already occupied in the out-of-order scheduler storage space are not less than 2,
[0063] then no storage resources need to be reserved for this thread.
[0064] This reservation rule is, for example, fixed during the operation of the processor, or can be changed according to circumstances during the operation of the processor. For example, a register can be used to store this reservation rule, and corresponding modification and adjustment can also be performed.
[0065] In one example, assume that the total amount of resources in the out-of-order scheduler storage space is 10 (e.g., 10 storage entries). Before the current operation, the occupied and reserved amounts of resources in the out-of-order scheduler storage space by threads T0 to T3 are as follows:
[0066] Thread T0: Occupied 0; Reserved 2;
[0067] Thread T1: Occupied 1; Reserved 1;
[0068] Thread T2: Occupied 3; Reserved 0;
[0069] Thread T3: Occupied 0; Reserved 2;
[0070] At this time, in the out-of-order scheduler storage space, the amount of occupied resources is 4, the amount of reserved resources is 5, and the amount of unoccupied and non-reserved resources is 1.
[0071] For example, in at least one example, the scheduling conditions include: the target thread is currently active, and the object resources have a first amount available for allocation after removing the amount reserved for other active threads other than the target thread.
[0072] Here, "active" is used to refer to a thread that is not suspended, paused, or aborted by the system, etc., and the corresponding instruction pipeline is in a state of continuous execution, so it is also in a schedulable state.
[0073] For example, when performing thread scheduling in the current operation cycle, the primary condition is to select candidate threads from active threads, that is, non-active threads will be excluded from candidate threads; secondly, it is also required that the current object resources have free resources available for allocation after removing the amount reserved for other active threads other than the target thread. The free resources available for allocation include the resources reserved for the target thread and any non-reserved free resources. Conversely, for a certain thread, if the resources reserved for the thread and any non-reserved free resources are both 0, then the thread will not be used as a candidate thread.
[0074] Assume that in the above example, in the current operation cycle, threads T0 to T3 are all active except for thread T2, that is, only thread T2 is non-active. Therefore, threads T0, T1, and T3 may all be selected for scheduling; on the other hand, assume that the amount of resources required by threads T0 to T3 in the current cycle is 3. Therefore, for threads T0, T1, and T3, assuming they are all scheduled, there will be at least 1 amount of free resources available. At this time, threads T0, T1, and T3 can all be used as candidate threads.
[0075] For example, in at least one example, selecting a target thread that meets the scheduling conditions for scheduling among multiple threads includes: performing arbitration based on a selection algorithm to select a target thread from one or more alternative threads that all meet the scheduling conditions.
[0076] On the other hand, it is also necessary to consider the relationships between threads and the operation history. For example, to achieve fair scheduling among multiple threads, arbitration can be performed based on the selection algorithm to select a target thread from all the one or more alternative threads that meet the scheduling conditions.
[0077] For example, in at least one example, the selection algorithm includes the Least Recently Used (LRU) algorithm, the Least Frequently Used (LFU) algorithm, or the Round Robin (RR) algorithm.
[0078] For example, for the case where the above threads T0, T1, and T3 are all alternative threads, assuming that among them, in a previous period of time, the number of times they were scheduled from most to least is: T3, T0, and T1. Therefore, in the example of using the LRU algorithm, the thread T1 that was least scheduled before will be selected from the threads T0, T1, and T3 as the target thread for scheduling.
[0079] For example, in at least one example, selecting a target thread that meets the scheduling conditions for scheduling among multiple threads includes: obtaining the target instruction to be sent by the target thread and the second quantity of the target instruction, and obtaining the first quantity that the object resource has available for allocation after removing the quantity reserved for other active threads except the target thread, where the first quantity is greater than 0; in response to the second quantity being less than or equal to the first quantity, sending all of the target instructions to use the object resource for instruction processing, or, in response to the second quantity being greater than the first quantity, sending a partial instruction of the first quantity in the target instructions to use the object resource for instruction processing.
[0080] In the above example of selecting thread T1 as the target thread for scheduling, the number of resources required by thread T1 is 3 (for example, 3 instructions need to perform register renaming operations), the number of resources reserved for thread T1 is 1, and the number of non-reserved and idle resources is 1. That is, currently, the number of resources required by thread T1 (3) is more than the number available to thread T1 (2). Then, 2 out of the 3 instructions that need to perform register renaming operations for thread T1 can be selected and sent to the out-of-order scheduler storage space for register renaming operations, and the remaining 1 instruction will be considered in the next operation cycle.
[0081] For example, in at least one example, the number of the multiple threads is greater than or equal to 4. For example, it can be 4, 6, 8, etc., that is, the corresponding processors are SMT4, SMT6, SMT8, etc.
[0082] For example, in at least one example, the object resources include a cache, a branch predictor storage space, an out-of-order scheduler storage space, a data prefetcher storage space, and the like.
[0083] Figure 5 FIG. shows a schematic structural diagram of a thread scheduling device provided by at least one embodiment of the present disclosure.
[0084] As Figure 5 shown, the thread scheduling device 500 includes an obtaining module 510 and a scheduling module 520. The obtaining module 510 is configured to obtain a plurality of threads to be scheduled for object resources in a processor, and obtain the current occupation quantity and reservation quantity of the object resources by the plurality of threads respectively; the scheduling module 520 is configured to select a target thread that meets the scheduling condition from the plurality of threads for scheduling based on the current occupation quantity and reservation quantity of the object resources by the plurality of threads respectively.
[0085] For example, in at least one embodiment, the scheduling condition includes that the target thread is currently active, and the object resource has a first quantity available for allocation after removing the quantity reserved for other active threads other than the target thread.
[0086] For example, in at least one embodiment, the scheduling module includes a thread selection module, and the thread selection module is configured to perform arbitration based on a selection algorithm to select a target thread from one or more alternative threads that all meet the scheduling condition.
[0087] For example, in at least one embodiment, the selection algorithm includes the Least Recently Used (LRU) algorithm, the Least Frequently Used (LFU) algorithm, or the Round Robin (RR) algorithm.
[0088] For example, in at least one embodiment, the scheduling module further includes an instruction selection module, and the instruction selection module is configured to: obtain a target instruction to be sent by the target thread and a second quantity of the target instruction, and obtain that the object resource has a first quantity available for allocation after removing the quantity reserved for other active threads other than the target thread, where the first quantity is greater than 0; and in response to the second quantity being less than or equal to the first quantity, send all of the target instructions to use the object resource for instruction processing, or in response to the second quantity being greater than the first quantity, send a portion of the target instructions equal to the first quantity to use the object resource for instruction processing.
[0089] Again, Figure 5 shown, for example, in at least one embodiment, the thread scheduling device further includes a resource reservation module 530, and the resource reservation module 530 is configured to determine a reservation quantity of the object resource for each of the plurality of threads.
[0090] For example, in at least one embodiment, the resource reservation module 530 is further configured to include: for the current thread among a plurality of threads, determine whether the current thread is active; in response to the current thread being inactive, do not reserve for the current thread, or, in response to the current thread being active, and based on the number of object resources occupied by the current thread, determine the reservation number of object resources for the current thread.
[0091] For example, in at least one embodiment, the number of the plurality of threads is greater than or equal to 4, such as 4, 6, 8, etc.
[0092] For example, in at least one embodiment, the above object resources may include caches, branch predictor storage spaces, out-of-order scheduler storage spaces, data prefetching storage spaces, etc., and the embodiments of the present disclosure do not limit this.
[0093] Figure 6 The structural schematic diagram of another thread scheduling device provided by at least one embodiment of the present disclosure is shown.
[0094] The thread scheduling device 600 is used for, for example, the out-of-order scheduler in an SMT4 processor, and includes an acquisition module 610, a scheduling module 620, four resource reservation modules 630, and an out-of-order scheduler storage space 640, and the scheduling module 620 includes a thread selection module 621 and an instruction selection module 622.
[0095] The acquisition module 610 is configured to acquire 4 threads T0 to T3 to be scheduled for the out-of-order scheduler storage space 640 in the processor, and acquire the current occupation numbers and reservation numbers of the 4 threads T0 to T3 for the out-of-order scheduler storage space 640 respectively; the scheduling module 620 is configured to select a target thread that meets the scheduling conditions from the 4 threads T0 to T3 based on the current occupation numbers and reservation numbers of the 4 threads T0 to T3 for the out-of-order scheduler storage space 640 respectively for scheduling.
[0096] As Figure 6 shown, the four resource reservation modules 630 are respectively configured to determine the reservation numbers for the out-of-order scheduler storage space 640 for each of the 4 threads T0 to T3. Here, the number of resource reservation modules 630 corresponds one-to-one to the number of threads, and each resource reservation module 630 is used for one thread, but the embodiments of the present disclosure do not limit this.
[0097] The thread selection module 621 is configured to perform arbitration based on a selection algorithm to select a target thread from one or more alternative threads that all meet the scheduling conditions. In this embodiment, the selection algorithm adopted is the RLU algorithm. The instruction selection module 622 is configured to: obtain the target instruction to be sent by the target thread and the second quantity of the target instruction, and obtain the first quantity available for allocation in the out-of-order scheduler storage space 640 after removing the quantity reserved for other active threads except the target thread, where the first quantity is greater than 0; and in response to the second quantity being less than or equal to the first quantity, send all of the target instructions to process the instructions using the object resources, or in response to the second quantity being greater than the first quantity, send a portion of the target instructions equal to the first quantity to process the instructions using the object resources.
[0098] Figure 7 FIG. shows a schematic structural diagram of another thread scheduling device provided by at least one embodiment of the present disclosure.
[0099] As Figure 7 shown, the thread scheduling device 700 includes a processing unit 720 and a memory 710. The memory 710 stores computer-readable instructions and is communicatively connected to the processing unit 720. The processing unit 720 executes the computer-readable instructions stored in the memory 710 to implement the thread scheduling method provided by any embodiment of the present disclosure.
[0100] For example, the memory 710 and the processing unit 720 may communicate with each other directly or indirectly. For example, in some examples, as Figure 7 shown, the thread scheduling device 700 may further include a system bus 730, and the memory 710 and the processing unit 720 may communicate with each other through the system bus 730. For example, the processing unit 720 may access the memory 710 through the system bus 730. For example, in other examples, components such as the memory 710 and the processing unit 720 may communicate through a network-on-chip (NOC) connection.
[0101] For example, the processing unit 720 may control other components in the thread scheduling device to perform desired functions. The processing unit 720 may be a central processing unit (CPU), a tensor processing unit (TPU), a network processor (NP), or a graphics processing unit (GPU) and other devices with data processing capabilities and / or program execution capabilities, and may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0102] For example, the memory 710 may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.
[0103] For example, one or more computer-readable instructions may be stored on the memory 710, and the processing unit 720 may run the computer-readable instructions to implement various functions. Various application programs and various data may also be stored in the computer-readable storage medium, such as instruction processing code and various data used and / or generated by the application programs.
[0104] For example, when some computer instructions stored in the memory 710 are executed by the processing unit 720, one or more steps in the thread scheduling method described above may be executed.
[0105] For example, as Figure 7 shown, the thread scheduling device 700 may further include an input interface 740 that allows external devices to communicate with the thread scheduling device 700. For example, the input interface 740 may be used to receive instructions from external computer devices, from users, etc. The thread scheduling device 700 may further include an output interface 750 that connects the thread scheduling device 700 and one or more external devices to each other. For example, the thread scheduling device 700 may pass through the output interface 750, etc.
[0106] It should be noted that the thread scheduling device provided in the embodiments of the present disclosure is exemplary and not restrictive. According to actual application needs, the thread scheduling device may further include other conventional components or structures. For example, to implement the necessary functions of the thread scheduling device, those skilled in the art may set other conventional components or structures according to specific application scenarios, and the embodiments of the present disclosure do not limit this.
[0107] At least one embodiment of the present disclosure further provides a processor, which is, for example, an SMT processor. The processor includes the thread scheduling device provided in any embodiment of the present disclosure. The maximum number of threads that the SMT processor can support may be, for example, 2, 4, 8, etc., and it may be a single-core or multi-core processor. For example, the processor core may adopt microarchitectures such as X86, ARM, RISC-V, etc., and may include one or more levels of cache. The embodiments of the present disclosure do not limit this.
[0108] At least one embodiment of the present disclosure further provides a computer-readable storage medium.Figure 8 Schematic diagram of a computer-readable storage medium provided by at least one embodiment of the present disclosure.
[0109] For example, as Figure 8 shown, the computer-readable storage medium 800 stores computer-readable instructions 810, and when the computer-readable instructions 810 are executed by a computer (including a processor), the thread scheduling method provided by any embodiment of the present disclosure can be implemented.
[0110] For example, one or more computer-readable instructions may be stored on the computer-readable storage medium 800. Some of the computer-readable instructions stored on the computer-readable storage medium 800 may be instructions for implementing one or more steps in the above thread scheduling method, for example.
[0111] For example, the computer-readable storage medium may include a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above computer-readable storage media, and may also be other applicable storage media. For example, the computer-readable storage medium 800 may include the memory 710 in the above thread scheduling device 700.
[0112] At least some embodiments of the present disclosure also provide an electronic device, and the electronic device includes the processor of any of the above embodiments. Figure 9 Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.
[0113] The electronic device in the embodiments of the present disclosure may be implemented as, but not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0114] Figure 9 The illustrated electronic device 900 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0115] For example, as Figure 9As shown, in some examples, the electronic device 900 includes a processor 901, which may include the processor of any of the above embodiments and may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 902 or a program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the computer system are also stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0116] For example, the following components may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; a communication device 909 including, for example, a network interface card such as a LAN card, a modem, etc. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 910 as needed so that a computer program read from it can be installed into the storage device 908 as needed. Although Figure 9 an electronic device 900 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.
[0117] For example, the electronic device 900 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 909 may communicate with the network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on the IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0118] Regarding the present disclosure, in addition to the above exemplary description, the following points need to be noted:
[0119] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures may refer to the general design.
[0120] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.
[0121] The above is only an exemplary implementation manner of the present disclosure, rather than being used to limit the protection scope of the present disclosure. The protection scope of the present disclosure is determined by the appended claims.
Claims
1. A thread scheduling method, comprising: Acquire multiple threads to be scheduled for object resources in the processor, and acquire the current occupied quantity and reserved quantity of the object resources by the multiple threads respectively; Based on the current occupation quantity and reserved quantity of the object resources respectively by the multiple threads, a target thread meeting the scheduling condition is selected from the multiple threads for scheduling; The scheduling condition includes that the target thread is currently active, and the object resource has a first quantity available for allocation after removing the quantity reserved for other active threads except the target thread, The step of selecting a target thread that meets the scheduling condition from among the multiple threads for scheduling includes: Obtaining a target instruction to be sent by the target thread and a second number of the target instructions, and obtaining a first number of the object resources available for allocation after deducting the number reserved for other active threads except the target thread, wherein the first number is greater than 0; In response to the second number being less than or equal to the first number, all of the target instructions are sent to use the object resources for instruction processing, or, in response to the second number being greater than the first number, part of the first number of instructions in the target instructions are sent to use the object resources for instruction processing.
2. The thread scheduling method according to claim 1, wherein: The step of selecting a target thread that meets the scheduling condition from among the multiple threads for scheduling includes: Arbitration is performed based on a selection algorithm to select the target thread from one or more candidate threads that meet the scheduling condition.
3. The thread scheduling method according to claim 2, wherein: The selection algorithm includes a least recently used algorithm (LRU), a least recently used algorithm (LFU), or a round robin algorithm (RH).
4. The thread scheduling method according to any one of claims 1 to 3, before obtaining the current occupied quantity and reserved quantity of the object resource by the plurality of threads respectively, further comprising: A reserved quantity for the object resource is determined for each of the multiple threads.
5. The thread scheduling method according to claim 4, wherein: The step of determining a reserved quantity for the object resource for each of the multiple threads includes: For a current thread among the multiple threads, determining whether the current thread is active; In response to the current thread being inactive, no resource is reserved for the current thread; or, in response to the current thread being active, the reserved quantity of the object resource by the current thread is determined based on the occupied quantity of the object resource by the current thread.
6. The thread scheduling method according to any one of claims 1 to 3, wherein: The number of the multiple threads is greater than or equal to 4.
7. The thread scheduling method according to any one of claims 1 to 3, wherein: The object resources include cache, branch predictor storage space, out-of-order scheduler storage space, and data prefetcher storage space.
8. A thread scheduling device, comprising: An acquisition module is configured to acquire multiple threads to be scheduled for object resources in the processor, and acquire the current occupied quantity and reserved quantity of the object resources by the multiple threads respectively; A scheduling module, configured to select a target thread that meets the scheduling condition from among the multiple threads for scheduling based on the current occupation quantity and reserved quantity of the object resources respectively by the multiple threads; The scheduling condition includes that the target thread is currently active, and the object resource still has a first quantity available for allocation after removing the quantity reserved for other active threads except the target thread; and Instruction selection module, configured as: Obtaining a target instruction to be sent by the target thread and a second number of the target instructions, and obtaining a first number of the object resources available for allocation after deducting the number reserved for other active threads except the target thread, wherein the first number is greater than 0; and In response to the second number being less than or equal to the first number, all of the target instructions are sent to use the object resources for instruction processing, or, in response to the second number being greater than the first number, part of the first number of instructions in the target instructions are sent to use the object resources for instruction processing.
9. The thread scheduling device according to claim 8, wherein: The scheduling module includes: The thread selection module is configured to perform arbitration based on a selection algorithm to select the target thread from one or more candidate threads that meet the scheduling condition.
10. The thread scheduling device according to claim 8, further comprising: The resource reservation module is configured to determine a reserved quantity for the object resource for each thread in the multiple threads.
11. A thread scheduling device, comprising a processing unit and a memory; wherein: The memory stores computer-readable instructions and is communicatively connected to the processing unit; The processing unit executes the computer-readable instructions stored in the memory to implement the thread scheduling method according to any one of claims 1 to 7.
12. A processor, comprising the thread scheduling device according to any one of claims 8 to 11.
13. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-readable instructions. When the processor executes the computer-readable instructions, the thread scheduling method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Resource distribution method and device
CN105591809A
Systems, Methods, and Apparatuses for Thread Selection and Reservation Station Binding
US20160378497A1
Cited By
Thread scheduling method and apparatus, processor, and computer-readable storage medium
EP4650960A1
Thread scheduling method and apparatus, processor, and computer-readable storage medium
WO2025194590A1