Scheduler, scheduling method, scheduling device and apparatus
Patent Information
- Application Number
- CN202510389675.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]若采用上述这种线程束的选择方式,若最老的线程束与优先级最高的线程束存在数据缓存冲突,则最老的线程束的执行将有可能替换优先级最高的线程束在数据缓存中的缓存信息,将会影响后续继续执行优先级最高的线程束,导致影响处理器的执行效率
[0019]本申请提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122838005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of task scheduling technology, and in particular to a scheduler, scheduling method, scheduling device and equipment. Background Technology
[0002] The GTO (Greedy-Then-Oldest) scheduling strategy is a strategy that prioritizes the execution of the highest-priority thread bundle until that thread bundle becomes unexecutable (e.g., due to data dependencies). Only then will the oldest thread bundle (or the highest-priority thread bundle) be selected for execution. The age of each thread bundle is determined by the time it was scheduled to the processor.
[0003] In related technologies, if the current thread bundle cannot be executed, the oldest thread bundle will be selected directly. If multiple thread bundles have the same age, the selection will be based on the thread bundle index (Warp Index, wid), such as selecting the thread bundle with the smallest or largest thread bundle index to continue execution.
[0004] If the above thread selection method is adopted, and there is a data cache conflict between the oldest thread and the highest priority thread, the execution of the oldest thread may replace the cached information of the highest priority thread in the data cache, which will affect the subsequent execution of the highest priority thread and thus affect the processor's execution efficiency. Summary of the Invention
[0005] This application provides a scheduler, a scheduling method, a scheduling device, and an equipment, the technical solutions of which are as follows:
[0006] According to one aspect of this application, a scheduler is provided, the scheduler including a scheduling unit and an update unit;
[0007] The scheduling unit is configured to obtain the reuse degree of each of at least two first thread bundles, wherein the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; and based on the reuse degree of each of the at least two first thread bundles, to issue a target thread bundle to the pipeline, wherein the target thread bundle is one of the at least two first thread bundles.
[0008] The update unit is used to update the reuse rate of the target thread bundle when the target thread bundle accesses the data cache.
[0009] According to one aspect of this application, a scheduling method is provided, the scheduling method being executed by a scheduler, the scheduler including a scheduling unit and an update unit;
[0010] The method includes:
[0011] The scheduling unit obtains the reuse degree of each of the at least two first thread bundles, where the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, a target thread bundle is emitted to the pipeline, where the target thread bundle is one of the at least two first thread bundles;
[0012] The update unit updates the reuse rate of the target thread bundle when the target thread bundle accesses the data cache.
[0013] According to one aspect of this application, a scheduling device is provided, the scheduling device comprising a scheduling module and an update module; the device includes:
[0014] The scheduling module obtains the reuse degree of each of the at least two first thread bundles, where the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, a target thread bundle is emitted to the pipeline, where the target thread bundle is one of the at least two first thread bundles;
[0015] The update module updates the reusability of the target thread bundle when the target thread bundle accesses the data cache.
[0016] According to one aspect of this application, a processor is provided, the processor including a scheduler.
[0017] According to one aspect of this application, a graphics card is provided, the graphics card including a scheduler.
[0018] According to one aspect of this application, a computer device is provided, the computer device including a scheduler.
[0019] The beneficial effects of the technical solution provided in this application include at least the following:
[0020] When scheduling the first thread bundle, the target thread bundle is determined based on its reuse rate. The reuse rate of the target thread bundle is updated when it accesses the data cache. The reuse rate of the first thread bundle indicates the number of times it uses data from the second thread bundle in the data cache. By selecting the target thread bundle to execute when the second thread bundle cannot, a first thread bundle with higher reuse rates can be chosen. This avoids the situation where cached data of the second thread bundle is affected by cached data of the first thread bundle during its execution, leading to cached data of the second thread bundle being replaced by the target thread bundle when the second thread bundle resumes execution. This would result in cached data of the second thread bundle being replaced by the target thread bundle, causing a cached data gap and requiring the second thread bundle to read the required data from the lower-level cache again, thus reducing its execution efficiency. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A schematic diagram of a GTO scheduler is shown;
[0023] Figure 2 A schematic diagram of the scheduling process of the GTO scheduler is shown;
[0024] Figure 3 This invention provides a schematic diagram of the structure of a scheduler according to an exemplary embodiment of the present application.
[0025] Figure 4 This invention provides a schematic diagram of the structure of a scheduler according to an exemplary embodiment of the present application.
[0026] Figure 5 This invention provides a schematic diagram of the structure of a scheduler according to an exemplary embodiment of the present application.
[0027] Figure 6 This invention provides a schematic diagram of the structure of a scheduler according to an exemplary embodiment of the present application.
[0028] Figure 7 This invention provides a schematic diagram of the structure of a scheduler according to an exemplary embodiment of the present application.
[0029] Figure 8 A schematic diagram of a scheduling process provided in an exemplary embodiment of this application is shown;
[0030] Figure 9 A schematic diagram of a scheduling process provided in an exemplary embodiment of this application is shown;
[0031] Figure 10 A schematic diagram of a scheduling process provided in an exemplary embodiment of this application is shown;
[0032] Figure 11 A schematic diagram of a scheduling process provided in an exemplary embodiment of this application is shown;
[0033] Figure 12 A flowchart of a scheduling method provided in an exemplary embodiment of this application is shown;
[0034] Figure 13 A structural block diagram of a scheduling apparatus provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0036] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0037] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0038] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the settings and operation information involved in this application were obtained with full authorization.
[0039] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0040] First, let me introduce the relevant terms used in this application:
[0041] GTO: A scheduling strategy that prioritizes the execution of the highest-priority thread bundle until that thread bundle becomes unexecutable (e.g., due to data dependencies), at which point the oldest thread bundle (or the highest-priority thread bundle) is selected for execution. Typically, the age of each thread bundle is determined by the time it was scheduled to the processor.
[0042] A thread warp (or warp for short) is a collection of threads. A thread warp can be understood as a basic execution unit in a processor, or a parallel execution unit within a processor. Multiple threads within a thread warp execute synchronously, or in other words, execute in parallel. When executing, multiple threads within a thread warp use the same or different computing resources to execute the same instruction. Generally, each thread warp includes 32 threads, but in other embodiments, depending on the processor architecture used, the number of threads in a thread warp can be more or less; this application does not limit this.
[0043] Pipeline: A collection of processing units used to execute a sequence of instructions, consisting of multiple processing units arranged in sequence. These processing units are responsible for different tasks in the instruction execution process, such as instruction fetch (IF), instruction decode (ID), execution (EX), memory access (MEM), and write back (WB).
[0044] Cache: A small-capacity, high-speed memory, part of the storage system, storing frequently used instructions and data. It can also be called cache memory. Processors typically include multiple levels of cache, each with a different capacity. Smaller caches generally have higher read / write efficiency because they process less data. Caches can be categorized into instruction caches and data caches based on the different types of information they store. Instruction caches store instructions, while data caches store data. This application primarily uses a data cache as an example, but it can also be used to support other types of caches for storing data; this application is not limiting in this regard.
[0045] Data caching: A type of cache used to store data.
[0046] GTO schedulers employing the GTO scheduling strategy, such as Figure 1 As shown, the GTO scheduler is used to select one thread bundle from multiple thread bundles and issue it to the processor for execution. Figure 1 The GTO scheduler in the system needs to select one thread bundle to fire from thread bundle 0, thread bundle 1, and thread bundle 2, and it will prioritize the highest priority thread bundle (i.e., thread bundle 0). Specifically, as follows... Figure 2 As shown, after the GTO scheduler issues thread bundle 0, thread bundle 0 is executed. If, during the execution of thread bundle 0, it needs to be temporarily suspended due to data dependencies or other issues, the GTO scheduler needs to schedule other low-priority thread bundles for execution, such as... Figure 2 As shown, while thread bundle 0 is suspended, thread bundle 1 and thread bundle 2 are emitted and executed respectively.
[0047] In this GTO scheduler in related technologies, if there is a serious data cache conflict between the selected low-priority thread bundle and the highest-priority thread bundle when the selected thread bundle is executed, that is, the selected thread bundle will use the data that the highest-priority thread bundle needs to access to replace the data cached in the data cache of the highest-priority thread bundle, then during the subsequent execution of the highest-priority thread bundle, a serious cache miss phenomenon will occur when the highest-priority thread bundle reads data. That is, the data cache is missing the data required by the thread bundle. At this time, it is necessary to read the data again from memory or lower-level cache, which greatly increases the data loading latency and reduces the processing efficiency of the highest-priority thread bundle.
[0048] Based on this, this application provides a scheduler that, when selecting other thread bundles for execution, will select based on the data cache conflict situation, giving priority to thread bundles that have data reuse with the highest priority thread bundle, reducing the probability of data cache conflicts, improving the execution efficiency of the highest priority thread bundle, and thus improving the utilization of the processor.
[0049] Figure 3 A schematic diagram of the structure of a scheduler 100 provided in an exemplary embodiment of this application is shown. The scheduler 100 includes a scheduling unit 110 and an update unit 120.
[0050] The scheduling unit 110 is used to obtain the reuse degree of each of the at least two first thread bundles, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, the target thread bundle is emitted to the pipeline, the target thread bundle is one of the at least two first thread bundles.
[0051] Optionally, the first thread bundle is a ready thread bundle; or, the first thread bundle is a thread bundle in a ready state; or, all threads in the first thread bundle are in a ready state.
[0052] Optionally, the second thread bundle is a blocked thread bundle; or, the second thread bundle is an interrupted thread bundle; or, the second thread bundle is a thread bundle in a blocked state; or, all or some of the threads in the second thread bundle are in a blocked state.
[0053] Optionally, the second thread bundle has a higher priority than any of the at least two first thread bundles; in other words, the second thread bundle is the highest priority thread bundle that the scheduler can schedule. The second thread bundle can also be called the highest priority thread bundle, and correspondingly, each of the at least two first thread bundles can be called a low priority thread bundle.
[0054] Optionally, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data used by the first thread bundle and data used by the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data used by the first thread bundle and data used by the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between data used by the first thread bundle and data used by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between data read into the data cache by the first thread bundle and data read into the data cache by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between data cached in the data cache by the first thread bundle and data used by the second thread bundle. The degree of reuse between the second thread bundle cached and the data in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between the data used by the first thread bundle and the data used by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between the data used by the first thread bundle and the data of the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the number of times a cache miss of the first thread bundle is resolved based on the data of the second thread bundle in the data cache (or cache line, cache, cache space, etc.); or, the reuse degree of the first thread bundle is used to indicate the number of times the first thread bundle hits the data of the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the difference between the number of times the first thread bundle uses the data of the second thread bundle in the data cache and the number of times it replaces the data of the second thread bundle in the data cache.
[0055] Optionally, the reuse rate of the first thread bundle is an estimated value. In other words, the reuse rate of the first thread bundle is calculated in real time based on whether a cache miss occurs after the first thread bundle is executed, and whether the cache miss causes the data of the second thread bundle in the data cache to be replaced. It can also be said that the reuse rate of the first thread bundle is not a precise value that can accurately describe the degree of data reuse between the first thread bundle and the second thread bundle, but rather describes the number of times the first thread bundle has hit the data of the second thread bundle in the data cache, that is, the number of times the first thread bundle has used the data of the second thread bundle in the data cache.
[0056] Optionally, the scheduling unit 110 selects a target thread bundle and issues the target thread bundle to the pipeline based on the reuse degree of each of the at least two first thread bundles. The target thread bundle may be the first thread bundle with the highest reuse degree among the at least two first thread bundles, or any first thread bundle among the at least two first thread bundles that meets the reuse degree condition, or the first thread bundle among the at least two first thread bundles that meets both the reuse degree and priority conditions, and so on.
[0057] Optionally, the priority of the first thread bundle is determined based on the age of the first thread bundle, such that the older the first thread bundle is, the higher its priority; or, the priority of the first thread bundle is determined based on the thread bundle index (wid) of the first thread bundle, such that the smaller the thread bundle index, the higher its priority, and so on.
[0058] Update unit 120 is used to update the reusability of the target thread bundle when the target thread bundle accesses the data cache.
[0059] Optionally, the update unit 120 updates the reuse of the target thread bundle when the target thread bundle accesses the data cache; or, the update unit 120 updates the reuse of the target thread bundle when a memory access instruction in the target thread bundle is executed.
[0060] In summary, the scheduler provided in this application determines the target thread bundle based on the reuse degree of the first thread bundle when scheduling the first thread bundle. It also updates the reuse degree of the target thread bundle when the target thread bundle accesses the data cache. The reuse degree of the first thread bundle indicates the number of times the first thread bundle uses data from the second thread bundle in the data cache. By selecting the target thread bundle to execute when the second thread bundle cannot execute, the scheduler can select a first thread bundle with a higher reuse degree than the second thread bundle. This avoids the situation where a cached data of the first thread bundle is missing during its execution, affecting the data cached by the second thread bundle in the data cache. This prevents the second thread bundle from resuming execution due to its cached data being replaced by the target thread bundle, leading to a cached data gap and requiring it to read the required data from the lower-level cache again, thus reducing its execution efficiency.
[0061] In other embodiments, the scheduler 100 is connected to the pipeline 200 and the data buffer 300. For example... Figure 4The schematic diagram of the processor structure is shown. The scheduling unit 110 in the scheduler 100 is connected to the pipeline 200, and the update unit 120 in the scheduler 100 is connected to the data cache 300. The pipeline 200 is used to execute a target thread bundle, or, execute a second thread bundle. When executing a memory access instruction for a target thread bundle, a memory access request is sent to the data cache 300. The memory access request includes at least one of the target thread bundle index (wid) and the memory access address of the target thread bundle, where the memory access address is the memory address of the data requested by the target thread bundle. The memory access address can be used to determine a cache tag, which can be used to quickly determine whether the data corresponding to the memory access address is in the data cache in the tag array (tag_array). Typically, the cache tag is the high-order bit of the memory access address, such as the high t bits of the memory access address, i.e., t bits from left to right. Optionally, the first bit from left to right of the memory access address can also be used as a valid bit, in which case the cache tag is from the second bit to the (t+1)th bit from left to right. Data cache 300 determines whether a memory access request has been hit based on the received request. Specifically, it checks if the data corresponding to the memory access request exists in data cache 300. If it does, the request is hit, and the data can be directly read for calculation. If it doesn't exist, the request is miss. In the case of a miss, it also checks if data cache 300 is full. If data cache 300 is not full, data is read directly from the lower-level cache (or memory) into a free cache line. If data cache 300 is full, a cache line needs to be selected from the data cache for replacement. This can be done using the Least Recently Used (LRU) algorithm, but other algorithms can also be used. This embodiment of the application does not limit the method of selecting a cache line. After executing the memory access instruction, data cache 300 generates relevant access information and informs update unit 120 of this information, so that update unit 120 updates the reuse rate of the target thread bundle based on the access information.
[0062] It should be noted that the structural schematic diagrams of the scheduler or processor shown in the embodiments of this application only show some of the connection relationships between the various units or devices. However, in specific implementations, there may be other connection relationships not shown in the embodiments of this application for the various units or devices shown in the embodiments of this application, or other units or other devices may be added to the connection lines shown in the embodiments of this application, etc. However, the protection scope of the embodiments of this application is not limited to this.
[0063] 1. Scheduling Unit
[0064] The following further illustrates how the scheduling unit 110 selects the target thread bundle.
[0065] In some embodiments, the scheduling unit 110 is configured to determine at least one first thread bundle among at least two first thread bundles that satisfies a first condition based on the reuse degree of each of the at least two first thread bundles, the first condition including the reuse degree of the first thread bundle being greater than or equal to a reuse degree threshold; and to send a target thread bundle to the pipeline, the target thread bundle being one of the at least one first thread bundles that satisfies the first condition.
[0066] Optionally, the scheduling unit 110 selects at least one first thread bundle that satisfies the first condition based on the reuse degree of each of the at least two first thread bundles. And it issues any one of the at least two first thread bundles as the target thread bundle to the pipeline.
[0067] Optionally, if at least two first thread bundles do not include a first thread bundle that satisfies the first condition, the scheduling unit 110 sends a prompt message to the scheduler 100 (or update unit 120) indicating that there is no first thread bundle that satisfies the first condition, so that the scheduler 100 (or update unit 120) can reset or initialize the reuse of all or part of the first thread bundles; or, the scheduling unit 110 resets or initializes the reuse of all or part of the first thread bundles. In other embodiments, if the first thread bundle satisfying the first condition is not included among the at least two second thread bundles, the scheduling unit 110 or the scheduler 100 may adopt other solutions, such as directly selecting any first thread bundle as the target thread bundle from the at least two first thread bundles, or selecting the oldest first thread bundle as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the smallest thread bundle index (wid) as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest priority as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest reuse rate as the target thread bundle from the at least two first thread bundles, or reducing the reuse rate threshold, or notifying the update unit 120 to reset the reuse rate of some first thread bundles, etc. In addition to the methods shown in the embodiments of this application, other methods can be used to solve the above-mentioned special cases of abnormal situations, combined with actual hardware design. The embodiments of this application do not limit this, but the protection scope of the embodiments of this application is not limited thereto.
[0068] Optionally, the reuse threshold can be set by the scheduler designer before shipment, or it can be modified by the user using an interface provided by the scheduler designer. That is, the reuse threshold can be changed according to actual usage, and different reuse thresholds can be set for different schedulers.
[0069] In summary, the scheduler provided in this application selects a target thread bundle from the first thread bundles that meet the first condition when selecting a target thread bundle. By comparing the relationship between the reuse degree of each first thread bundle and the reuse degree threshold, the scheduler determines the first thread bundle with a reuse degree greater than the reuse degree threshold. This ensures that the final selected target thread bundle has a high degree of reuse with the second thread bundle, thus reducing the probability that the determined target thread bundle can avoid cache misses and replace the corresponding data of the second thread bundle in the data cache. Furthermore, by first determining at least one first thread bundle with a reuse degree greater than the reuse degree threshold, rather than directly selecting the first thread bundle with the highest reuse degree as the target thread bundle, the scheduler ensures that the first thread bundle that meets the first condition among at least two first thread bundles has a probability of being selected for execution. This minimizes the possibility of repeatedly issuing a first thread bundle, causing other first thread bundles to have no opportunity to execute and leaving most first thread bundles in a long-term starved state, which is detrimental to processor load balancing.
[0070] The aforementioned scheduling unit 110 has two functions: condition judgment and scheduling. Considering the principle of functionality, i.e., the circuit design process is complex, it needs to be divided according to the functions of each component. The scheduling unit 110 can be divided into a multiplexing comparison subunit 111 and a scheduling subunit 112, as follows: Figure 5 As shown. That is, the scheduling unit 110 includes a reuse degree comparison subunit 111 and a scheduling subunit 112.
[0071] Optionally, the reuse comparison subunit 111 is used to obtain the i-th first thread bundle among at least two first thread bundles, where i is a positive integer; if the reuse of the i-th first thread bundle is greater than or equal to the reuse threshold, determine that the i-th first thread bundle is a first thread bundle that satisfies the first condition; let i = i + 1, and repeat the process starting from obtaining the i-th first thread bundle until at least two first thread bundles have been judged; send at least one first thread bundle that satisfies the first condition to the scheduling subunit 112. Here, the initial value of i is 1, and when at least two first thread bundles are n first thread bundles, i is less than or equal to n, where n is a positive integer.
[0072] Optionally, if the reuse degree comparison subunit 111 determines the i-th first thread bundle as a first thread bundle satisfying the first condition when the reuse degree of the i-th first thread bundle is greater than or equal to the reuse degree threshold, for example, the i-th first thread bundle is set to valid. If the reuse degree of the i-th first thread bundle is less than the reuse degree threshold, the i-th first thread bundle is determined to be a first thread bundle that does not satisfy the first condition, for example, the i-th first thread bundle is set to invalid.
[0073] For example, if at least two first thread bundles constitute n first thread bundles, then an n-bit information can be used, which can be represented by a register, other unit, or other device. The multiplexing comparison subunit 111 is responsible for determining whether at least two thread bundles meet the first condition and updating the n-bit information. For example, each bit in the n-bit information corresponds to a first thread bundle. When a bit is 0, it indicates that the first thread bundle corresponding to that bit is invalid, or does not meet the first condition; when a bit is 1, it indicates that the first thread bundle corresponding to that bit is valid, or meets the first condition. Of course, the specific meaning of each bit can also be reversed, i.e., a bit of 0 indicates that the first thread bundle corresponding to that bit is valid, a bit of 1 indicates that the first thread bundle corresponding to that bit is invalid, and so on. This application embodiment does not limit this.
[0074] Optionally, the reuse degree comparison subunit 111 sends at least one first thread bundle that satisfies the first condition to the scheduling subunit 112. This can be implemented by the reuse degree comparison subunit 111 sending n bits of information to the scheduling subunit, which is used by the scheduling subunit 112 to determine the first thread bundle that is allowed to be scheduled; or, the reuse degree comparison subunit 111 sends the thread bundle index (wid) of each of the at least one first thread bundle that satisfies the first condition to the scheduling subunit 112; or, the reuse degree comparison subunit 111 sends n bits of information and the thread bundle index of each of the at least two first thread bundles to the scheduling subunit 112; or, the reuse degree comparison subunit 111 sends n bits of information and at least two first thread bundles to the scheduling subunit 112.
[0075] Optionally, the scheduling subunit 112 is used to send the target thread bundle to the pipeline.
[0076] Optionally, the scheduling subunit 112 selects a target thread bundle from at least one first thread bundle that satisfies the first condition. For example, it may directly select any first thread bundle as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the oldest first thread bundle as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the first thread bundle with the smallest thread bundle index as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the first thread bundle with the highest priority as the target thread bundle from at least one first thread bundle that satisfies the first condition.
[0077] Optionally, the scheduling subunit 112 sends the target thread bundle to the pipeline; or, the scheduling subunit 112 sends the thread bundle index of the target thread bundle to the pipeline.
[0078] In summary, the scheduler provided in this application illustrates the functional division of the scheduling unit, as well as the specific implementation process of the reuse comparison subunit and the scheduling subunit within the scheduling unit. Functional division facilitates the relatively independent design of each functional unit, making the design process clearer, reducing the possibility of design errors, and helping to reduce mutual interference between different functional modules. Furthermore, the implementation process of the reuse comparison subunit and the scheduling subunit is also shown. The reuse comparison subunit compares the reuse of the first thread bundle with a reuse threshold to block or allow the first thread bundle. Dynamic adjustments can be made for different implementation scenarios by adjusting the reuse threshold. For example, when the number of at least one first thread bundle satisfying the first condition is small, the reuse threshold can be lowered to allow more first thread bundles to run, thereby avoiding repeated execution of some first thread bundles and improving the processor's load balancing performance. It also ensures that the allowed first thread bundle has a low probability of replacing data of the second thread bundle in the data cache, reducing the probability of cache misses in the first and second thread bundles and improving the overall execution efficiency of the first and second thread bundles.
[0079] The following further illustrates the scheduling method for at least one first thread bundle that satisfies the first condition, i.e., how to select the target thread bundle from at least one first thread bundle that satisfies the first condition.
[0080] In some embodiments, such as Figure 6 As shown, the scheduling subunit 112 includes a polling arbitrator 1121 and a strict priority arbitrator 1122.
[0081] Optionally, the polling arbitrator 1121 is used to select a target thread bundle from at least one first thread bundle that satisfies the first condition.
[0082] Optionally, the polling arbitrator 1121 is used to select a target thread bundle from at least one first thread bundle that satisfies the first condition, according to a polling scheduling strategy.
[0083] In some embodiments, the polling scheduling strategy is as follows: at least two first thread bundles are n first thread bundles, where n is a positive integer; a polling arbiter 1121 is used to obtain a polling flag, which indicates that the first thread bundle previously selected by the polling arbiter 1121 is the j-th first thread bundle among at least two first thread bundles, where j is a positive integer; if at least one first thread bundle satisfying the first condition includes the k-th first thread bundle, the k-th first thread bundle is selected as the target thread bundle, and the polling flag is updated to indicate the k-th first thread bundle, where k is a positive integer and the initial value of k is j+1; if at least one first thread bundle satisfying the first condition does not include the k-th first thread bundle, k = k+1; if k is greater than n, k = kn.
[0084] Optionally, the polling arbitrator 1121 selects the next first thread bundle pointed to by the polling flag from at least one first thread bundle that satisfies the first condition, based on the polling flag, as the target thread bundle. The target thread bundle is not necessarily the (j+1)th first thread bundle. In other words, starting from the (j+1)th first thread bundle, it determines whether it satisfies the first condition (or whether it is valid, or whether it is not masked), and selects the first first thread bundle that satisfies the first condition as the target thread bundle.
[0085] Optionally, if there is no first thread bundle that satisfies the first condition, or if k = j, a prompt message is sent to indicate to scheduler 100 (or update unit 120) that there is no first thread bundle that satisfies the first condition, so that scheduler 100 (or update unit 120) can reset or initialize the reuse of all or part of the first thread bundle; or, scheduler unit 110 resets or initializes the reuse of all or part of the first thread bundle. In other embodiments, if the first thread bundle satisfying the first condition is not included among the at least two second thread bundles, the polling arbitrator 1121 or scheduler 100 may also adopt other solutions, such as directly selecting any first thread bundle (e.g., the j-th or j+1-th first thread bundle) as the target thread bundle from the at least two first thread bundles, or selecting the oldest first thread bundle as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the smallest thread bundle index (wid) as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest priority as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest reuse rate as the target thread bundle from the at least two first thread bundles, or reducing the reuse rate threshold, etc. In addition to the methods shown in the embodiments of this application, other methods can be used to solve the above-mentioned special cases of abnormal situations, combined with actual hardware design. The embodiments of this application do not limit this to such cases, but the scope of protection of the embodiments of this application is not limited thereto.
[0086] Optionally, in the initial case, the polling flag is used to indicate the first first thread bundle, or the initial polling flag is used to indicate the first first thread bundle, or the initial polling flag is used to indicate any one of at least two first thread bundles, or the initial polling flag is used to indicate the last first thread bundle, or the initial polling flag is null.
[0087] Optionally, if the polling flag is null, determine whether the first thread bundle satisfies the first condition. If a bit of 1 indicates that the corresponding first thread bundle is valid, then determine whether the first bit in the n-bit information is 1. If the first thread bundle does not satisfy the first condition, let j = j + 1; if the first thread bundle satisfies the first condition, determine the first thread bundle as the target thread bundle and update the polling flag to indicate the first thread bundle.
[0088] Optionally, the determination of whether the kth first thread bundle satisfies the first condition, as shown above, employs different determination methods depending on the information exchange method between the reuse comparison subunit 111 and the polling arbitrator 1121. For example, if the reuse comparison subunit 111 sends n bits of information to the scheduling subunit 112, which is used by the scheduling subunit 112 to determine the first thread bundles that can be scheduled, then determining whether the kth first thread bundle satisfies the first condition can be achieved by determining whether the kth bit in the n bits of information is 1; or, if the reuse comparison subunit 111 sends the thread bundle index of each of the at least one first thread bundles that satisfies the first condition to the scheduling subunit 112, then determining whether the kth first thread bundle satisfies the first condition can be achieved by determining whether the thread bundle index of the kth first thread bundle has been received (this thread bundle index is read by the polling arbitrator 1121 itself); or, the reuse comparison... Subunit 111 sends n bits of information and the thread bundle index of each of the at least two first thread bundles to scheduling subunit 112. Then, it can be determined whether the kth first thread bundle satisfies the first condition by judging whether the kth bit in the n bits of information is 1, or by judging whether the thread bundle index of the kth first thread bundle has been received (the thread bundle index is read by the polling arbiter 1121 itself); or, if the reuse degree comparison subunit 111 sends n bits of information and at least two first thread bundles to scheduling subunit 112, it can be determined whether the kth first thread bundle satisfies the first condition by judging whether the kth bit in the n bits of information is 1.
[0089] For example, n is 4, meaning at least two first thread bundles include 4 first thread bundles. Assume the initial polling flag is null. If, in the first poll, the first n bits of information received are 0011, then starting from the first bit, the process is as follows: the first bit indicates that the first first thread bundle does not meet the first condition, so let k = k + 1 = 2; the second bit indicates that the second first thread bundle does not meet the first condition, so let k = k + 1 = 3; the third bit indicates that the third first thread bundle meets the first condition, thus determining the third first thread bundle as the target thread bundle, and updating the polling flag so that it indicates the third first thread bundle. For example, if the polling flag is updated to 3, the current polling ends. In the second polling, assuming the received n bits of information are 0110, according to the polling flag, j is 3, so the initial value of k is 3+1=4. We determine to start judging from the 4th bit. The 4th bit indicates that the 4th first thread bundle does not meet the first condition, so let k=k+1=5. Since k=5>n=4, let k=kn=5-4=1. Continuing to judge from the 1st bit, the 1st bit indicates that the 1st first thread bundle does not meet the first condition, so let k=k+1=2. Judging from the 2nd bit, the 2nd bit indicates that the 2nd first thread bundle meets the first condition. We determine the 2nd first thread bundle as the target thread bundle and update the polling flag to indicate the 2nd first thread bundle, such as updating the polling flag to 2.
[0090] In other embodiments, a polling flag is used to indicate that the first thread bundle previously selected by the polling arbiter is the j-th first thread bundle among at least one first thread bundle that satisfies the first condition. If there are p first thread bundles that satisfy the first condition this time, if j+1 is greater than p, the first first thread bundle among the first thread bundles that satisfy the first condition this time is selected as the target thread bundle, and the polling flag is updated to 1; if j+1 is less than or equal to p, the (j+1)-th first thread bundle among the first thread bundles that satisfy the first condition this time is selected as the target thread bundle, and the polling flag is updated to j+1. That is, polling for at least two first thread bundles is no longer strictly performed, but the next first thread bundle among at least one first thread bundle that satisfies the first condition is selected according to the polling flag.
[0091] For example, there are at least two first thread bundles, totaling eight first thread bundles, denoted as first thread bundle 0, first thread bundle 1, first thread bundle 2, ..., first thread bundle 7. When a first thread bundle needs to be scheduled for the i-th time, the first thread bundles that satisfy the first condition are {first thread bundle 0, first thread bundle 3, first thread bundle 5, first thread bundle 6}, that is, the number of first thread bundles that satisfy the first condition at the i-th time is p = 3. If the polling flag at this time is 2, k = 2 + 1 = 3 < p, therefore the third first thread bundle, i.e., first thread bundle 5, is determined as the target thread bundle, and the polling flag is updated to 3. When the first thread bundle needs to be scheduled for the (i+1)th time, the first thread bundles that meet the first condition are {first thread bundle 0, first thread bundle 3, first thread bundle 6}. That is, the number of first thread bundles that meet the first condition at the (i+1)th time is p = 3. If the polling mark at this time is 3, k = 3 + 1 = 4 > p, then start selecting from the beginning, select the first first thread bundle, i.e., first thread bundle 0, as the target thread bundle, and update the polling mark to 1.
[0092] Optionally, in the above embodiments, when j+1 is greater than p, the first first thread bundle that satisfies the first condition is selected as the target thread bundle, and the polling flag is updated to 1. Alternatively, when j+1 is greater than p, let k = j+1-p, select the kth first thread bundle that satisfies the first condition as the target thread bundle, and update the polling flag to k.
[0093] Optionally, the strict priority arbiter 1122 is used to emit the target thread bundle if the second thread bundle satisfies a second condition, the second condition including that the second thread bundle is suspended.
[0094] Optionally, the strict priority arbiter 1122 fires the second thread bundle if the second thread does not meet the second condition. That is, if the second thread bundle can execute normally, it is fired first, i.e., the highest priority thread bundle. If the second thread bundle cannot execute normally, the target thread bundle is fired, i.e., the first thread bundle selected from at least two first thread bundles, i.e., the low priority thread bundle.
[0095] Optionally, the second condition includes at least one of the following: the second thread bundle is suspended; the second thread bundle is interrupted; the second thread bundle is blocked; the second thread bundle is in a blocked state; or all or some of the threads in the second thread bundle are in a blocked state.
[0096] In summary, the scheduler provided in this application further illustrates the method by which the scheduling subunit implements thread bundle scheduling. Instead of directly selecting the most extreme first thread bundle based on priority, reuse, or age (e.g., selecting the highest priority, highest reuse, or oldest first thread bundle), the target thread bundle is selected from at least one first thread bundle that meets the first condition through round-robin scheduling. This ensures fair scheduling of the first thread bundles that meet the first condition, better guaranteeing processor load balancing. Furthermore, the basis of round-robin scheduling is that the first thread bundle being round-robin meets the first condition, i.e., it has a high reuse rate with the second thread bundle. This ensures that the target thread bundle obtained through round-robin scheduling also has a low probability of cache misses, reducing the probability of the second thread bundle experiencing a cache miss due to the execution of the first thread bundle, thus guaranteeing the execution efficiency of both the first and second thread bundles. Moreover, by selecting the first thread bundle that meets the first condition as the target thread bundle, it is also possible to ensure that data read from the data cache is reused as much as possible during the execution of the target thread bundle and the second thread bundle, improving the data reuse rate in the data cache.
[0097] 2. Reusability
[0098] To reduce the delay in the scheduler reading these two parameters, the reuse degree and reuse threshold can be stored using registers.
[0099] In some embodiments, the scheduling unit 110 is configured to read a reuse threshold from a reuse threshold register.
[0100] In some embodiments, the scheduling unit 110 is configured to read the reuse degree of each first thread bundle from the reuse degree register corresponding to each of the at least two first thread bundles.
[0101] In addition, when each target thread bundle is executed, the reusability of the target thread bundle needs to be updated in a timely manner based on the execution status of the target thread bundle, as shown below.
[0102] (1) Reusability update
[0103] In some embodiments, the update unit 120 is configured to increase the reuse rate of the target thread bundle when the target thread bundle accesses data corresponding to the second thread bundle in the data cache. Alternatively, the reuse rate of the target thread bundle is not updated when the target thread bundle accesses data corresponding to the second thread bundle in the data cache.
[0104] Optionally, the update unit 120 increases the reusability of the target thread bundle when the target thread bundle accesses the data corresponding to the second thread bundle in the data cache; or, the update unit 120 increases the reusability of the target thread bundle when the target thread bundle hits the data corresponding to the second thread bundle in the data cache; or, the update unit 120 increases the reusability of the target thread bundle when the target thread bundle hits the cache line corresponding to the second thread bundle in the data cache; or, the update unit 120 increases the reusability of the target thread bundle when the target thread bundle accesses the cache line corresponding to the second thread bundle in the data cache.
[0105] Optionally, the step size for increasing reusability is the first step size, such as a first step size of 1.
[0106] Optionally, the update unit 120 determines whether the target thread bundle has accessed the data corresponding to the second thread bundle based on the access information sent by the data cache. Specifically, the access information includes at least one of the following: the thread bundle index of the target thread bundle, the access status, and the thread bundle index corresponding to the accessed cache line. The access status indicates whether the access was successful; if the access status is successful and the thread bundle index corresponding to the accessed cache line is the thread bundle index of the second thread bundle, it is determined that the target thread bundle accessed the data corresponding to the second thread bundle in the data cache, increasing the reusability of the target thread bundle. If the access status is unsuccessful and the thread bundle index corresponding to the accessed cache line is the thread bundle index of the second thread bundle, it means that the target thread bundle replaced the data corresponding to the second thread bundle in the data cache.
[0107] In some embodiments, the update unit 120 is configured to reduce the reuse rate of the target thread bundle when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache. Alternatively, when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache, the reuse rate of the target thread bundle is not updated.
[0108] Optionally, the step size for reducing reusability is the second step size, such as 1.
[0109] Optionally, the step size for increasing reuse and the step size for decreasing reuse can be the same or different. For example, the step size for increasing reuse can be 1, and the step size for decreasing reuse can be 2, and so on. Developers can adjust the first and second step sizes according to specific strategies. For instance, in scenarios where it is necessary to strictly avoid cache misses in the second thread bundle, the second step size can be set to be greater than the first step size, that is, a more severe penalty is applied to the first thread bundle that causes a cache miss in the second thread bundle.
[0110] Optionally, the first step length and the second step length can be determined based on the number of at least one first thread bundles that satisfy the first condition (e.g., denoted as the first quantity). For example, if the first quantity is greater than the first threshold, the second step length is set to be greater than the first step length, such as a second step length of 2 and a first step length of 1. In this case, since the number of at least one first thread bundles satisfying the first condition is relatively large, a larger second step length can be used. If the first quantity is less than the second threshold, the first step length is set to be greater than or equal to the second step length, such as both the first step length and the second step length being 1. In this case, the number of at least one first thread bundles satisfying the first condition is relatively small. To avoid unfair thread scheduling caused by a small number of at least one first thread bundles satisfying the first condition, the second step length can be reduced, or a larger first step length can be used.
[0111] In summary, the scheduler provided in this application embodiment illustrates a scenario where reuse is updated. When the target thread bundle accesses data of the second thread bundle in the data cache, the reuse of the target thread bundle can be increased for further recording; when the target thread bundle replaces data of the second thread bundle in the data cache, the reuse of the target thread bundle can be decreased. For the above two update scenarios, only one can be used, or both can be used. That is, when the target thread bundle accesses data of the second thread bundle in the data cache, the reuse of the target thread bundle is increased; when the target thread bundle replaces data of the second thread bundle in the data cache, the reuse of the target thread bundle is not updated. Alternatively, when the target thread bundle accesses data of the second thread bundle in the data cache, the reuse of the target thread bundle is not updated; when the target thread bundle replaces data of the second thread bundle in the data cache, the reuse of the target thread bundle is decreased. Alternatively, when the target thread accesses data in the data cache of the second thread, the reuse of the target thread is increased; when the target thread replaces data in the data cache of the second thread, the reuse of the target thread is decreased. By updating the reuse of the target thread in real time during its execution, the timeliness of the reuse used when selecting the target thread is ensured. This guarantees the accuracy of the reuse level between the selected target thread and the second thread, thus ensuring a high probability of data reuse between them. It also reduces the possibility of the target thread replacing data in the data cache of the second thread due to cache misses, thereby ensuring the execution efficiency of both the target and second thread.
[0112] The above illustrates some scenarios of reuse reset. For example, if there is no first thread bundle that meets the first condition, reuse may need to be reset to ensure the scheduler functions correctly. The following sections further illustrate more reuse reset scenarios and the methods for resetting reuse.
[0113] (2) Reuse rate reset
[0114] In some embodiments, the update unit 120 is configured to reset the reuse rate of each of at least two first thread bundles when a third condition is met. Alternatively, it may reset the reuse rate of some of the first thread bundles when the third condition is met. Or, it may update the reuse rate of all or some of the first thread bundles when the third condition is met.
[0115] The third condition includes at least one of the following: the reset time reaches a reset time threshold, where the reset time is used to indicate the time between the current moment and the moment of the last reset of reuse; a first thread bundle is identified as a new second thread bundle; and the number of at least one first thread bundle that satisfies the first condition is less than a number threshold.
[0116] Optionally, resetting the reuse of the first thread bundle means updating the reuse of the first thread bundle to an initial value, such as 0. Updating the reuse of the first thread bundle can be done by increasing or decreasing the reuse of the first thread bundle by a third step size, which can be set by the developer or by the designer of the scheduler 100.
[0117] Optionally, the third condition includes the reset time reaching a reset time threshold, which can be set by the developer or by the designer of the scheduler 100. That is, the reuse rate is reset every time the reset time threshold is reached. Optionally, if the reset time reaches the reset time threshold but the pipeline is executing the target thread bundle, the reset is not performed until the pipeline executes the second thread bundle.
[0118] Optionally, the third condition includes the existence of a first thread being identified as a second thread bundle. That is, after a second thread bundle is executed, it is necessary to select the first thread bundle with the highest priority from at least two first thread bundles as the second thread bundle. At this time, it is necessary to update the reuse of all thread bundles in the remaining first thread bundles. This is because the old reuse is the reuse between the old and the old second thread bundles. After the new second thread bundle is identified, the reuse between the new and the new second thread bundles should be recorded again.
[0119] Optionally, the third condition includes the number of at least one first thread bundle satisfying the first condition being less than a quantity threshold, such as a quantity threshold of 1, meaning that first thread bundles satisfying the first condition are not included. The quantity threshold can also be other thresholds, and this application embodiment does not limit this. The determination that the number of at least one first thread bundle satisfying the first condition is less than the quantity threshold can also be made by the scheduling unit 110. When the scheduling unit 110 confirms that the number of at least one first thread bundle satisfying the first condition is less than the quantity threshold, it sends an indication message to the update unit 120. In this case, the third condition can be understood as including receiving the indication message sent by the scheduling unit 110. This indication message can be an electrical signal, such as the scheduling unit 110 indicating to the update unit 120 via a high-level signal that the reuse of at least two first thread bundles needs to be reset or updated.
[0120] For example, if the number of first thread bundles satisfying the first condition is less than a quantity threshold, the reuse of some of the first thread bundles among at least two first thread bundles is reset or updated, such as resetting or updating the reuse of the first thread bundles among at least two first thread bundles whose reuse is greater than a second threshold. Resetting or updating the reuse of some first thread bundles instead of resetting or updating the reuse of all first thread bundles can prevent first thread bundles with low reuse from being reset, which would cause the execution of the first thread bundle to replace the data of the second thread bundle in the data cache, resulting in a cache miss for the second thread bundle.
[0121] In summary, the scheduler provided in this application embodiment illustrates a reuse level reset scenario. Resetting the reuse level can, on the one hand, prevent the scheduler from functioning properly when there is no first thread bundle satisfying the first condition and the second thread bundle is blocked. On the other hand, resetting the reuse level can also prevent the target thread bundle from becoming fixed. For example, resetting the reuse level periodically or when the number of first thread bundles satisfying the first condition is too small can expand the range of selectable first thread bundles that satisfy the first condition, i.e., expand the range of selectable target thread bundles. This allows first thread bundles that originally did not satisfy the first condition to potentially satisfy it after the reuse level reset, thus making them possible to be selected for execution. This ensures the fairness of target thread bundle selection and guarantees the load balancing of the scheduler.
[0122] Specifically, Figure 7 A structural block diagram of a scheduling circuit provided in an exemplary embodiment of this application is shown. The scheduling circuit includes a scheduler 100, a pipeline 200, and a data buffer 300. The scheduler 100, pipeline 200, and data buffer 300 are respectively connected, and pipeline 200 and data buffer 300 are connected.
[0123] The scheduler 100 includes a reuse comparison subunit 111, a round-robin (RR) arbitrator 1121, a strict priority (SP) arbitrator 1122, and an update unit 120. The scheduler 100 is connected to the pipeline 200. The scheduler 100 selects thread bundles to be executed and sends them to the pipeline 200 for execution. The thread bundle may include memory access instructions, which trigger the pipeline 200 to read cached data from the data cache 300.
[0124] For this scheduler 100, there are at least two first thread bundles and a second thread bundle, and the second thread bundle has a higher priority than any of the at least two first thread bundles. That is, the second thread bundle is a high-priority warp (HPW), and the first thread bundle is a low-priority warp (LPW).
[0125] The second thread bundle is the highest priority thread bundle. During each execution, the strict priority arbiter 1122 prioritizes executing the second thread bundle. If the second thread bundle does not meet the execution conditions (e.g., data dependencies are not resolved), one of the at least two first thread bundles is selected for execution. The scheduler 100 maintains a reuse register for each of the at least two first thread bundles; that is, each of the at least two first thread bundles corresponds to a reuse degree. The reuse degree corresponding to each of the at least two first thread bundles is used to count the data cache conflict or reuse status of the current first thread bundle and the second thread bundle. Specifically, the reuse degree of the first thread bundle indicates the number of times a cache miss of the first thread bundle is resolved based on the data (or cache line, cache, cache space, etc.) of the second thread bundle in the data cache; or, the reuse degree of the first thread bundle indicates the number of times the first thread bundle hits the data of the second thread bundle in the data cache.
[0126] The reuse degree comparison subunit 111 in the scheduler 100 is used to mask first thread bundles that do not meet the first condition based on the reuse degree of each of the at least two first thread bundles. Alternatively, the reuse degree comparison subunit 111 is used to set the first thread bundles that meet the first condition to be valid, and to set the first thread bundles that do not meet the first condition to be invalid. That is, the reuse degree comparison subunit 111 is responsible for comparing and masking the reuse degree (or reuse degree register) of each first thread bundle. The comparison threshold comes from a threshold register, which can be preset through a software interface. The reuse degree comparison subunit 111 compares the reuse degree register value with the threshold register value; all first thread bundles greater than or equal to the threshold register value are set to valid, and the remaining first thread bundles are set to invalid. Valid first thread bundles can participate in subsequent polling arbitration. Optionally, the validity information of each first thread bundle among the at least two first thread bundles can also be stored in a register, such as using an n-bit register, with each bit corresponding to one first thread bundle.
[0127] Specifically, the first condition includes a reuse degree greater than or equal to a reuse degree threshold. The reuse degree threshold is read from the threshold register by the reuse degree comparison subunit 111.
[0128] For example, such as Figure 8 As shown, assuming the reuse degree of the first thread bundle 0 is -1, the reuse degree of the first thread bundle is 2, and the reuse degree of the first thread bundle 2 is 1, the reuse degree threshold is set to 0. The reuse degree comparison subunit 111 blocks all first thread bundles with reuse degrees less than the reuse degree threshold, and sets the remaining first thread bundles to be valid so that they can participate in the next round of polling arbitration.
[0129] For example, if at least two first thread bundles constitute n first thread bundles, then an n-bit information can be used, which can be represented by a register, other unit, or other device. The multiplexing comparison subunit 111 is responsible for determining whether at least two thread bundles meet the first condition and updating the n-bit information. For example, each bit in the n-bit information corresponds to a first thread bundle. When a bit is 0, it indicates that the first thread bundle corresponding to that bit is invalid, or does not meet the first condition; when a bit is 1, it indicates that the first thread bundle corresponding to that bit is valid, or meets the first condition. Of course, the specific meaning of each bit can also be reversed, i.e., a bit of 0 indicates that the first thread bundle corresponding to that bit is valid, a bit of 1 indicates that the first thread bundle corresponding to that bit is invalid, and so on. This application embodiment does not limit this.
[0130] Optionally, the reuse degree comparison subunit 111 sends valid first thread bundles to the polling arbitrator 1121, which then selects one first thread bundle as the target thread bundle and sends it to the priority strict arbitrator 1122. Alternatively, the reuse degree comparison subunit 111 sends at least two first thread bundles to the polling arbitrator 1121, which then selects one valid first thread bundle as the target thread bundle and sends it to the priority strict arbitrator 1122.
[0131] Optionally, the reuse degree comparison subunit 111 sends at least one first thread bundle that satisfies the first condition to the scheduling subunit 112. This can be implemented by the reuse degree comparison subunit 111 sending n bits of information to the scheduling subunit, which is used by the scheduling subunit 112 to determine the first thread bundle that is allowed to be scheduled; or, the reuse degree comparison subunit 111 sends the wid of each of the at least one first thread bundle that satisfies the first condition to the scheduling subunit 112; or, the reuse degree comparison subunit 111 sends n bits of information and the wid of each of at least two first thread bundles to the scheduling subunit 112; or, the reuse degree comparison subunit 111 sends n bits of information and at least two first thread bundles to the scheduling subunit 112.
[0132] The polling arbiter 1121 is used to select a first thread bundle from a valid first thread bundle (or a thread bundle that satisfies a first condition); or, to select a valid first thread bundle from at least two first thread bundles. For example, according to the polling scheduling policy, a first thread bundle is selected as the target thread bundle.
[0133] The strict priority arbiter 1122 is used to determine whether to send the second thread bundle or the first thread bundle. Specifically, the strict priority arbiter 1122 sends the target thread bundle to the pipeline 200 if the second thread bundle meets the second condition; otherwise, it sends the second thread bundle to the pipeline 200.
[0134] The second condition is that the second thread bundle is in a blocked state, or the second condition is that the second thread bundle is suspended, or the second condition is that the second thread bundle is interrupted.
[0135] Pipeline 200 is used to execute the received second thread bundle or target thread bundle.
[0136] When a memory access instruction exists in the first thread bundle, pipeline 200 accesses data cache 300 according to the memory access instruction. Specifically, pipeline 200 sends a memory access request to data cache 300, where req is the request information and req_wid is the thread bundle corresponding to the request.
[0137] Specifically, the tag array in data cache 300 records tag information and thread bundle index (wid) information. The thread bundle index information indicates which thread bundle the cache line corresponds to. The tag information is used to determine whether the data corresponding to the memory access address is in data cache 300. When accessing the tag array, the tag of which line to access is determined based on the memory access address, and then compared with the tag information in each way. If the tags are the same, it means a cache hit. If the tags in each way are different from the input tag (i.e., the tag corresponding to the memory access address), it means that the input request is cached missing, and it is necessary to search from the lower-level cache and then select a way to replace the cache line. For data cache 300, a cache line is the smallest storage unit managed by the cache, also called a cache block. Each cache line can contain a Flag (i.e., valid bit, also called valid), a Tag, and Data. The Data size is usually 64 bytes, but the Flag and Tag may be different for different CPU models. Data is loaded from memory to the cache line by line, with one cache line corresponding to a memory block of the same size. In some embodiments, such as Figure 9 As shown, the cache is arranged in a matrix (M×N), with sets horizontally and paths vertically. Each element is a cache line. Given the memory address, first find the set number it belongs to, then traverse all paths in that set, continuing until the label in the cache line matches the label in the memory address. If no match is found across all paths, then a cache miss occurs.
[0138] If data cache 300 includes the data corresponding to the memory access instruction, the access is successful. If the wid corresponding to the accessed cache line is the second thread bundle, meaning there is data reuse between the target thread bundle and the second thread bundle, the reuse degree of the target thread bundle needs to be incremented by 1. If data cache 300 does not include the data corresponding to the memory access instruction, the access is unsuccessful, i.e., a cache miss occurs. The data corresponding to the memory access instruction needs to be read from the lower-level cache or memory into data cache 300. If the read data replaces the data corresponding to the second thread bundle in data cache 300, the reuse degree of the target thread bundle needs to be decremented by 1.
[0139] Specifically, a replacement status register exists in the data cache 300, which records access information in the data cache 300. This access information is used by the update unit 120 in the scheduler 100 to determine how to update the reuse of each thread bundle in at least two first thread bundles (or the reuse of the target thread bundle).
[0140] For example, such as Figure 10As shown, assuming that data cache 300 includes two paths, if the request label is the same as the label in the row corresponding to the first path, it means that the cache has been hit, and the thread bundle label corresponding to this cache row is 3, which is the cache row corresponding to the second thread bundle. This indicates that the current accessed request and the second thread bundle have data reuse, so the reuse degree of the corresponding target thread bundle is incremented by 1.
[0141] For example, such as Figure 11 As shown, the request label is compared with the labels of the two corresponding lines. They are found to be different, which means that a cache miss has occurred. According to the replacement algorithm, the cache line of one of the lines needs to be replaced. However, the thread bundle labels corresponding to these two lines are both the second thread bundle. This means that the current cache replacement process has replaced a cache line of the second thread bundle, resulting in a cache conflict. The reuse of the corresponding target thread bundle needs to be reduced by 1.
[0142] In summary, the scheduler provided in this application statistically analyzes the data cache conflict / reuse situation between low-priority and high-priority thread bundles. This is reflected in the reuse degree of each thread bundle. A higher reuse degree indicates a higher data reuse rate between the thread bundle and the highest-priority thread bundle, while a lower reuse degree indicates a greater cache conflict. Each time a lower-priority thread is selected for execution, it is chosen only from thread bundles with higher reuse degrees, minimizing cache replacement of the highest-priority thread bundle and improving its execution efficiency. By selecting based on data cache conflict, thread bundles with data reuse with the highest-priority thread bundle are prioritized, reducing the probability of data cache conflicts and improving the execution efficiency of the highest-priority thread bundle, thereby improving processor utilization. Furthermore, this scheduler boasts high performance and consumes a small chip area, effectively enhancing product competitiveness.
[0143] Figure 12 A flowchart illustrating a scheduling method provided in an exemplary embodiment of this application is shown. This method can be executed by the scheduler 100 described above. The scheduler 100 includes a scheduling unit and an update unit.
[0144] Step 410: The scheduling unit obtains the reuse degree of each of the at least two first thread bundles. The reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache. The second thread bundle is the highest priority thread bundle. Based on the reuse degree of each of the at least two first thread bundles, the target thread bundle is emitted to the pipeline. The target thread bundle is one of the at least two first thread bundles.
[0145] Optionally, the first thread bundle is a ready thread bundle; or, the first thread bundle is a thread bundle in a ready state; or, all threads in the first thread bundle are in a ready state.
[0146] Optionally, the second thread bundle is a blocked thread bundle; or, the second thread bundle is an interrupted thread bundle; or, the second thread bundle is a thread bundle in a blocked state; or, all or some of the threads in the second thread bundle are in a blocked state.
[0147] Optionally, the second thread bundle has a higher priority than any of the at least two first thread bundles; in other words, the second thread bundle is the highest priority thread bundle that the scheduler can schedule. The second thread bundle can also be called the highest priority thread bundle, and correspondingly, each of the at least two first thread bundles can be called a low priority thread bundle.
[0148] Optionally, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data used by the first thread bundle and data used by the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data used by the first thread bundle and data used by the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between data used by the first thread bundle and data used by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between data read into the data cache by the first thread bundle and data read into the data cache by the second thread bundle; or The reuse degree of the first thread bundle is used to indicate the degree of reuse between the data cached in the data cache by the first thread bundle and the data cached in the data cache by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between the data used by the first thread bundle and the data used by the second thread bundle; or, the reuse degree of the first thread bundle is used to indicate the degree of reuse between the data used by the first thread bundle and the data of the second thread bundle in the data cache; or, the reuse degree of the first thread bundle is used to indicate the number of times a cache miss of the first thread bundle is resolved based on the data of the second thread bundle in the data cache (or cache line, cache, cache space, etc.); or, the reuse degree of the first thread bundle is used to indicate the number of times the first thread bundle hits the data of the second thread bundle in the data cache.
[0149] Optionally, the reuse rate of the first thread bundle is an estimated value. In other words, the reuse rate of the first thread bundle is calculated in real time based on whether a cache miss occurs after the first thread bundle is executed, and whether the cache miss causes the data of the second thread bundle in the data cache to be replaced. It can also be said that the reuse rate of the first thread bundle is not a precise value that can accurately describe the degree of data reuse between the first thread bundle and the second thread bundle, but rather describes the number of times the first thread bundle has hit the data of the second thread bundle in the data cache, that is, the number of times the first thread bundle has used the data of the second thread bundle in the data cache.
[0150] Optionally, the scheduling unit selects a target thread bundle and issues the target thread bundle to the pipeline based on the reuse degree of each of the at least two first thread bundles. The target thread bundle may be the first thread bundle with the highest reuse degree among the at least two first thread bundles, or any first thread bundle among the at least two first thread bundles that meets the reuse degree condition, or the first thread bundle among the at least two first thread bundles that meets both the reuse degree and priority conditions, and so on.
[0151] Optionally, the priority of the first thread bundle is determined based on the age of the first thread bundle, such that the older the first thread bundle is, the higher its priority; or, the priority of the first thread bundle is determined based on the thread bundle index (wid) of the first thread bundle, such that the smaller the thread bundle index, the higher its priority, and so on.
[0152] Step 420: The update unit updates the reusability of the target thread bundle when the target thread bundle accesses the data cache.
[0153] Optionally, the update unit updates the reuse of the target thread bundle when the target thread bundle accesses the data cache; or, the update unit updates the reuse of the target thread bundle when a memory access instruction in the target thread bundle is executed.
[0154] In summary, the method provided in this application determines the target thread bundle based on its reuse rate when scheduling the first thread bundle. Furthermore, it updates the reuse rate of the target thread bundle when the target thread bundle accesses the data cache. The reuse rate of the first thread bundle indicates the number of times the first thread bundle uses data from the second thread bundle in the data cache. By selecting the target thread bundle to execute when the second thread bundle cannot execute, a first thread bundle with a higher reuse rate than the second thread bundle can be selected. This avoids the situation where a cached data of the second thread bundle is lost during the execution of the first thread bundle, affecting the data cached by the second thread bundle in the data cache. This prevents the second thread bundle from resuming execution due to its cached data being replaced by the target thread bundle, leading to a cached data loss and requiring it to read the required data from the lower-level cache again, thus reducing its execution efficiency.
[0155] 1. Scheduling Unit
[0156] The following section further illustrates how the scheduling unit selects the target thread bundle.
[0157] In some embodiments, step 410 above may be implemented as follows: the scheduling unit determines at least one first thread bundle among at least two first thread bundles that satisfies a first condition based on the reuse degree of each of the at least two first thread bundles, the first condition including the reuse degree of the first thread bundle being greater than or equal to a reuse degree threshold; and sends a target thread bundle to the pipeline, the target thread bundle being one of the at least one first thread bundles that satisfies the first condition.
[0158] Optionally, if at least two first thread bundles do not include a first thread bundle that satisfies the first condition, the scheduling unit sends a prompt message to the scheduler (or update unit) indicating that there is no first thread bundle that satisfies the first condition, so that the scheduler (or update unit) can reset or initialize the reuse of all or part of the first thread bundles; or, the scheduling unit resets or initializes the reuse of all or part of the first thread bundles. In other embodiments, if the first thread bundle satisfying the first condition is not included among the at least two second thread bundles, the scheduling unit or scheduler may also adopt other solutions, such as directly selecting any first thread bundle as the target thread bundle from the at least two first thread bundles, or selecting the oldest first thread bundle as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the smallest thread bundle index (wid) as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest priority as the target thread bundle from the at least two first thread bundles, or selecting the first thread bundle with the highest reuse rate as the target thread bundle from the at least two first thread bundles, or reducing the reuse rate threshold, or notifying the update unit 120 to reset the reuse rate of some first thread bundles, etc. In addition to the methods shown in the embodiments of this application, other methods can be used to solve the above-mentioned special cases of abnormal situations, combined with actual hardware design. The embodiments of this application do not limit this to such cases, but the protection scope of the embodiments of this application is not limited thereto.
[0159] In some embodiments, the scheduling unit includes a reuse degree comparison subunit and a scheduling subunit.
[0160] Optionally, the reuse comparison subunit obtains the i-th first thread bundle among at least two first thread bundles; if the reuse of the i-th first thread bundle is greater than or equal to the reuse threshold, the i-th first thread bundle is determined to be a first thread bundle that satisfies the first condition; let i = i + 1, and repeat the process starting from obtaining the i-th first thread bundle until at least two first thread bundles have been judged; send at least one first thread bundle that satisfies the first condition to the scheduling subunit.
[0161] Optionally, the reuse degree comparison subunit sending at least one first thread bundle satisfying the first condition to the scheduling subunit can be implemented as follows: the reuse degree comparison subunit sends n bits of information to the scheduling subunit, which is used by the scheduling subunit to determine the first thread bundle that can be scheduled; or, the reuse degree comparison subunit sends the wid of each of the at least one first thread bundle satisfying the first condition to the scheduling subunit; or, the reuse degree comparison subunit sends n bits of information and the wid of each of at least two first thread bundles to the scheduling subunit; or, the reuse degree comparison subunit sends n bits of information and at least two first thread bundles to the scheduling subunit.
[0162] Optionally, the scheduling subunit is used to send the target thread bundle to the pipeline.
[0163] Optionally, the scheduling subunit selects a target thread bundle from at least one first thread bundle that satisfies the first condition. For example, it may directly select any first thread bundle as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the oldest first thread bundle as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the first thread bundle with the smallest thread bundle index (wid) as the target thread bundle from at least one first thread bundle that satisfies the first condition, or select the first thread bundle with the highest priority as the target thread bundle from at least one first thread bundle that satisfies the first condition.
[0164] The following further illustrates the scheduling method for at least one first thread bundle that satisfies the first condition, i.e., how to select the target thread bundle from at least one first thread bundle that satisfies the first condition.
[0165] In some embodiments, the scheduling subunit includes a polling arbitrator and a strict priority arbitrator.
[0166] Optionally, the polling arbitrator selects the target thread bundle from at least one first thread bundle that satisfies the first condition. Optionally, the polling arbitrator selects the target thread bundle from at least one first thread bundle that satisfies the first condition according to a polling scheduling strategy.
[0167] In some embodiments, the polling scheduling strategy is as follows: at least two first thread bundles are n first thread bundles, where n is a positive integer; the polling arbiter obtains a polling flag, which indicates that the first thread bundle previously selected by the polling arbiter is the j-th first thread bundle among at least two first thread bundles, where j is a positive integer; if at least one first thread bundle satisfying the first condition includes the k-th first thread bundle, the k-th first thread bundle is selected as the target thread bundle, and the polling flag is updated to indicate the k-th first thread bundle, where k is a positive integer and the initial value of k is j+1; if at least one first thread bundle satisfying the first condition does not include the k-th first thread bundle, k = k+1; if k is greater than n, k = kn.
[0168] Optionally, the strict priority arbiter fires the target thread bundle if the second thread bundle satisfies a second condition, the second condition including that the second thread bundle is suspended.
[0169] Optionally, the strict priority arbiter fires the second thread bundle if the second thread does not meet the second condition. That is, if the second thread bundle can execute normally, it is fired first, i.e., the highest priority thread bundle. If the second thread bundle cannot execute normally, the target thread bundle is fired, i.e., the first thread bundle selected from at least two first thread bundles, i.e., the low priority thread bundle.
[0170] Optionally, the second condition includes at least one of the following: the second thread bundle is suspended; the second thread bundle is interrupted; the second thread bundle is blocked; the second thread bundle is in a blocked state; or all or some of the threads in the second thread bundle are in a blocked state.
[0171] For details, please refer to "1. Scheduling Unit" shown in the exemplary embodiment of the scheduler above, which will not be repeated here.
[0172] 2. Reusability
[0173] To reduce the delay in the scheduler reading these two parameters, the reuse degree and reuse threshold can be stored using registers.
[0174] In some embodiments, the method further includes: the scheduling unit reading a reuse threshold from a reuse threshold register.
[0175] In some embodiments, the method further includes: the scheduling unit reading the reuse degree of each first thread bundle from the reuse degree register corresponding to each of the at least two first thread bundles.
[0176] In addition, when each target thread bundle is executed, the reusability of the target thread bundle needs to be updated in a timely manner based on the execution status of the target thread bundle, as shown below.
[0177] (1) Reusability update
[0178] In some embodiments, step 420 above can be implemented as follows: the updating unit increases the reuse of the target thread bundle when the target thread bundle accesses the data corresponding to the second thread bundle in the data cache. Alternatively, the target thread bundle does not update its reuse when it accesses the data corresponding to the second thread bundle in the data cache.
[0179] In some embodiments, step 420 above can be implemented as follows: When the target thread bundle replaces the data corresponding to the second thread bundle in the data cache, the updating unit reduces the reuse rate of the target thread bundle. Alternatively, when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache, the reuse rate of the target thread bundle is not updated.
[0180] For details, please refer to “2. Reusability” in “(1) Reusability Update” shown in the exemplary embodiment of the scheduler above, which will not be repeated here.
[0181] The above illustrates some scenarios of reuse reset. For example, if there is no first thread bundle that meets the first condition, reuse may need to be reset to ensure the scheduler functions correctly. The following sections further illustrate more reuse reset scenarios and the methods for resetting reuse.
[0182] (2) Reuse rate reset
[0183] In some embodiments, the method further includes: the updating unit resetting the reuse rate of each of the at least two first thread bundles if a third condition is met; or, resetting the reuse rate of some of the at least two first thread bundles if the third condition is met; or, updating the reuse rate of all or some of the at least two first thread bundles if the third condition is met.
[0184] The third condition includes at least one of the following: the reset time reaches a reset time threshold, where the reset time is used to indicate the time between the current moment and the moment of the last reset of reuse; a first thread bundle is identified as a new second thread bundle; and the number of at least one first thread bundle that satisfies the first condition is less than a number threshold.
[0185] For details, please refer to “2. Reusability” in “(2) Resetting of Reusability” shown in the exemplary embodiment of the scheduler above, which will not be repeated here.
[0186] Figure 13A structural block diagram of a scheduling apparatus provided in an exemplary embodiment of this application is shown. The apparatus has the functionality to implement the scheduling method example described above; the functionality can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the scheduler described above, or it can be located within a scheduler. The scheduler includes a scheduling module 510 and an update module 520.
[0187] The scheduling module 510 is used to obtain the reuse degree of each of the at least two first thread bundles, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, a target thread bundle is emitted to the pipeline, the target thread bundle is one of the at least two first thread bundles.
[0188] The update module 520 is used to update the reuse rate of the target thread bundle when the target thread bundle accesses the data cache.
[0189] In some embodiments, the scheduling module 510 is configured to determine at least one first thread bundle among the at least two first thread bundles that satisfies a first condition based on the reuse degree of each of the at least two first thread bundles, the first condition including the reuse degree of the first thread bundle being greater than or equal to a reuse degree threshold; and to send the target thread bundle to the pipeline, the target thread bundle being one of the first thread bundles among the at least one first thread bundle that satisfies the first condition.
[0190] In some embodiments, the scheduling module 510 includes a reuse degree comparison submodule and a scheduling submodule.
[0191] The reuse comparison submodule is used to obtain the i-th first thread bundle among the at least two first thread bundles; if the reuse of the i-th first thread bundle is greater than or equal to the reuse threshold, determine that the i-th first thread bundle is a first thread bundle that satisfies the first condition; let i = i + 1, and repeat the process from obtaining the i-th first thread bundle until all at least two first thread bundles have been judged; and send at least one first thread bundle that satisfies the first condition to the scheduling subunit.
[0192] The scheduling submodule is used to send the target thread bundle to the pipeline.
[0193] In some embodiments, the scheduling submodule includes a polling submodule and a strict priority arbitration submodule.
[0194] The polling submodule is used to select the target thread bundle from at least one first thread bundle that satisfies the first condition.
[0195] The strict priority arbitration submodule is used to launch the target thread bundle when the second thread bundle meets a second condition, the second condition including that the second thread bundle is suspended.
[0196] In some embodiments, at least two first thread bundles are n first thread bundles, where n is a positive integer.
[0197] The polling submodule is used to obtain a polling flag, which indicates that the first thread bundle previously selected by the polling arbiter is the j-th first thread bundle among the at least two first thread bundles, where j is a positive integer; if at least one first thread bundle satisfying the first condition includes the k-th first thread bundle, the k-th first thread bundle is selected as the target thread bundle, and the polling flag is updated to indicate the k-th first thread bundle, where k is a positive integer and the initial value of k is j+1; if at least one first thread bundle satisfying the first condition does not include the k-th first thread bundle, k = k+1; if k is greater than n, k = kn.
[0198] In some embodiments, the update module 520 is used to increase the reusability of the target thread bundle when the target thread bundle accesses the data corresponding to the second thread bundle in the data cache.
[0199] In some embodiments, the update module 520 reduces the reusability of the target thread bundle when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache.
[0200] In some embodiments, the update module 520 is configured to reset the reuse of each of the at least two first thread bundles when a third condition is met.
[0201] The third condition includes at least one of the following: the reset time reaches a reset time threshold, where the reset time indicates the time between the current moment and the last time the reuse was reset; or a first thread bundle is identified as a new second thread bundle.
[0202] In some embodiments, the scheduling module 510 is configured to read the reuse threshold from the reuse threshold register.
[0203] In some embodiments, the scheduling unit 510 is configured to read the reuse degree of each first thread bundle from the reuse degree register corresponding to each of the at least two first thread bundles.
[0204] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0205] On the other hand, embodiments of this application provide a processor that includes the scheduler described in the above embodiments.
[0206] On the other hand, embodiments of this application provide a graphics card that includes the scheduler described in the above embodiments.
[0207] On the other hand, embodiments of this application provide a computer device, which includes the scheduler described above. This computer device can be at least one of a portable computer, a desktop computer, a server, a server cluster, an artificial intelligence (AI) computing cluster, or a cloud computing cluster. The AI computing cluster can also be simply referred to as an intelligent computing cluster or a smart computing cluster.
[0208] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0209] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A scheduler, characterized in that, The scheduler includes a scheduling unit and an update unit; The scheduling unit is used to obtain the reuse degree of each of at least two first thread bundles, the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; Based on the reusability of each of the at least two first thread bundles, a target thread bundle is emitted into the pipeline, wherein the target thread bundle is one of the at least two first thread bundles; The update unit is used to update the reuse rate of the target thread bundle when the target thread bundle accesses the data cache.
2. The scheduler according to claim 1, characterized in that, The scheduling unit is configured to determine, based on the reuse degree of each of the at least two first thread bundles, at least one first thread bundle that satisfies a first condition, wherein the first condition includes the reuse degree of the first thread bundle being greater than or equal to a reuse degree threshold; and to send the target thread bundle to the pipeline, wherein the target thread bundle is one of the at least one first thread bundles that satisfies the first condition.
3. The scheduler according to claim 2, characterized in that, The scheduling unit includes a reuse degree comparison subunit and a scheduling subunit; The reuse degree comparison subunit is used to obtain the i-th first thread bundle among the at least two first thread bundles, where i is a positive integer; If the reuse degree of the i-th first thread bundle is greater than or equal to the reuse degree threshold, the i-th first thread bundle is determined to be a first thread bundle that satisfies the first condition. Let i = i + 1, and repeat the process starting from obtaining the i-th first thread bundle until all at least two first thread bundles have been judged. Send at least one first thread bundle that satisfies the first condition to the scheduling subunit; The scheduling subunit is used to send the target thread bundle to the pipeline.
4. The scheduler according to claim 3, characterized in that, The scheduling subunit includes a round-robin arbitrator and a strict priority arbitrator; The polling arbiter is used to select the target thread bundle from at least one first thread bundle that satisfies the first condition; The strict priority arbiter is used to launch the target thread bundle if the second thread bundle satisfies a second condition, the second condition including that the second thread bundle is suspended.
5. The scheduler according to claim 4, characterized in that, The at least two first thread bundles are n first thread bundles, where n is a positive integer; The polling arbiter is used to obtain a polling flag, which indicates that the first thread bundle previously selected by the polling arbiter is the j-th first thread bundle among the at least two first thread bundles, where j is a positive integer; if at least one first thread bundle satisfying the first condition includes the k-th first thread bundle, the k-th first thread bundle is selected as the target thread bundle, and the polling flag is updated to indicate the k-th first thread bundle, where k is a positive integer and the initial value of k is j+1; if at least one first thread bundle satisfying the first condition does not include the k-th first thread bundle, k = k+1; if k is greater than n, k = kn.
6. The scheduler according to any one of claims 1 to 5, characterized in that, The update unit is used to increase the reusability of the target thread bundle when the target thread bundle accesses the data corresponding to the second thread bundle in the data cache.
7. The scheduler according to any one of claims 1 to 6, characterized in that, The update unit is used to reduce the reuse of the target thread bundle when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache.
8. The scheduler according to any one of claims 1 to 7, characterized in that, The update unit is used to reset the reuse degree of each of the at least two first thread bundles when the third condition is met. The third condition includes at least one of the following: The reset time has reached the reset time threshold. The reset time is used to indicate the time between the current moment and the last time the reuse degree was reset. There exists a first thread bundle that is identified as a new second thread bundle.
9. The scheduler according to any one of claims 2 to 5, characterized in that, The scheduling unit is used to read the reuse threshold from the reuse threshold register.
10. The scheduler according to any one of claims 1 to 9, characterized in that, The scheduling unit is configured to read the reuse degree of each first thread bundle from the reuse degree register corresponding to each of the at least two first thread bundles.
11. A scheduling method, characterized in that, The scheduling method is executed by a scheduler as described in any one of claims 1 to 10, the scheduler comprising a scheduling unit and an update unit; The method includes: The scheduling unit obtains the reuse degree of each of the at least two first thread bundles, where the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, a target thread bundle is emitted to the pipeline, where the target thread bundle is one of the at least two first thread bundles; The update unit updates the reuse rate of the target thread bundle when the target thread bundle accesses the data cache.
12. The method according to claim 11, characterized in that, The scheduling unit obtains the reuse degree of each of at least two first thread bundles. The reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache. The second thread bundle is the highest priority thread bundle. Based on the reusability of each of the at least two first thread bundles, a target thread bundle is emitted into the pipeline, wherein the target thread bundle is one of the at least two first thread bundles, including: The scheduling unit determines at least one first thread bundle among the at least two first thread bundles that satisfies a first condition based on the reuse degree of each of the at least two first thread bundles, the first condition including that the reuse degree of the first thread bundle is greater than or equal to a reuse degree threshold; and sends the target thread bundle to the pipeline, the target thread bundle being one of the at least one first thread bundles that satisfies the first condition.
13. The method according to claim 12, characterized in that, The scheduling unit includes a reuse degree comparison subunit and a scheduling subunit; The scheduling unit determines at least one of the at least two first thread bundles that satisfies a first condition based on the reuse degree of each of the at least two first thread bundles, wherein the first condition includes the reuse degree of the first thread bundle being greater than or equal to a reuse degree threshold. The target thread bundle is emitted to the pipeline, wherein the target thread bundle is one of the first thread bundles satisfying the first condition, including: The reuse degree comparison subunit obtains the i-th first thread bundle among the at least two first thread bundles, where i is a positive integer; if the reuse degree of the i-th first thread bundle is greater than or equal to the reuse degree threshold, the i-th first thread bundle is determined to be a first thread bundle that satisfies the first condition; let i = i + 1, and repeat the process from obtaining the i-th first thread bundle until all at least two first thread bundles have been judged; send at least one first thread bundle that satisfies the first condition to the scheduling subunit; The scheduling subunit sends the target thread bundle to the pipeline.
14. The method according to any one of claims 11 to 13, characterized in that, The update unit updates the reuse rate of the target thread bundle when the target thread bundle accesses the data cache, including: The update unit increases the reusability of the target thread bundle when the target thread bundle accesses the data corresponding to the second thread bundle in the data cache.
15. The method according to any one of claims 11 to 14, characterized in that, The update unit updates the reuse rate of the target thread bundle when the target thread bundle accesses the data cache, including: The update unit reduces the reusability of the target thread bundle when the target thread bundle replaces the data corresponding to the second thread bundle in the data cache.
16. The method according to any one of claims 11 to 15, characterized in that, The method further includes: If the third condition is met, the update unit resets the reuse degree of each of the at least two first thread bundles; The third condition includes at least one of the following: The reset time has reached the reset time threshold. The reset time is used to indicate the time between the current moment and the last time the reuse degree was reset. There exists a first thread bundle that is identified as a new second thread bundle.
17. A scheduling device, characterized in that, The scheduling device includes a scheduling module and an update module; the device includes: The scheduling module obtains the reuse degree of each of the at least two first thread bundles, where the reuse degree of the first thread bundle is used to indicate the degree of reuse of data between the first thread bundle and the second thread bundle in the data cache, and the second thread bundle is the highest priority thread bundle; based on the reuse degree of each of the at least two first thread bundles, a target thread bundle is emitted to the pipeline, where the target thread bundle is one of the at least two first thread bundles; The update module updates the reusability of the target thread bundle when the target thread bundle accesses the data cache.
18. A processor, characterized in that, The processor includes the scheduler as described in any one of claims 1 to 10.
19. A graphics card, characterized in that, The graphics card includes a scheduler as described in any one of claims 1 to 10.
20. A computer device, characterized in that, The computer device includes a scheduler as described in any one of claims 1 to 10.