AUTONOMOUS JOB QUEUE SYSTEM FOR HARDWARE ACCELERATORS
A hardware-based job queue management system using TCBs and DMA infrastructure addresses the integration challenges of hardware accelerators in DSPs, enhancing their utilization and efficiency by allowing core tasks to be offloaded without waiting for completion and managing queues based on priority.
Patent Information
- Application Number
- DE102020117505
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-23
- Filing Date
- 2020-07-02
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2040-07-02
AI Technical Summary
Hardware accelerators in digital signal processors (DSPs) face integration challenges due to the need for core-driven data stream management, leading to inefficiencies and underutilization unless they perform significantly faster than the core's execution time, making seamless integration difficult.
A hardware-based autonomous job queue management system using a chained DMA infrastructure and task control blocks (TCBs) to efficiently offload tasks to hardware accelerators, allowing the core to commit tasks without waiting for completion, and manage queues based on priority levels.
Enhances the utilization of hardware accelerators by enabling efficient task offloading and management, reducing computational and memory intensity, and improving overall processing efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Area of Revelation
[0001] This disclosure relates generally to the field of computing and, more particularly, though not exclusively, to a system and method for queuing jobs in a computing device. background
[0002] Hardware accelerators are ubiquitous in digital signal processors (DSPs) to power the most common DSP routines such as finite impulse response (FIR), infinite impulse response (IIR), fast Fourier transform (FFT), and so on. While the idea may be to significantly increase computational performance, there is often a high barrier to their effective use due to the need for the instruction processing core to build data streams and track and manage execution. Such core-driven management often also requires significant instruction flows. This makes them very difficult to integrate seamlessly into the data stream through the DSP, where the core and accelerator must efficiently share the processing load. As a result, accelerators remain unused unless the accelerator can perform a computational task many times faster than the core's execution time and the associated overhead.
[0003] US 4 658 351 A relates to a task control system for a multi-tasking data processing system that controls the simultaneous execution of multiple tasks. The system includes a task manager that organizes tasks according to their priority. Each task is placed in a queue that corresponds to the task's priority. The task manager ensures that tasks are executed in order of priority, taking the task's status (active or inactive) into account. The system is intended to enable efficient task management in a multi-user system and ensure optimal task execution.
[0004] US 6 105 127 A relates to a multithreaded processor that executes multiple instruction streams simultaneously. The processor comprises multiple functional units that execute instructions, as well as multiple decode units, each assigned to an instruction stream. These units decode instructions and generate requests to assign the instructions to the corresponding functional units. A control unit decides which instructions are sent to which functional unit when multiple requests are present simultaneously, taking into account the priority levels of the instruction streams. The goal is to improve the efficiency of instruction processing by flexibly adjusting the priorities to optimize the processing speed of the instruction streams.
[0005] US 2003 / 0 208 521 A1 relates to a system and method for thread scheduling using a weak preemption approach. The scheduler receives requests from newly prepared threads. The scheduler adds a "preempt value" to the priority of the current work to make it more difficult for newly prepared threads to interrupt the ongoing work. A less strict preemption procedure is applied, so that threads with slightly higher priority do not immediately interrupt the ongoing work, but are executed after it has completed. This contributes to reducing system overhead by avoiding unnecessary interruptions.
[0006] US 2004 / 0 160 446 A1 relates to systems and methods for scheduling the processing of tasks by a coprocessor, such as a graphics unit. Applications submit tasks to a scheduler, which decides how much computing power each application receives and in what order they are processed. The system uses techniques to ensure system security by preventing applications from modifying important memory areas required for system operations. The scheduler is designed to ensure that tasks are executed efficiently without unnecessary waiting times. Brief description of the drawings Fig. 1 depicts an example structure of a task control block (TCB) in accordance with various embodiments. Fig.2 depicts an example of queuing jobs in accordance with various embodiments. Fig. 3 depicts an example application programming interface (API) that may be used to perform job queuing in a hardware accelerator, in accordance with various embodiments. Fig. 4 depicts another example API that may be used to perform queuing of TCB jobs, in accordance with various embodiments. Fig. 5 depicts another example API that may be used to perform queuing of TCB jobs, in accordance with various embodiments. Fig.6 depicts another example API that may be used to perform queuing of TCB jobs, in accordance with various embodiments. Fig. 7 depicts another example API that may be used to perform queuing of TCB jobs, in accordance with various embodiments. Fig. 8 is a block diagram of an example electrical device that may include a hardware accelerator and a processor configured to perform queuing of TCB jobs, in accordance with various embodiments. Summary of Revelation
[0007] A computing device according to claim 1, a processor according to claim 8, and one or more non-transitory computer-readable media according to claim 14 are claimed. Preferred embodiments are claimed in the subclaims. Detailed description
[0008] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, wherein like reference characters designate like parts throughout, and in which is shown by way of illustration embodiments in which the subject matter of the present disclosure may be practiced. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.
[0009] For the purposes of this disclosure, the term "A or B" means (A), (B), or (A and B). For the purposes of this disclosure, the term "A, B, or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0010] The description may use the terms "in one embodiment" or "in embodiments," each of which may refer to one or more of the same or different embodiments. Furthermore, the terms "comprising," "including," "having," and the like, as used with respect to embodiments of the present disclosure, are synonymous.
[0011] The term "coupled to," along with its derivatives, may be used herein. "Coupled" may mean one or more of the following. "Coupled" may mean that two or more elements are in direct physical or electrical contact. However, "coupled" may also mean that two or more elements are in indirect contact with each other, yet they still cooperate or interact with each other, and it may mean that one or more other elements are coupled or connected between the elements said to be coupled. The term "directly coupled" may mean that two or more elements are in direct contact. In some embodiments, "coupled" may refer to two elements being physically placed in close proximity to each other and having exclusive and very fast access to each other.
[0012] Different operations may be described as a plurality of discrete operations in a manner best suited to understanding the claimed subject matter. However, the order of description should not be interpreted to imply that these operations are necessarily order-dependent.
[0013] As used herein, the term “module” may refer to, be a part of, or include an application-specific integrated circuit (ASIC), electronic circuit, processor (shared, dedicated, or group), or memory (shared, dedicated, or group) that executes one or more software or firmware programs, combinational logic circuitry, or other suitable components that provide the described functionality.
[0014] As mentioned above, hardware accelerators can be unused for a variety of reasons. However, a management system that allows the core to efficiently offload specific DSP tasks to a dedicated hardware accelerator and subsequently be notified by the hardware accelerator upon task completion can be useful. The applicability of such a system can increase if the management system allows the core to commit tasks to a pool without waiting for the current job to complete.
[0015] Embodiments herein may provide a variety of advantages. For example, a hardware accelerator managed by such a management system may allow a processor or processor core to offload one or more DSP tasks to the hardware accelerator. Additionally, a hardware-based accelerator queuing system may be advantageous over a software-based management system because a fully software-based management system may be computationally or memory-intensive because the processor must manage the task pool and assign the next task to the accelerator when the current task terminates.
[0016] In general, embodiments herein relate to a hardware-based autonomous job queue management structure. Embodiments may include a number of elements that may be used to implement the management structure. One such element may relate to a chained DMA infrastructure that may be used to manage the hardware accelerator. Another such element may relate to a submission of a new job within the management structure. In particular, the submission of a new job may be performed by creating a new entry in the DMA chain. The DMA chain may, for example, be a linked list of TCBs. In general, a TCB may be considered an element that relates to a specific job to be executed by the hardware accelerator.Each TCB may include one or more data elements or subfields that may contain specific configuration data for use by the hardware accelerator in executing the job. It should be understood that although the TCB is referred to herein as a "task control block," in other embodiments the TCB may be referred to by a different name, such as a "transmission control block."
[0017] Various TCBs or fields thereof can be modified to include information that can be used to support the queuing mechanism. For example, a TCB may include a pointer to an input (I / P) buffer. The TCB may also include a pointer to an output (O / P) buffer. The TCB may also include a pointer to a "previous" or "next" descriptor, which may be a reference to a previous or subsequent TCB in the DMA chain. The TCB may also include a placeholder that can be used to store a system address related to that TCB. The TCB may also include a job-specific hardware accelerator computation configuration. The TCB may also include, repurpose, or reuse a placeholder to store completion status.
[0018] In some embodiments, one or more APIs may be used to perform queue management tasks such as creating or submitting new jobs, identifying the status of an existing job, deleting jobs, etc. The APIs may perform these tasks by traversing the linked list of TCBs (e.g., using information related to identifying a last job or TCB of a given priority group), reading or modifying various TCBs or fields thereof, inserting or removing entries from the linked list, etc.
[0019] Jobs can then be executed in the order in which they appear in the TCB's linked list. In general, the TCB's linked list can have separate sections based on the different priorities of the submitted jobs. As a high-level example, a high-priority TCB section can be at the top of the TCB's linked list, and the jobs related to that TCB can be executed first. A medium-priority TCB section can follow the high-priority TCB section. The jobs related to that medium-priority TCB can be executed after the jobs related to the high-priority TCB have been executed. If a new job is received with a high-priority flag, it can be inserted at the end of the high-priority section of the TCB's linked list.If the high priority section of the linked list is finished, then the high priority TCB can be inserted immediately after the currently running job that refers to the TCB in the medium priority section.
[0020] Fig. 1 depicts an example TCB 105 that may be provided for execution by a hardware accelerator. As can be seen, the TCB 105 may include a number of fields. In some embodiments, the various fields of the TCB 105 may be referred to as "words," and the TCB 105 may include up to 16 fields.
[0021] Another such field may be the chain pointer 104, which may be labeled, for example, "FWD_CP" to indicate that it is a forward chain pointer. The chain pointer 104 may include an address of a subsequent TCB in the linked list of TCBs. After the hardware accelerator completes the job referenced by the TCB 105, the hardware accelerator may read the address in the chain pointer 104 and advance to the next TCB identified by the chain pointer 104. In some embodiments, the subsequent job identifier may be a "TCBYS_ADDR" identifier, which may be a system address of the subsequent TCB.
[0022] Another such field may be the JOB_ID field 106. In general, the JOB_ID field 106 may be a placeholder that may be used to store identification information related to the TCB 105 or the job to which the TCB 105 is associated. Each job described by the TCB is uniquely identified by the value of field 106. Such information may be used by the processor or hardware accelerator, for example, to identify a status of the job related to the TCB 105. The identification information may include, for example, the TCBYS_ADDR identifier.
[0023] Another such field may be the BACK_CP field 108. The BACK_CP field 108 may store an identifier of a previous TCB in the linked list of TCBs and may thus be referred to as a backward chain pointer. The identifier of the previous TCB may, for example, be the identifier described above with reference to the JOB_ID field 106. If each TCB in the linked list of TCBs has a BACK_CP field 108 identifying a previous job and a chain pointer identifying a subsequent job, then the corresponding TCBs may form the linked list by pointing to the previous or subsequent TCB. These fields may be used in combination by the APIs to traverse the linked list of TCBs, which may be useful for priority-based modification of the linked list, as described in more detail below.
[0024] Another such field may be the COMP_CTL field 112. The COMP_CTL field 112 may include a number of subfields 110 that may support job-specific calculation configuration. Although the term "subfields" is used herein, in other embodiments, the subfield may be referred to as a "flag" or a "data item."
[0025] One such subfield may be the FXD subfield 122. The FXD subfield 122 may indicate whether the job should use fixed-point computation, floating-point computation, or another type of computation. The available options may depend, for example, on how many bits are used for the FXD subfield 122. Another such subfield may be the TC subfield 124. The TC subfield 124 may be used to indicate to the accelerator the type of data it should operate on or the type of computation to be used by the hardware accelerator during job execution. Another such subfield may be the RND subfield 126, which may specify a rounding mode to be used during completion of the job associated with the TCB 105.
[0026] Another field of the TCB 105 may be a COEFF_P field 114. The COEFF_P field 114 may contain a reference to the table of coefficients that may be used by the hardware accelerator during execution of the job associated with the TCB 105. For example, in some embodiments, multiple data words may be associated with the COEFF_P field 114. The data words may specify the layout of the coefficient table in system memory.
[0027] Another field of TCB 105 may be the DATA_O / P_P field 116. The DATA_O / P_P field 116 may be or include a pointer to an O / P buffer. The O / P buffer may be, for example, a buffer of the hardware accelerator or a buffer of an electronic device of which the hardware accelerator is a part or to which the hardware accelerator is communicatively coupled. The O / P buffer may be where the hardware accelerator is to store data related to the execution or completion of the job with which TCB 105 is associated.
[0028] Another field of the TCB 105 may be the DATA_I / P_P field 118. The DATA_I / P_P field 118 may be or include a pointer to an I / P buffer. The I / P buffer may be, for example, a buffer of the hardware accelerator or a buffer of an electronic device of which the hardware accelerator is a part or to which the hardware accelerator is communicatively coupled. The I / P buffer may be where the hardware accelerator is to receive data related to the execution of the job to which the TCB 105 is assigned.
[0029] Another such field of the TCB 105 may include the CONFIG_CTL field 120. The CONFIG_CTL field 120 may include additional subfields 115 that may be used to implement the priority-based hardware acceleration scheme described herein. One such subfield may include the ACC_CONFIG subfield 102. The ACC_CONFIG subfield 102 may include information related to techniques that may be used by the hardware accelerator to execute the job related to the TCB.For example, the ACC_CONFIG subfield 102 may instruct the hardware accelerator to implement a finite impulse response (FIR) filter; a biquadratic infinite impulse response (IRR) filter of a specific order; a specific time window related to job execution, where the window length may relate to the number of output sample points to be produced; or a sample rate translation that may be used. However, it should be understood that these are examples of what the ACC_CONFIG subfield 102 may include, and other embodiments may include additional or fewer elements.
[0030] Another such subfield may include the PRIO subfield 128. The PRIO subfield 128 may include a number of bits that may indicate a priority level of the TCB 105 within the linked list of TCBs. For example, the PRIO subfield 128 may be a 1-bit flag that may indicate whether the TCB 105 is of a "high" or "low" priority. Alternatively, the PRIO subfield 128 may be a 2-bit flag that may indicate whether the TCB 105 is of a "very high," "high," "medium," or "low" priority.
[0031] Another such subfield may include the IMASK subfield 130. The IMASK subfield 130 may indicate whether an interrupt is to be sent to the processor upon completion of the job associated with the TCB 105. Generally, it may not be desirable to send an interrupt to the processor upon completion of every job, and thus, this subfield may serve as a flag to indicate whether an interrupt is desired.
[0032] Another such subfield may include the TMASK subfield 132. Similar to the IMASK subfield 130, the TMASK subfield 132 may indicate whether a trigger is to be sent to the processor subsystem upon completion of the job with which the TCB 105 is associated. Generally, it may not be desirable to send a trigger to the processor subsystem upon completion of every job, and thus, this subfield may serve as a flag to indicate whether a trigger is desired.
[0033] Another such subfield may include the TWAIT subfield 134. In general, before the processor can submit a new TCB to the job pool by adding it to the linked list, both a data buffer containing one or more coefficients (the location of which may be specified by the COEFF_P field 114) and the input data buffer (which may be the data in the I / P buffer specified by the DATA_I / P_P field 118) may be necessary. In general, the coefficient may be static, but the input data may be dynamic. If the input data is unavailable, the hardware accelerator may process old or corrupted data during the execution of the job specified by the TCB 105.If the TWAIT subfield for TCB 105 is set, then when the hardware accelerator executes the job related to that TCB, the hardware accelerator can wait for an indication of I / P buffer completion in the form of a trigger, which may be issued by any of the components of the processor subsystem, such as a DMA, which may be responsible for preparing the buffer. Once the trigger is received by the processor, the hardware accelerator can start the input DMA to transfer data to its internal buffer so that it can execute the relevant job. As a result, a TCB can be committed by the processor to the linked list even if the corresponding input data is not available in the I / P buffer at the time of the commit by the processor.
[0034] In general, it should be understood that the TCB 105 and the relevant fields / subfields, etc., are depicted and discussed herein as examples of such fields / subfields, and other embodiments may vary from those depicted. For example, some TCBs may have additional or fewer fields than depicted. In some embodiments, the fields described may differ from the exact parameters described. As a specific example, in some embodiments, the PRIO subfield 128 may have more or fewer bits than the described 1-bit or 2-bit arrangements. Additionally, it should be understood that the specific names given in this example are names that could be used in one embodiment, but other names for the specific fields or subfields may be used in other embodiments while still performing the described functionality.Other variations may be present in other embodiments.
[0035] Fig. Figure 2 depicts a high-level example of queuing jobs in accordance with various embodiments. In particular, Fig. 2 shows a high-level example of queuing a newly introduced TCB into a linked list of TCBs. Each of the discussed TCBs can have a priority level such as the one discussed with reference to the PRIO subfield 128. In this particular example, three different priority levels are present, which are Fig.2 based on different shading. The three different priority levels are high priority 202, medium priority 204, and low priority 206. Generally, TCBs with high priority levels should be executed before TCBs with medium or low priority levels. Similarly, TCBs with medium priority levels should be executed before TCBs with low priority levels.
[0036] The system may include a processor 210 communicatively coupled to the hardware accelerator 205 by a system bus 203. The system bus 203 may be a communicative coupling that enables one or more elements of an electronic device to send or receive data signals to or from each other. The processor 210 may be, for example, a central processing unit (CPU), a multi-core or single-core processor, a core of a multi-core processor, or another type of processor. Similarly, the hardware accelerator 205 may be a processing unit, such as a hardware logic block configured to perform specific functions or a group of functions, a CPU, a multi-core or single-core processor, a core of a multi-core processor, or another type of processor.The processor 210 may further include logic for executing one or more computer-executable instructions, logic for performing one or more mathematical processes or calculations, etc.
[0037] The system may further include a memory 201. The memory 201 may be volatile or non-volatile memory, such as double data rate (DDR) memory, flash memory, random access memory (RAM), or another type of memory. In particular, the memory 201 may include one or more buffers, such as the I / P buffers, O / P buffers, or COEFF buffers discussed above. Additionally, the memory 210 may be configured to store one or more TCBs in the linked list of TCBs, in accordance with various embodiments discussed herein. The memory 201 may be communicatively coupled to the processor 210 and the hardware accelerator 205 through the system bus 203.
[0038] One or more of the processor 210, the memory 201, and the hardware accelerator 205 may be configured as a system-on-chip (SoC), a system-on-a-chip (SiP), or another configuration. In some embodiments, two or more of the processor 210, the memory 201, and the hardware accelerator 205 may be elements of the same substrate, e.g., elements of an interposer or printed circuit board (PCB), while in other embodiments, each of the processor 210, the memory 201, and the hardware accelerator 205 may be on a different substrate. In some embodiments, two or more of the memory 201, the processor 210, and the hardware accelerator 205 may be logical or physical divisions within a single semiconductor or a single chip, while in other embodiments, each of the memory 201, the processor 210, and the hardware accelerator 205 may be physically separate from each other.
[0039] The hardware accelerator 205 may be configured to execute jobs related to a number of TCBs in a linked list of TCBs stored in the memory 210. The respective TCBs of the linked list of TCBs may, for example, be similar to TCB 105. The linked list of TCBs is Fig. 2 is depicted as having previously executed TCB 214, a currently executed TCB 218, and TCB 220 yet to be executed.
[0040] A TCB 212 may be identified or generated by processor 210. Processor 210 may identify that the TCB is to be inserted into the linked list of TCBs. Providing TCB 212 to hardware accelerator 205 may then include a number of messages to be exchanged between the processor and hardware accelerator 205. In particular, if hardware accelerator 205 is currently executing a job related to a queued TCB, it may not be desirable to add a TCB to the queue because modification of the TCB by either processor 210 or hardware accelerator 205 may lead to consistency issues. Processor 210 at 236 may send a halt command to hardware accelerator 205. The halt command may include, for example, writing a bit to a register of hardware accelerator 205.The pause command at 235 may indicate to the hardware accelerator 205 that the hardware accelerator 205 should pause processing as soon as feasible. In some embodiments, the pause may require a number of cycles in the case where, for example, the hardware accelerator 205 may have initiated a memory access operation and is thereafter awaiting completion of the memory access operation. In response to the pause command, the hardware accelerator 205 may send a pause acknowledgment at 238 at 236. The pause acknowledgment at 238 may indicate that the hardware accelerator 205 has paused the operation and is awaiting further instructions from the processor 210. The processor 210 may then insert the TCB 212 into the TCB queue in the memory 210.For example, placing a TCB in the queue at a predetermined location relative to the other TCBs may involve modifying the FWD_CP-104 and BACK_CP-106 fields in one or more of the TCBs in the linked list of TCBs.
[0041] Continuing with the TCB queue, it can be seen that a number of previously executed TCBs 214 are present. The previously executed TCBs 214 may include two high-priority TCBs 208 (i.e., TCBs with a PRIO subfield indicating that they are to be executed in accordance with the "high" priority). The previously executed TCBs 214 may also include a medium-priority TCB 216 (i.e., a TCB with a PRIO subfield indicating that it is to be executed after the high-priority TCBs in accordance with the "medium" priority). Further, the hardware accelerator 205 may execute a job related to another medium-priority TCB 218. The TCBs 220 to be executed may include a medium priority TCB 224 and a low priority TCB 228 (i.e., a TCB with a PRIO subfield indicating that it is to be executed in accordance with the "low" priority after the high priority and medium priority TCBs).
[0042] The TCBs with the dashed borders indicate different locations where TCB 212 can be placed in the queue according to its priority. As mentioned, it is not possible to pause a job that refers to a currently executing TCB. Therefore, if TCB 212 is a high-priority TCB, it would be placed at the location indicated by TCB 222. Thus, it would be placed in the queue so that it would be processed, and any jobs that refer to it would be executed before the execution of any other jobs that refer to medium-priority TCBs, such as TCB 224. In other words, because the currently executing TCB 218 is a medium-priority TCB, the high-priority TCB 212 would be next in the queue.
[0043] However, if TCB 212 is a medium priority TCB, then it would be placed at the queue location indicated by TCB 226. This location would be at the end of the medium priority TCB, but before any low priority TCBs. Finally, if TCB 212 is a low priority TCB, then it would be placed at the queue location indicated by TCB 230. This location would be at the end of the low priority TCB.
[0044] The following table, Table 1, provides various examples of how priority-based task insertion may be achieved in an embodiment with three priority levels (high, medium, and low). However, it should be understood that this table is intended as a high-level example, and other embodiments may include variations of this table. Table 1 Priority level of the currently expiring TCB 210 Priority level of the TCB to be inserted, TCB 212 Place of insertion High High Insert TCB 212 after the last high priority TCB High Medium Insert TCB 212 after the last TCB with medium priority level High Low Insert TCB 212 after the last TCB with low priority level Medium High Insert TCB 212 immediately after TCB 218 Medium Medium Insert TCB 212 after the last TCB with medium priority level Medium Low Insert TCB 212 after the last TCB with low priority level Low High Insert TCB 212 immediately after TCB 218 Low Medium Insert TCB 212 immediately after TCB 218 Low Low Insert TCB 212 after the last TCB with low priority level
[0045] It is to be understood that this example of Fig. 2 is intended as a simplified example of priority-based queue management and other examples may vary from the illustrated example. For example, some embodiments may have more or fewer priority levels than in Fig. 2. The number of priority levels may depend on the PRIO subfield 128 discussed above. In addition, the number of TCBs at different priority levels may vary in different use cases.
[0046] In addition, the use case of Fig.2 as inserting a TCB 212 into the queue. However, it should be understood that other modifications to the queue may be possible. Such modifications may include deleting TCBs from the queue, querying the status of a job related to a TCB in the queue, modifying a TCB in the queue, etc. In general, queue management may be performed based on various APIs that may be used by processor 210 to manage the queue. Fig.3-7 depict various APIs that may be used for queue management. Such queue management by processor 210 may include one or more of the following: creating a "job" (e.g., positioning a TCB in memory), submitting the job to the queue, inserting the job at a specified location in the queue relative to other existing TCBs, deleting the job from the queue, determining the completion status of a job, etc. To perform these management tasks, processor 210 may need to store specific queue-related information, such asthe memory reference of the first TCB in the queue, the first TCB in the queue at a given priority level, the last TCB in the queue, the last TCB in the queue at a given priority level, or a combination thereof, and is capable of traversing the linked list of TCBs using the two chain pointer fields in each TCB (e.g., the chain pointer field 104 and the BACK_CP field 108).
[0047] Generally describes Fig. 3 an API that can be used to submit a job to the queue. The API of Fig. 3 may, for example, be executed by processor 210. The syntax of the submit command referring to the API may be: submit_job(ACC_TCB TCB_x, ACC_TCB TCB_n).
[0048] "ACC_TCB TCB_x" may contain information relating to a TCB to be inserted into the queue. The TCB may be named "TCB_x." The information may be a pointer to TCB_x, or it may be all or a portion of the TCB. Similarly, "ACC_TCB TCB_n" may contain information relating to a TCB in the queue after which TCB_x is to be placed into the queue. The information may be a pointer to TCB_n, or it may be all or a portion of the TCB.
[0049] Initially, the API may include checking at 305 whether the hardware accelerator (e.g., hardware accelerator 205) is enabled. As used herein, "enabled" may refer to whether the hardware accelerator is configured to accept TCBs / execute jobs related to the TCB. If the hardware accelerator is not enabled, then the processor may provide the hardware accelerator 205 with an indication of the first TCB to be processed in the linked list of TCBs. In particular, the processor may update a chain pointer register CP within the hardware accelerator with the address of the first TCB in the queue. The hardware accelerator may then be enabled at 315 so that it will process the TCB and execute the job with which the TCB is associated. In some embodiments, the TCB submission API depicted at 305 may then be restarted.
[0050] However, if the hardware accelerator is identified as enabled at 305, it may be desirable for the processor to stop the hardware accelerator at 320. Stopping the hardware accelerator may, for example, involve sending a message with reference to element 236 of Fig.2. The processor may wait for acknowledgment that the accelerator has paused by monitoring the pause acknowledgment described above with reference to element 238. The processor may then determine at 325 whether the hardware accelerator is free. As used herein, "free" may refer to whether the hardware accelerator has finished processing all TCBs in an existing queue of TCBs. If the hardware accelerator is determined to be free at 325, then the hardware accelerator may be disabled at 330 by the processor by writing one or more bit fields to the configuration registers within the hardware accelerator. This disabling may serve as a "reset" of the hardware accelerator.The chain pointer register CP within the hardware accelerator may be updated with the address of the TCB at 335 by the processor, and then the hardware accelerator may be re-enabled at 340. Re-enabled of the hardware accelerator may include, for example, changing the bit in the register indicating that the hardware accelerator should stop, as described above with reference to element 236. Once the hardware accelerator is re-enabled at 340, the hardware accelerator may resume its functions, and in addition, the processor may resume executing application code related to the API as described herein, another API, or another processing function.
[0051] However, if it is determined at 325 that the hardware accelerator is not free, then at 345 the processor may execute an API related to inserting the TCB into the TCB queue. The syntax of the insert command related to the API may be: insert_job(ACC_TCB TCB_x, ACC_TCB TCB_n). "ACC_TCB TCB_x" and "ACC_TCB TCB_n" may be similar to those described above with respect to Fig. 3 are described. Fig. 4 depicts an example of the insert API. It should be understood that in other embodiments, the insert_job API may be separate from the API of Fig. 3 can be executed.
[0052] Initially, the API may include identifying at 405 whether the pointer to TCB_n in the insert_job API is a NULL (0) value. If TCB_n is not a NULL value, then the API may include identifying at 410 whether TCB_n has completed. In particular, the API may include identifying at 410 whether the job referencing TCB_n has completed, or whether TCB_n is in the list of previously executed TCBs described at 214. If so, then the hardware accelerator may provide a fault indication to the processor at 415. The fault indication may be due to the fact that a TCB inserted after a TCB whose processing has completed will not be accessed by the hardware accelerator, and thus the associated job may never be executed. The hardware accelerator or the processor may resume operation as described above.
[0053] However, if TCB_n is not identified as completed at 410, then the processor may locate TCB_n in the list of TCBs at 420. Specifically, the processor may query one or more of the TCBs in the list of TCBs until TCB_n is identified. The processor may then update the TCB connection chain at 425. Updating the connection chain may include, for example, updating the chain pointer of the TCB to be inserted, updating the chain pointers of one or more TCBs in the chain, updating a hardware accelerator buffer, updating the processor, etc. The updates may serve the purpose of correctly identifying TCB_n. The processor and hardware accelerator may then resume operation as described above.
[0054] If at 405 the value of the pointer to TCB_n in the API is a NULL value, this may indicate that TCB_x is to be inserted into the queue based on its priority, as specified in the PRIO field 128 of the TCB structure 105. As a result, the processor may then perform priority-based insertion of the TCB at 430. Priority-based insertion may include using information about a first or last TCB in a given priority bin, or queue in general, to identify a location where the TCB should be inserted. As an example, priority-based insertion at 435 may include traversing a priority table to insert the new job or new TCB into the list of TCBs.For example, the processor may have or have access to a table that refers to the list of TCBs, specifying a priority level for each TCB, or an indication of a first or last TCB of a given priority bin. In some embodiments, the processor may have or have access to tags or flags that indicate a last TCB of a given priority level. These flags or tags may help reduce the period of time the system or hardware accelerator is halted while TCB management is performed. Alternatively, the processor 210 may maintain the chain of TCBs in memory 201 in . Fig.2, starting from the first TCB in the queue, the currently executing TCB, or another TCB, using the TCB's chain pointer fields 104 or 108. The traversal may, for example, be used to identify a last TCB in a given priority section. It may determine the priority level of each TCB by reading the TCB's PRIO field 128. It should be understood that although a priority table is described herein, in other embodiments, the data structure relating to the priority-based organization of the TCB may be different.
[0055] The API can then return at 440 and begin again. Generally, APIs can be called by a higher-level application running on the processor. When an API returns, the processor can "resume" executing the remaining portion of the application code. The hardware accelerator can also resume its operation from the "paused" state, as described above.
[0056] The processor 210 may traverse the linked list (e.g., using information about a last TCB of a given priority bin or by another technique) to identify a location where the TCB_x should be placed, as described above with reference to Fig.2. Once the location is identified, the TCB can be inserted into the list of TCBs. The API can then indicate that the TCB has been successfully inserted into the list of TCBs, for example, by returning a SUCCESS flag to the higher-level application that may have required or initiated the TCB, and resume performing other functions of the higher-level application.
[0057] In other embodiments, it may be desirable to modify an existing job or TCB. The modification may include, for example, replacing an existing TCB in the list of TCBs with another TCB that has different coefficients, a pointer to a different input buffer, etc. Fig.Figure 5 depicts an example API that can be used for a job or TCB modification command. The API can be executed by the processor. The syntax of the modification command that refers to the API can be: modifyjob(ACC_TCB TCB_n, ACC_TCB TCB_o). "ACC_TCB TCB_n" can contain information related to a new TCB ("TCB_n") to be inserted into the queue. "ACC_TCB TCB_o" can contain information related to a TCB ("TCB_o") to be removed from the queue. In particular, TCB_n can be a TCB that should be inserted into the queue instead of TCB_o. In this way, an existing TCB (TCB_o) can be modified by replacing that TCB with a new TCB (TCB_n) that has updated information.
[0058] The API may include pausing the hardware accelerator at 505. The pausing may occur in response to a pausing command, such as the pausing command described above with reference to element 236. The API may then include querying the status of a given job at 510. In particular, the processor 210 may identify at 510 whether a job is pending, completed, expiring, etc. In some embodiments, the query at 510 may include running a query API, which is described below with reference to Fig. 7 is discussed in more detail.
[0059] The API may then include determining at 515 whether the result of the query at 510 indicates that TCB_o is not yet completed. A "not yet completed" TCB may refer to a TCB that is queued for future execution / processing, and thus, a job related to that TCB is not currently being executed by the hardware accelerator or has not already been executed. In other words, a not yet completed TCB may be in the queue of TCBs to be executed at 220, as described with reference to Fig.2. If TCB_o is not yet completed, the API may return an error at 520. The error condition may exist because, if TCB_o is not yet completed, it may not be possible to modify or replace the job related to TCB_o by replacing it with TCB_n because that job may have already been executed. After the error condition is identified and returned, the hardware accelerator's operation may restart, for example, by processing the next TCB in the TCB queue.
[0060] However, if a query at 510 returns at 515 that the TCB_o is not yet completed, then the API may include finding the TCB_o in the list of uncompleted TCBs at 525. For example, the processor 210 may identify the location or address of the TCB_o in the TCBs to be executed 220. Once the TCB_o is found, the processor may pop out the TCB_o and place the TCB_n in its place at 530. Such a pop out operation may include, for example, modifying the chain pointer FWD_CP field of the TCB that precedes the TCB_o in the queue and updating that field with the memory address of the TCB_n. The pop out operation may further include updating the BACK_CP field of the TCB_n with the same value in the BACK_CP field of the TCB_o.The detach operation may also include updating the FWD_CP field of TCB_n with the same value as the FWD_CP field of TCB_o and updating the BACK_CP field of the TCB referenced by TCB_n's FWD_CP with the memory address of TCB_n. In some embodiments, the API may also release the memory used to store TCB_o and return the memory to the free memory block pool. Finally, at 535, the API may include returning a status flag that TCB_o has been replaced by TCB_n, and the hardware accelerator may restart operation. For example, a flag may be sent to another logic or processing element of the hardware accelerator or to the processor indicating that the replacement is complete, and then the hardware accelerator may resume processing TCB in the linked list of TCBs.For example, the return and resume function of element 535 may be similar to the return and resume function of element 440.
[0061] In some embodiments, it may be desirable to remove a TCB in the linked list of TCBs rather than replacing it. Fig. 6 depicts an API that may be used by a component of an electronic device, such as the processor 210, to remove a TCB from the list of TCBs to be executed 220. The syntax of the API may be: delete_job(ACC_TCB TCB_d), where TCB_d may be a TCB to be discarded, and "ACC_TCB TCB_d" may contain information related to the TCB_d.
[0062] The API in Fig. 6 may comprise one or more elements similar to elements of Fig.5 or have one or more properties in common with them. In particular, elements 605, 610, 615, 620, 625, and 635 may each be similar to elements 505, 510, 515, 520, 525, and 535 of Fig. 5 or have one or more properties in common with them. Details of these elements are not repeated to avoid redundancy.
[0063] However, as can be seen, the API of Fig.6, after TCB_d is found in the list of TCBs to be executed 220, at 630, detaching TCB_d and joining the TCBs in the list that were before and after TCB_d. In particular, the FWD_CP field of the TCB that precedes TCB_d may be modified to point to the TCB that follows TCB_d. Similarly, the BACK_CP field of the TCB that follows TCB_d may be modified to point to the TCB that precedes TCB_d. In this way, TCB_d may be effectively removed from the list of TCBs to be executed at 220. In some embodiments, the API may also release the memory used to store TCB_d and return the memory to the pool of free memory blocks.
[0064] As mentioned, in some embodiments, it may be desirable to query the status of a TCB in the linked list of TCBs. For example, application code running on processor 210 may request a status update for a specific TCB or job from the hardware accelerator. As another example, the APIs of the Fig. 5 or Fig. 6 issue a query at elements 510 or 610. Fig. 7 shows a sample API that can be used in response to such a query. The API of Fig. 7 may be executed by processor 210. The query syntax may be, for example, query_job(ACC_TCB TCB_q). "ACC_TCB TCB_q" may contain information such as a pointer and address, etc., relating to a TCB ("TCB_q").
[0065] It should be noted that in some embodiments, the APIs of one or more of the Fig.3-6 can affect the last TCB at a given priority level. For example, a new TCB with a "high" priority can be inserted after the existing last TCB with a "high" priority. Alternatively, the last TCB with a given priority can be replaced or removed. In this situation, various updates can be performed on the various chain pointers, linked lists of TCBs, etc., to accurately reflect the updated "last TCB" of that priority level.
[0066] The API may include stopping the hardware accelerator at 705. Stopping the hardware accelerator may be similar to the stopping described with respect to elements 320, 505, 605, etc. The API may then include identifying whether the TCB_q exists at 710. In other words, the API may identify whether the TCB_q can be located in the linked list of TCBs. If it is found at 710 that the TCB_q does not exist, then the API may include returning an error at 715 and resuming the operation, which may, for example, be similar to element 440 above.
[0067] If the TCB_q is identified at 710, then the API may include identifying whether the TCB_q is completed at 720. In other words, the API may include identifying whether the TCB_q is in the previously executed TCB 214. Additionally or alternatively, identifying whether the TCB_q is completed may include checking a completion flag related to, or part of, the TCB_q. Other embodiments may include other flags related to the completion of the TCB. If so, then the API may include returning an indication that the TCB_q is completed at 725 and resuming the operation, as described above, for example, with reference to element 440.
[0068] If it is identified at 720 that the TCB_q has not been completed, then the API may include identifying at 730 whether the TCB_q is expiring. As used herein, "expiring" may indicate whether the TCB was in the active state being processed or the job to which the TCB refers was executed by the hardware accelerator before the hardware accelerator was stopped. If it is identified at 730 that the TCB_q is expiring, then the API may include returning at 735 an indication that the TCB_q is expiring and resuming operation, for example, as described above with reference to element 440. However, if at 730 it is not identified that the TCB_q is expiring, then at 740 the API may include returning an indication that the TCB_q is not done and resuming the operation, as described above, for example, with reference to element 440.
[0069] It should be understood that the APIs described above are provided as examples of one embodiment, and other embodiments may vary from those described. For example, other embodiments may have more or fewer elements, different syntax, etc. In some embodiments, specific elements may be presented in a different order than that described in the Fig. 3-7. In addition, although particular APIs or elements thereof may be described as being executed by the hardware accelerator or the processor, in some embodiments, one or more of the APIs or elements thereof described as being executed by the hardware accelerator may be executed by the processor, and vice versa. Other variations may be present in other embodiments.
[0070] Fig.Figure 8 is a block diagram of an exemplary electrical device 1800 that may include one or more hardware accelerators, in accordance with any of the embodiments disclosed herein. A number of components are shown in Fig. 8 as being included in the electrical device 1800, however, one or more of these components may be omitted or duplicated as appropriate for the application. In some embodiments, some or all of the components included in the electrical device 1800 may be connected to one or more motherboards. In some embodiments, some or all of these components are fabricated on a single SoC package.
[0071] Additionally, in various embodiments, the electrical device 1800 may include one or more of the Fig.8, but the electrical device 1800 may include interface circuitry for coupling to the one or more components. For example, the electrical device 1800 may not include a display device 1806, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 1806 may be coupled. In another group of examples, the electrical device 1800 may not include an audio input device 1824 or an audio output device 1808, but may include audio input or audio output interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 1824 or an audio output device 1808 may be coupled.
[0072] The electrical device 1800 may include a processing device 1802 (e.g., one or more processing devices). As used herein, the term "processing device" or "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory. The processing device 1802 may include one or more DSPs, ASICs, CPUs, graphics processing units (GPUs), cryptoprocessors (specialized processors that execute cryptographic algorithms within hardware), server processors, or any other suitable processing devices. The electrical device 1800 may include a memory 1804, which may itself include one or more storage devices, such asvolatile memory (e.g., dynamic random access memory (DRAM)), non-volatile memory (e.g., read-only memory (ROM)), flash memory, solid-state memory, and / or a hard disk drive. In some embodiments, memory 1804 may include memory that shares a building block with processing device 1802. This memory may be used as cache memory and may include embedded dynamic random access memory (eDRAM) or spin momentum transfer magnetic random access memory (STT-MRAM). For example, processing device 1802 may be similar to processor 210 and may be coupled to a hardware accelerator, such as hardware accelerator 205.
[0073] In some embodiments, the electrical device 1800 may include a communication chip 1812 (e.g., one or more communication chips). For example, the communication chip 1812 may be configured to manage wireless communication for the transmission of data to and from the electrical device 1800. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc., that can communicate data using modulated electromagnetic radiation over a non-solid medium. The term does not imply that the associated devices do not include wires, although this could be the case in some embodiments.
[0074] The communication chip 1812 may implement any of a number of wireless standards or protocols, including, but not limited to, Institute for Electrical and Electronic Engineers (IEEE) standards, Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards (e.g., IEEE 802.16-2005 amendment), Long Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., LTE evolved project, Ultra Mobile Broadband (UMB) project (also referred to as "3GPP2"), etc.). IEEE 802.16 compliant Broadband Wireless Access (BWA) networks are commonly referred to as WiMAX networks, an acronym that stands for Worldwide Interoperability for Microwave Access, which is a certification mark for products that pass tests for conformance and interoperability to the IEEE 802.16 standards.The communication chip 1812 can operate in accordance with the Global System for Mobile Communications (GSM), the General Packet Radio Service (GPRS), the Universal Mobile Telecommunications System (UMTS), the High Speed Packet Access (HSPA), the Evolved HSPA (E-HSPA), or LTE networks. The communication chip 1812 can operate in accordance with Enhanced Data for GSM Evolved (EDGE), the GSM EDGE Radio Access Network (GERAN), the Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). The communication chip 1812 can operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless (DECT), Evolved Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols designated as 3G, 4G, 5G, and beyond.In other embodiments, the communication chip 1812 may operate in accordance with other wireless protocols. The electrical device 1800 may include an antenna 1822 for supporting wireless communication and / or for receiving other wireless communications (such as AM or FM radio broadcasts).
[0075] In some embodiments, the communication chip 1812 may manage wired communication, such as electrical, optical, or any other suitable communication protocols (e.g., Ethernet). As mentioned above, the communication chip 1812 may include multiple communication chips. For example, a first communication chip 1812 may be dedicated to shorter-range wireless communication, such as Wi-Fi or Bluetooth, and a second communication chip 1812 may be dedicated to longer-range wireless communication, such as the Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 1812 may be dedicated to wireless communication, and a second communication chip 1812 may be dedicated to wired communication.
[0076] The electrical device 1800 may include a battery / power supply circuitry 1814. The battery / power supply circuitry 1814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the electrical device 1800 to a power source separate from the electrical device 1800 (e.g., AC power line).
[0077] The electrical device 1800 may include a display device 1806 (or corresponding interface circuitry, as discussed above). The display device 1806 may include any visual indicator, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0078] The electrical device 1800 may include an audio output device 1808 (or corresponding interface circuitry, as discussed above). The audio output device 1808 may include any device that produces an audible indicator, such as speakers, headphones, or earbuds.
[0079] The electrical device 1800 may include an audio input device 1824 (or corresponding interface circuitry, as discussed above). The audio input device 1824 may include any device that generates a signal representative of sound, such as microphones, microphone assemblies, or digital instruments (e.g., instruments having a digital musical instrument interface (MIDI) output).
[0080] The electrical device 1800 may include a GPS device 1818 (or corresponding interface circuitry, as discussed above). The GPS device 1818 may be in communication with a satellite-based system and may receive a location of the electrical device 1800, as is known in the art.
[0081] The electrical device 1800 may include another output device 1810 (or corresponding interface circuitry, as discussed above). Examples of the other output device 1810 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.
[0082] The electrical device 1800 may include another input device 1820 (or corresponding interface circuitry, as discussed above). Examples of the other input device 1820 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touch-sensitive contact surface, a barcode reader, a quick response (QR) code reader, or a radio frequency identification (RFID) reader.
[0083] Electrical device 1800 may have any desired form factor, such as a portable or mobile electrical device (e.g., a cellular phone, a smartphone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA), an ultramobile personal computer, etc.), a desktop electrical device, a server device or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital video recorder, or a wearable electrical device. In some embodiments, electrical device 1800 may be any other electronic device that processes data. EXAMPLES OF DIFFERENT EMBODIMENTS
[0084] Example 1 includes a method comprising: identifying, by a component of a computing device, a TCB related to a job to be executed by the component; identifying, based on the TCB, a priority level of the job; inserting the job into a queue of jobs to be executed based on the identified priority level; and executing the job based on the identified priority level.
[0085] Example 2 includes the method of Example 1 or any other example herein, wherein the component is a core or a hardware accelerator of the computing device.
[0086] Example 3 includes the method of Example 1 or any other example herein, where the job is a job related to a DMA request.
[0087] Example 4 includes the method of any of Examples 1-3 or any other example herein, wherein the job is a job having a higher priority level and queuing the job having the higher priority level includes queuing the job having the higher priority level so that it executes before executing a job having a lower priority level.
[0088] Example 5 includes the method of Example 4 or any other example herein, wherein the job having the lower priority level was identified by the component of the computing device prior to the identification of the TCB.
[0089] Example 6 includes the method of Example 4 or any other example here, where the job with the lower priority level was inserted into the queue before the job with the highest priority level was inserted.
[0090] Example 7 includes the method of any of Examples 1-3 or any other example herein, wherein inserting the job into the queue of jobs to be executed is based on the performance of an API related to the jobs.
[0091] Example 8 includes the method of any of Examples 1-3 or any other example herein, wherein the queue of jobs to be executed includes a queue having four priority levels.
[0092] Example 9 includes one or more non-transitory computer-readable media comprising instructions that, when executed by one or more elements of a computing device, are to cause a component of the computing device to: identify a TCB related to a job to be executed by the component; identify, based on the TCB, a priority level of the job; insert the job into a queue of jobs to be executed based on the identified priority level; and execute the job based on the identified priority level.
[0093] Example 10 includes one or more non-transitory computer-readable media according to Example 9 or any other example herein, wherein the component is a core or a hardware accelerator of the computing device.
[0094] Example 11 includes one or more non-transitory computer-readable media according to Example 9 or any other example herein, where the job is a job related to a DMA request.
[0095] Example 12 includes one or more non-transitory computer-readable media according to any one of Examples 9-11 or any other example herein, wherein the job is a job with a higher priority level and queuing the job with the higher priority level includes queuing the job with the higher priority level so that it executes before executing a job with a lower priority level.
[0096] Example 13 includes one or more non-transitory computer-readable media according to Example 12 or any other example herein, wherein the job having the lower priority level was identified by the component of the computing device prior to identifying the TCB.
[0097] Example 14 includes one or more non-transitory computer-readable media according to Example 12 or any other example herein, wherein the job with the lower priority level was inserted into the queue before the job with the highest priority level was inserted.
[0098] Example 15 includes one or more non-transitory computer-readable media according to any of Examples 9-11 or any other example herein, wherein the insertion of the job into the queue of jobs to be executed is based on the performance of an API related to the jobs.
[0099] Example 16 includes one or more non-transitory computer-readable media according to any one of Examples 9-11 or any other example herein, wherein the queue of jobs to be executed includes a queue having four priority levels.
[0100] Example 17 includes an apparatus comprising: means for identifying a TCB related to a job to be executed by the component; means for identifying, based on the TCB, a priority level of the job; means for inserting the job into a queue of jobs to be executed based on the identified priority level; and means for executing the job based on the identified priority level.
[0101] Example 18 includes the apparatus of Example 17 or any other example herein, wherein the component is a core or a hardware accelerator of the computing device.
[0102] Example 19 includes the facility of Example 17 or any other example here, where the job is a job related to a DMA request.
[0103] Example 20 includes the apparatus of any of Examples 17-19 or any other example herein, wherein the job is a job having a higher priority level, and means for inserting the job having the higher priority level into the queue comprises means for inserting the job having the higher priority level into the queue so that it is executed before a job having a lower priority level is executed.
[0104] Example 21 includes the apparatus of Example 20 or any other example herein, wherein the job having the lower priority level was identified by the component of the computing device prior to the identification of the TCB.
[0105] Example 22 contains the setup of Example 20 or any other example here, where the job with the lower priority level was inserted into the queue before the job with the highest priority level.
[0106] Example 23 includes the setup of any of Examples 17-19 or any other example here, where insertion of the job into the queue of jobs to be executed is based on the performance of an API related to the jobs.
[0107] Example 24 includes the facility of any of Examples 17-19 or any other example herein, wherein the queue of jobs to be executed includes a queue with four priority levels.
[0108] Example 25 includes a device comprising: a memory; and a component communicatively coupled to the memory, the component to: identify a TCB related to a job to be executed by the component; identify, based on the TCB, a priority level of the job; insert the job into a queue of jobs to be executed based on the identified priority level; and execute the job based on the identified priority level.
[0109] Example 26 includes the apparatus of Example 25 or any other example herein, wherein the component is a core or a hardware accelerator of the computing device.
[0110] Example 27 includes the facility of Example 25 or any other example here, where the job is a job related to a DMA request.
[0111] Example 28 includes the facility of any of Examples 25-27 or any other example herein, where the job is a job with a higher priority level and placing the job with the higher priority level in the queue includes placing the job with the higher priority level in the queue so that it executes before a job with a lower priority level executes.
[0112] Example 29 includes the apparatus of Example 28 or any other example herein, wherein the job having the lower priority level was identified by the component of the computing device prior to the identification of the TCB.
[0113] Example 30 contains the setup of Example 28 or any other example here, where the job with the lower priority level was inserted into the queue before the job with the highest priority level.
[0114] Example 31 includes the setup of any of Examples 25-27 or any other example here, where insertion of the job into the queue of jobs to be executed is based on the performance of an API related to the jobs.
[0115] Example 32 includes the facility of any of Examples 25-27 or any other example herein, wherein the queue of jobs to be executed includes a queue with four priority levels.
[0116] Example 33 includes a computing device comprising: a processor to: identify in a TCB an indication of a priority level of the TCB; identify, based on the indication of the priority level, a location in a queue of one of a plurality of TCBs; and insert the TCB into the queue at the identified location; and a hardware accelerator communicatively coupled to the processor, the hardware accelerator to execute jobs related to TCBs in the queue of the plurality of TCBs, the hardware accelerator to execute a job related to a TCB in accordance with an ordering of the TCB in the queue of the plurality of TCBs.
[0117] Example 34 includes the computing device of Example 33 or any other example herein, wherein the priority level indication is a 2-bit indicator in the TCB.
[0118] Example 35 includes the computing device of Example 33 or any other example herein, wherein the processor is further to send a stop request before sending the TCB.
[0119] Example 36 includes the computing device of any of Examples 33-35 or any other example herein, wherein the queue comprises a first portion of TCB at a first priority level and a second portion of TCB at a second priority level lower than the first priority level, and wherein the hardware accelerator is to execute jobs related to the first portion of the TCB before the hardware accelerator executes jobs related to the second portion of the TCB.
[0120] Example 37 includes the computing device of Example 36 or any other example herein, wherein the TCB is in the first priority level, and wherein the processor is further to: identify a priority level of a currently executing TCB; and identify the location in the queue based on the priority level of the currently executing TCB.
[0121] Example 38 includes the computing device of Example 37 or any other example herein, wherein the priority level of the currently running TCB is the first priority level and wherein the location is at the end of the first subsection of the TCB.
[0122] Example 39 includes the computing device of Example 37 or any other example herein, wherein the priority level of the currently executing TCB is the second priority level and wherein the location is after the currently executing TCB and before other TCBs of the second subsection of TCB.
[0123] Example 40 includes a processor having: a communication interface to communicate with a hardware accelerator; and logic coupled to the communication interface, the logic to: identify in a TCB an indication of a priority level of the TCB; and insert a TCB into a location in a queue of a plurality of TCBs based on the indication of the priority level; wherein the hardware accelerator is to execute corresponding ones of the plurality of TCBs based on their location within the queue.
[0124] Example 41 includes the processor of Example 40 or any other example herein, where the priority level indication is a 2-bit indicator in the TCB.
[0125] Example 42 includes the processor of Example 40 or any other example herein, where the priority level indication indicates that the TCB is at one of several possible priority levels.
[0126] Example 43 includes the processor of any of Examples 40-42 or any other example herein, wherein the TCB further comprises an indication that the hardware accelerator should wait to receive a trigger from the processor before processing the TCB, an indication that the hardware accelerator should send an interrupt related to completion of the TCB to the processor, or an indication that the hardware accelerator should send a trigger related to completion of the TCB to the processor.
[0127] Example 44 includes the processor of any of Examples 40-42 or any other example herein, wherein the TCB further comprises an indication of a type of computation to be used to perform the job related to the TCB.
[0128] Example 45 includes the processor of any of Examples 40-42 or any other example herein, wherein the TCB further comprises an indication of a previous TCB in the queue of the plurality of TCBs.
[0129] Example 46 includes one or more non-transitory computer-readable media comprising instructions that, when executed by an electronic device, are to cause a processor of the electronic device to: identify a queue of a plurality of TCBs, wherein the respective TCBs of the plurality of TCBs include an indication of a priority level of the respective TCB within the queue, and wherein a hardware accelerator is to execute a job related to a corresponding TCB in accordance with the priority level of the TCB; execute an application programming interface (API) related to the queue; and modify the queue based on execution of the API.
[0130] Example 47 includes the one or more non-transitory computer-readable media of Example 46 or any other example herein, wherein the API relates to inserting a TCB into the queue of the plurality of TCBs based on an indication of a priority level of the TCB.
[0131] Example 48 includes the one or more non-transitory computer-readable media of Example 47 or any other example herein, wherein the insertion of the TCB refers to a flag indicating a last TCB at a priority level.
[0132] Example 49 includes the one or more non-transitory computer-readable media of Example 46 or any other example herein, where the API relates to dequeuing a TCB.
[0133] Example 50 includes the one or more non-transitory computer-readable media of Example 46 or any other example herein, wherein the API relates to a modification of a TCB of the queue.
[0134] Example 51 includes the one or more non-transitory computer-readable media of Example 46 or any other example herein, wherein the API relates to identifying an execution status of a TCB within the queue.
[0135] Example 52 includes the one or more non-transitory computer-readable media of Example 46, wherein the API relates to a hardware accelerator reset.
[0136] Example 53 includes a method for performing the subject matter of any of Examples 1-52, or a subset or combination thereof.
[0137] Example 54 includes one or more non-transitory computer-readable media having instructions that, when executed by one or more elements of a computing device, are operable to cause a component of the computing device to perform the subject matter of any of Examples 1-52, or a subset or combination thereof.
[0138] Example 55 includes a device having circuitry for performing or causing to be performed the subject matter of any of Examples 1-52, or a subset or combination thereof.
[0139] Example 56 includes an apparatus comprising means for performing or causing to be performed the subject matter of any of Examples 1-52, or a subset or combination thereof.
[0140] Various embodiments may include any suitable combination of the embodiments described above, including alternative (or) embodiments of embodiments described above in conjunction (and) form (e.g., the "and" may be an "and / or"). Furthermore, some embodiments may include one or more articles of manufacture (e.g., non-transitory computer-readable media) having instructions stored thereon that, when executed, result in actions of any of the embodiments described above. Furthermore, some embodiments may include devices or systems having any suitable means for performing the various operations of the embodiments described above.
[0141] The foregoing description of illustrated embodiments, including what is described in the abstract, is not intended to be exhaustive or to be limited to the precise forms disclosed. Although specific implementations of and examples of various embodiments or concepts are described herein for illustrative purposes, various equivalent modifications may be possible, as will be appreciated by those skilled in the relevant art. These modifications may be made in light of the foregoing detailed description, the abstract, the figures, or the claims.
[0142] According to one aspect, embodiments may relate to an electronic device comprising a processor communicatively coupled to a hardware accelerator. The processor may be configured to identify, based on an indication of a priority level in a TCB, a location at which the TCB should be inserted into a queue of TCBs. The hardware accelerator may execute jobs related to the queue of TCBs in an order related to the order of TCBs within the queue. Other embodiments may include fewer features and may be described or claimed.
Claims
[1] Calculation device comprising: a processor (210) for: Identifying, in a task control block (218), an indication of a priority level (202; 204; 206) of the task control block (218); Identifying, based on the indication of the priority level (202; 204; 206), a location in a queue of multiple task control blocks (218), wherein each task control block (218) in the queue of multiple task control blocks (218) is specifically associated with a hardware accelerator (205) to execute a job associated with the corresponding task control block (218); and Inserting (345) the task control block (218) into the queue at the identified location, the inserted task control block (218) having a subfield indicating whether completion of the job associated with the inserted task control block (218) should be signaled to the processor (210); wherein the hardware accelerator (205) is communicatively coupled to the processor (210), wherein the hardware accelerator (205) is for executing jobs related to task control blocks (218) in the queue of the plurality of task control blocks (218), wherein the hardware accelerator (205) is for executing a job related to a task control block in accordance with an order of the task control block (218) in the queue of the plurality of task control blocks (218). [2] The computing device of claim 1, wherein the indication of the priority level (202; 204; 206) is a 2-bit indicator in the task control block (218). [3] A computing device according to any one of the preceding claims, wherein the processor (210) is further operable to send a stop request before sending the task control block (218). [4] Computing device according to one of the preceding claims, wherein the queue comprises a first subsection of task control blocks (218) at a first priority level (202; 204) and a second subsection of task control blocks (218) at a second priority level (204; 206) which is lower than the first priority level (202; 204), and wherein the hardware accelerator (205) is for executing jobs relating to the first subsection of the task control blocks (218) before the hardware accelerator (205) executes jobs relating to the second subsection of the task control blocks (218). [5] The computing device of claim 4, wherein the task control block (218) is at the first priority level (202; 204) and wherein the processor (210) is further for: Identifying a priority level (202; 204; 206) of a currently executing task control block (218); and Identifying the location in the queue based on the priority level (202; 204; 206) of the currently executing task control block (218). [6] A computing device according to claim 5, wherein the priority level (202; 204; 206) of the currently executing task control block (218) is the first priority level (202; 204) and wherein the location is at the end of the first subsection of the task control block (218). [7] The computing device of claim 5, wherein the priority level (202; 204; 206) of the currently executing task control block (218) is the second priority level (204; 206) and wherein the location is after the currently executing task control block (218) and before other task control blocks (218) of the second subsection of task control blocks (218). [8] Processor (210) comprising: a communication interface for communicating with a hardware accelerator (205); and Logic coupled to the communication interface, where the logic is used to: Identifying in a task control block (218) an indication of a priority level (202; 204; 206) of the task control block (218); and Inserting, based on the specification of the priority level (202; 204; 206), the task control block (218) at a location in a queue of multiple task control blocks (218); wherein the hardware accelerator (205) is for executing corresponding ones of the plurality of task control blocks (218) based on their position within the queue, and wherein the inserted task control block (218) has a subfield indicating whether completion of the job associated with the inserted task control block (218) should be signaled to the processor (210). [9] The processor (210) of claim 8, wherein the priority level indication (202; 204; 206) is a 2-bit indicator in the task control block (218), and wherein the priority level indication (202; 204; 206) indicates that the task control block (218) is at one of a plurality of priority levels (202; 204; 206). [10] The processor (210) of claim 8 or 9, wherein the inserted task control block (218) comprises the subfield comprising: an indication that the hardware accelerator (205) should send an interrupt relating to the completion of the task control block (218) to the processor (210), or an indication that the hardware accelerator (205) should send a trigger related to the completion of the task control block (218) to the processor (210). [11] The processor (210) of any one of claims 8-10, wherein the task control block (218) further comprises an indication that the hardware accelerator (205) should wait to receive a trigger from the processor (210) before processing the task control block (218). [12] The processor (210) of any of claims 8-11, wherein the task control block (218) further comprises an indication of a type of computation to be used to execute the job related to the task control block (218). [13] The processor (210) of any of claims 8-12, wherein the task control block (218) further comprises an indication of a previous task control block (218) in the queue of the plurality of task control blocks (218). [14] One or more non-transitory computer-readable media comprising instructions that, when executed by an electronic device, cause a processor (210) of the electronic device to perform the following steps: Identifying a queue of multiple task control blocks (218), wherein the respective task control blocks (218) of the multiple task control blocks (218) have an indication of a priority level (202; 204; 206) of the respective task control blocks (218) within the queue, and a hardware accelerator (205) is used to execute a job related to a corresponding task control block (218) in accordance with the priority level (202; 204; 206) of the task control block (218), and wherein at least the first task control block (218) of the multiple task control blocks (218) has a subfield indicating whether completion of the job associated with the inserted task control block (218) should be signaled to the processor (210); Executing an application programming interface (API) related to the queue; and Modify the queue based on API execution. [15] One or more non-transitory computer-readable media according to claim 14, wherein the API relates to inserting a task control block (218) into the queue of the plurality of task control blocks (218) based on an indication of a priority level (202; 204; 206) of the task control block (218). [16] One or more non-transitory computer-readable media according to claim 15, wherein the insertion of the task control block (218) refers to a flag indicating a last task control block (218) at a priority level (202; 204; 206). [17] One or more non-transitory computer-readable media according to any one of claims 14-16, wherein the API relates to dequeuing a task control block (218). [18] One or more non-transitory computer-readable media according to any one of claims 14-17, wherein the API relates to the modification of a task control block (218) of the queue. [19] One or more non-transitory computer-readable media according to any one of claims 14-18, wherein the API relates to identifying an execution status of a task control block (218) within the queue. [20] One or more non-transitory computer-readable media according to any one of claims 14-19, wherein the API relates to resetting the hardware accelerator (205).
Citation Information
Patent Citations
System and method for thread scheduling with weak preemption policy
US20030208521A1
Multithreaded kernel for graphics processing unit
US20040160446A1
Task control means for a multi-tasking data processing system
US4658351A
Multithreaded processor for processing multiple instruction streams independently of each other by flexibly controlling throughput in each instruction stream
US6105127A