A new efficient memory representation of pass dependencies for hetrogeneous task execution

By decoupling workgraph descriptions from execution schedules using a pass dependency schema, the method enhances flexibility and efficiency in heterogeneous task execution on mobile System on Chip platforms, reducing latency and improving performance.

WO2025131261A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/086800
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current approaches to heterogeneous task execution in mobile System on Chip platforms tightly couple workgraph descriptions with their execution schedules, making it difficult to experiment with schedule changes and limiting performance and utility.

Method used

A device and method that separate and decouple workgraph descriptions from their underlying execution schedules, using a pass dependency schema that captures event information and cyclic properties, allowing for efficient memory encoding and dynamic scheduling.

Benefits of technology

This approach enables flexible adaptation of task processing without modifying the workgraph, improves scheduling efficiency, and reduces latency by allowing domain specific accelerators to execute passes independently of the host.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023086800_26062025_PF_FP_ABST
    Figure EP2023086800_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A device for processing an application pipeline, the device comprising one or more processor configured to: receive a workgraph comprising one or more pass to be performed for processing the application pipeline, each of the one or more pass comprising one or more task formed of one or more command for processing the application pipeline; generate a pass dependency schema representing an order in which the one or more pass in the workgraph is to be processed when the application pipeline is executed; wherein the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema comprises event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A NEW EFFICIENT MEMORY REPRESENTATION OF PASS DEPENDENCIES FOR HETROGENEOUS TASK EXECUTION

[0002] FIELD OF THE INVENTION

[0003] This present disclosure relates to a device for processing an application pipeline. The present disclosure also relates to a method of doing the same.

[0004] BACKGROUND

[0005] Modem mobile system on chip platforms routinely contain multiple Domain Specific Accelerators such as a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural Processing Unit (NPU), and Image Signal Processor (ISP) device which can be orchestrated to work together in processing a workload dataset and share system resources coherently during this process. A single domain specific accelerator device may support executing multiple types of domain operations such as a GPU that is able to execute both render and compute domain operations.

[0006] Such domain specific accelerators may be configured to execute application pipelines (AP) that may include of one or more heterogenous domain operations connected together in a particular order and are used to transform an input dataset using one or more of these domain operations such as compute, image processing, machine learning and / or graphics rendering techniques to generate an output dataset. In some cases, heterogeneous task (HT) execution (HTE) means utilization of different types of hardware processing units or domain specific accelerator such as CPU, NPU and GPU within a single system to execute application pipelines APs in order to achieve improved system performance and energy efficiency.

[0007] In order to execute application pipelines, it may be beneficial to utilise one or more workgraphs (WG), which describe how a system of application pipelines should execute across a set of domain specific accelerators in a specific sequence in order to process a given input workload dataset. A workgraph may be data that captures a list of graph node elements that may represent domain operations and a description of the related resources. Workgraphs, which may be known as workgraph algorithms, capture the user’s algorithmic intent and overall structure, without burdening the developer to know too much about the specific hardware it will run on. Workgraphs may therefore provide a description of the tasks that are to be performed as part of an application pipeline and may include one or more passes or descriptions of other node elements that may form the application processing pipeline.

[0008] A pass may be used to represent the input and output resource descriptions along with a set of domain operation descriptions and associated dataset. Some examples of passes may be render pass, machine learning pass, transfer pass. A further example is shown in Figure 1, where segmentation and detection are two separate passes running on the NPU device in parallel. Each pass may have associated with it one resource A, B or C and each resource may have an associated resource lifetime for that pass. Furthermore, a workgraph with multiple passes may be executed on a single domain specific accelerator device, for example a GPU is typically capable of executing both graphics rendering and compute domain operations so a workgraph containing only these two types of passes can be fully executed on a GPU device. Figure 1 illustrates an example of a workgraph for augmented reality filtering wherein a camera captured image is input to an ISP capture process 101. After the ISP capture 101 a series of passes are to be performed by a NPU in order to perform segmentation 102 and detection 103. The workgraph of Figure 1 also includes a post processing pass 104 performed by a GPU in consideration of the passes performed by the NPU. A scan-out pass 105 is also included in the example workgraph to provide a screen output. A pass dependency (PD) specifies the ordering of a set of passes based on input and output dependencies connection to other passes with a workgraph (WG) which controls the scheduling and sequencing of these pass execution. A pass dependency list (PDL) captures the entire set of pass dependencies for a given workgraph. In other words, the pass dependency list may provide the actual values of the passes that are dependent on each other and the event information that may be used to trigger the processing of a pass. The pass dependency list can be explicitly specified by the user or can be inferred by the runtime during pass dependency analysis by comparing each pass dataflow requirement. A pass dependency lists are typically expressed as an acyclic graph where nodes represent passes and edges represents pass dependencies.

[0009] A pass dependency list not only captures the dependencies of passes used for dataset transformations, it can also capture other dependencies when executing a workgraph such as resource provisioning and data movement kind of pass dependencies to minimize access latency, both of these are required to make efficient use of available system resource and delivery high performance heterogeneous task execution.

[0010] On a mobile System on Chip platform with a host processor, a set of Domain Specific Accelerators (DSAs) devices, memory and an interrupt controller, any workgraph lowered down for execution must be managed by a Task Scheduler (TS) running on the host. The main role of the task scheduler is to perform Heterogenous Task Scheduling (HTS), which sequences workgraph pass executions in the order described by the pass dependency list (PDL) schedule and effects the dispatch of corresponding domain operations to the domain specific accelerator devices when its pass dependencies have been met, an example of such a workflow is illustrated in Figure 2. The task scheduler also dispatches Command Buffers (CB) to domain specific accelerators for execution, these command buffers contain the set of pass domain operations that have been assigned during workgraph lowering by the driver to the same domain specific accelerator device on the system on chip platform. In Figure 2 it can be shown that the workgraph can be transmitted between host devices and therefore the original host device does not need to be consulted by each domain specific accelerator device linked to subsequent host devices. As shown in Figure 2, the domain specific accelerators are only required to return to the previous host device to obtain the workgraph information via a completion interrupt request (IRQ) and pass dispatch commands, it is not necessary to return to the source host.

[0011] Despite the advantages that may be achieved by the above approaches, there are a number of problems with these approaches. Current approaches to heterogenous task execution do not separate workgraph descriptions from its underlying execution schedule, these two definitions are almost always interleaved and tightly coupled together such that experimenting with changing the execution schedule requires rewriting large portions of the workgraph with every schedule change permutation. Furthermore, by coupling workgraph descriptions with its execution schedule limits its performance scope and utility as users are ill equipped when exploring its schedule space in order to find the most optimal schedule configuration that overlaps computation and communication for a given system on chip platform without changing the workgraph definition. The solution space for implementing heterogenous task scheduling is also severely limited in environments that use a coupled workgraph definition as it is no longer possible to orchestrate heterogenous task execution by using a central hardware which consumes dependencies to sequence execution schedules.

[0012] Performing heterogeneous task execution on the host processor requires constantly circling back to the host after any pass execution is completed by any domain specific accelerator device, this is required in order for the task scheduler to be able to calculate and dispatch the next set of dependent pass executions unblocked by the just completed pass to their assigned domain specific accelerator devices. This was briefly described above in relation to Figure 2.

[0013] These repeated relaying of signals between domain specific accelerator devices and the host causes delays between successive pass dispatch in the workgraph, which leads to the domain specific accelerator devices idling during this roundtrip period and not performing useful work. Furthermore, in such an example scenario the host unable to idle during workgraph execution leading to increased power consumption.

[0014] By using a domain specific accelerator adapter (DSAA) mechanism, the task scheduler now expresses a workgraph as a set of command buffers plus an event list (EL) so that the input dispatch event for a pass assigned to a domain specific accelerator can be consumed directly by its domain specific accelerator adapter to dispatch said pass execution if and only if the Output Completion Event (OCE) for its pass dependency was also generated by the same domain specific accelerator adapter.

[0015] This approach finds utility in workgraphs where passes that are back-to-back are assigned to same domain specific accelerator so they can be directly executed by the domain specific accelerator adapter without any host participation especially when a single domain specific accelerator device (i.e. GPU) is being used to execute multiple domain operations types (Render, Compute, machine learning, etc.) on a system on chip platform as shown in Figure 3. Figure 3 illustrates workgraph execution using a domain specific accelerator adapter such that the domain specific accelerator does not need to return to the original host as in Figure 2.

[0016] Whilst this approach improves upon the host centric heterogenous task scheduling workflow, its utility is a function of a workgraph topology and degenerates to a host centric heterogenous task scheduling workflow if there exist no back-to-back passes in workgraphs.

[0017] As mobile system on chip platforms original equipment manufacturers (OEMs) source domain specific accelerator devices typically from multiple intellectual property (IP) vendors, it is very common for these devices to use different command synchronization methods such as hardware-only, software-only with / without operating system support or mixed-approach mechanisms. These synchronization primitives are further abstracted (i.e. Semaphore, Pipeline Barrier, Interrupts, Fences) to the user application leading to very inefficient memory space representation which impacts task scheduler performance when building and scheduling command buffers.

[0018] Furthermore, workgraph management by host increases during heterogenous task scheduling because the task scheduler uses a mixed-bag of synchronization primitives, which requires different handling requirements on the host, for composing its pass dependency list. As each domain specific accelerator must be synchronized using its own vendor specific and / or domain specific mechanisms, this leads to highly inefficient scheduling workflows by the task scheduler causing high latency forward progress execution of the workgraph.

[0019] Some of these problems are identified in the art, for example Vulkan API uses the concept of subpass dependencies to order pass executions, it also provides general queueing primitives (i.e. semaphores & fences) to synchronize command buffer (CB) submissions. Pipeline barrier primitives are used to order sets of domain operations whilst event primitives are used to effect fine grained dependency and synchronization control inside command buffer. Vulkan supports workgraphs in a limited sense in the Dispatch Dependency extension which is not heterogenous and targets only compute, it provides primitives for specifying how to lunch different interdependent compute kernels.

[0020] A further example is the use of OpenCL API which provides a "wait list" abstraction that enables developers to express complex workflows, and optimize ordering of command execution, it is a critical mechanism for managing command dependencies and a primitive for structuring efficient parallel computing applications in OpenCL. Heterogenous workgraphs can be manually composed by the user using wait list primitives such that they may then be applied to specific hardware devices through compiling them into command buffers specific to the hardware selected. Such an approach may include some key aspects such as, dependency management: Wait list allows users to specify dependencies between different commands, which can include tasks like kernel executions, data transfers, or memory operations. By creating a wait list, users define the order in which different commands execute based on their dependencies. Synchronization: As multiple commands may run concurrently on different devices (e.g., CPUs and GPUs), wait list ensures that specific commands do not begin execution until all specified prerequisite commands have completed, this helps prevent data races and ensures data is available when needed by dependent commands. Optimizing Execution: Wait list are used to optimize the execution of commands by scheduling data transfers in parallel with kernel executions which helps to minimize device idle time and improving overall performance. Resource Management: Wait lists are used ensures that resources (e.g., buffers, images, or devices) are not accessed or modified by subsequent commands until they are ready, avoiding contention and potential conflicts. Asynchronous Workflows: Wait lists provides a structured way to coordinate the asynchronous execution of commands, making it easier to express intricate workflows that execute in very controlled sequence.

[0021] In an alternate known approach OneAPI provides a C++ language (Single-source Heterogeneous Programming) for programming multi-processors and accelerators such as GPUs and Field Programmable Gate Arrays. Threading Building Blocks Library (oneTBB) breaks computations into parallel tasks and Collective Communications Library (oneCCL) provides communication primitives for scaling computations across multiple devices with support for both scale-up for platforms with multiple OneAPI devices and scale-out for high performance computing clusters with multiple compute notes. The Deep Neural Network Library (oneDNN) provides primitives for compiling and executing deep learning computation graphs.

[0022] A further approach uses a Single-source Heterogeneous Programming for C++Task Graph. The Single-source Heterogeneous Programming for C++ API uses the concept of task graphs as a high level and portable abstraction for managing and scheduling heterogenous task execution. Task graphs can leverage multiple processing devices such as CPUs, GPUs, and Field Programmable Gate Arrays while managing dependencies and optimizing performance. The key aspects of this approach are as follows:

[0023] 1. Parallel Task Management: Task graphs allow developers to express their algorithms as a set of parallel tasks which can represent different parts of a computation and be executed concurrently on available devices.

[0024] 2. Heterogeneous Execution: Task graphs enables the utilization of diverse hardware resources, including CPUs and GPUs by providing a way to describe how tasks can be distributed and executed across these devices.

[0025] 3. Dependency Management: Task graphs allow developers to specify dependencies between tasks so that tasks may only execute after its dependent tasks have completed.

[0026] 4. Dynamic Scheduling: Task graphs can dynamically schedule tasks based on runtime conditions or available resources, this can improve load balancing and optimize execution for various hardware configurations.

[0027] 5. Low-Level Abstraction: Task graphs abstracts the underlying hardware-specific APIs (like CUD A or OpenCL) and provides a higher-level interface for task scheduling, making it easier for developers to write portable and efficient code.

[0028] 6. Portability: Task graphs, combined with Single-source Heterogeneous Programming’s portability features, allows applications to run on different platforms without extensive code modifications using a single codebase.

[0029] 7. Performance Optimization: By expressing computations as a task graph, developers optimize task executions to make the best use of available hardware, potentially leading to improved application performance. An additional approach utilises Heterogenous Interface for Portability / Radeon Open Compute Platform streams. HIP (Heterogenous Interface for Portability) API on ROCm (Radeon Open Compute Platform) provide Stream primitives which can be used to specify how a task’s computation and communication operations are dispatched and concurrently scheduled on single / multi-GPU systems. Stream primitives are used to launch concurrent kernels, and kernels in the same stream context have sequential dependency.

[0030] Furthermore, streams are the basic primitive for local platform multi-GPU dispatch support, Message Passing Interface (MPI) API is used for distributed multi-GPU dispatch. Local GP-GPU communication patterns are provided using first-party Heterogenous Interface for Portability / Radeon Open Compute Platform libraries.

[0031] Heterogenous task graph computations and communications can be manually composed and scheduled for execution by using any combination of the above set of Heterogenous Interface for Portability / Radeon Open Compute Platform primitives, this approach provides the benefit of explicit control. Explicit control is achieved by providing several low-level primitives for CPU-GPU compute dispatch on both local and remote platforms with GPU-GPU communications, a great deal of control can be exercised by the developer when composing higher level abstractions for delivering robust, flexibility and performance control.

[0032] Finally, one further approach is the use of a DirectX12 API which uses a workgraph abstraction to provide a system for GPU- only autonomy for general compute workload dispatch directly from the GPU device in order to unlock latent GPU capabilities. The D3D12 workgraph is a graph of nodes where shader code at each node can request invocations of other nodes, without waiting for them to launch.

[0033] While the above approaches provide certain benefits, there are a number of disadvantages to using each of these approaches. For example, Vulkan Dependency abstractions suffer from portability challenges as its implementations varies across different platforms so may not provide optimal performance across all implementations. OpenCL wait lists provides low-level control over tasks and dependencies, developers seeking a more abstract way of expressing parallelism must use alternative programming models, this limit wait lists functional scope. Wait lists adds complexity to codebase making it prone to coding errors when coding many interdependent tasks. OneAPI Single-source Heterogeneous Programming for C++’s task graph mechanism introduce performance overhead due to using a host centric heterogenous task scheduling model, causing devices to excessively idle during heterogenous task execution which reduces overall performance and throughput.

[0034] Heterogenous Interface for Portability / Radeon Open Compute Platform streams have limited concurrency because it imposes a certain level of sequential execution, this limitation can hinder the full exploitation of parallelism on devices especially when composing heterogenous task dependencies. DirectX12 Work Graphs abstractions may not provide enough control over the hardware for very performance critical applications as performance / behaviour can differ between hardware platforms leading to inconsistencies.

[0035] SUMMARY

[0036] According to a first aspect of this disclosure there is provided a for processing an application pipeline, the device comprising one or more processor configured to: receive a workgraph comprising one or more pass to be performed for processing the application pipeline, each of the one or more pass comprising one or more task formed of one or more command for processing the application pipeline; generate a pass dependency schema representing an order in which the one or more pass in the workgraph is to be processed when the application pipeline is executed; wherein the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema comprises event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed. This allows the processing of the application pipeline to be easily adapted due to the decoupling of the workgraph with the pass dependency schema, and further allows the processing of the commands that make up the tasks of the workgraph to be controlled based on the event information.

[0037] The device as described above, wherein the pass dependency schema is encoded in a 32-bit or 64-bit hardware descriptor. This allows the pass dependency schema to be easily implemented on hardware.

[0038] The device as described above, wherein the pass dependency schema further comprises pass information representing the cyclic properties of the one or more pass. This allows for identical chains of task executions implementing workload processing strategies to be iterated.

[0039] The device as described above, wherein the one or more processor is further configured to execute the workgraph on a domain specific accelerator. This allows the workgraph to be processed after having been adapted to be specific to the appropriate domain specific accelerator.

[0040] The device as described above, wherein the one or more processor is further configured to: segment the workgraph into one or more slices, and execute each of the one or more slices of the segmented workgraph. This allows for the workgraph to be segmented and provided in parts to different domain specific accelerators in order to increase the efficiency of processing.

[0041] The device as described above, wherein the one or more processor is further configured to: generate a pass dependency list representing the dependencies of the one or more pass in the workgraph; provide the pass dependency list and / or the pass dependency schema to a hardware device suitable for execution. This allows the priority of the passes comprised within the workgraph to be described for execution.

[0042] The device as described above, wherein one or more processors are configured to embed the workgraph in one or more command buffer associated with the hardware device. This provides the advantage that the workgraph need not be obtained from the source host for every pass of the workgraph.

[0043] The device as described above, wherein the one or more processor is further configured to: modify the workgraph to comprise standard domain operations configured to adapt the workgraph to be executable on non-host domain specific accelerators, and / or modify the pass dependency list, independently of the workgraph, to represent a further order in which the one or more pass in the application pipeline is to be executed. This provides the advantage that the user can adapt the pass dependency list included associated with the pass dependency schema while not adapting the workgraph in order to change the priority of commands executed by the device.

[0044] According to a further aspect of this disclosure there is provided a method for processing an application pipeline, the method comprising: receiving a workgraph comprising one or more pass to be performed for processing the application pipeline, each of the one or more pass comprising one or more task formed of one or more command for processing the application pipeline; generating a pass dependency schema representing an order in which the one or more pass in the workgraph is to be processed when the application pipeline is executed; wherein the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema comprises event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed. This allows the processing of the application pipeline to be easily adapted due to the decoupling of the workgraph with the pass dependency schema, and further allows the processing of the commands that make up the tasks of the workgraph to be controlled based on the event information. The method as described above, wherein the pass dependency schema is encoded in a 32-bit or 64-bit hardware descriptor. This allows the pass dependency schema to be easily implemented on hardware.

[0045] The method as described above, wherein the pass dependency schema further comprises pass information representing the cyclic properties of the one or more pass. This allows for identical chains of task executions implementing workload processing strategies to be iterated.

[0046] The method as described above, wherein the method further comprises executing the workgraph on a domain specific accelerator. This allows the workgraph to be processed after having been adapted to be specific to the appropriate domain specific accelerator.

[0047] The method as described above, wherein the method further comprises: segmenting the workgraph into one or more slices, and executing each of the one or more slices of the segmented workgraph. This allows for the workgraph to be segmented and provided in parts to different domain specific accelerators in order to increase the efficiency of processing.

[0048] The method as described above, wherein the method further comprises: generating a pass dependency list representing the dependencies of the one or more pass in the workgraph; providing the pass dependency list to a hardware device suitable for execution. This allows the priority of the passes comprised within the workgraph to be described for execution.

[0049] The method as described above, wherein the method further comprises: embedding the pass dependency list and / or workgraph in one or more command buffer associated with the hardware device. This provides the advantage that the workgraph need not be obtained from the source host for every pass of the workgraph.

[0050] BRIEF DESCRIPTION OF THE FIGURES

[0051] The embodiments of the present disclosure will now be described by way of example with reference to the accompanying drawings. In the drawings:

[0052] Figure 1 illustrates an example workgraph for augmented reality filtering;

[0053] Figure 2 illustrates an example of workgraph scheduling performed by the task scheduler;

[0054] Figure 3 illustrates an example of workgraph execution using a domain specific accelerator adapter;

[0055] Figure 4 illustrates an example of 2 to 1 pass dependency in a workgraph;

[0056] Figure 5 illustrates an example of the pass dependency in a 32-bit hardware descriptor;

[0057] Figure 6 illustrates an example of a hardware task scheduler compatible system on chip platform;

[0058] Figure 7 illustrates an example of a hardware task scheduler based heterogeneous task scheduling;

[0059] Figure 8 illustrates an example of an image filtering workflow task dependency graph;

[0060] Figure 9 illustrates an example of an image filtering task dependency graph; and

[0061] Figure 10 illustrates an example of an expanded pass dependency 64-bit hardware descriptor.

[0062] DETAILED DESCRIPTION OF THE INVENTION

[0063] This disclosure provides methods and techniques to overcome the above problems discussed above. In particular, this disclosure provides a workgraph that is separate and decoupled from a pass dependency schema. In addition, the present disclosure provides a unified schema representation (pass dependency schema) that may be implemented by any domain specific accelerator. Generally, workgraphs may capture the user’s algorithmic intent and overall structure, without burdening the developer to know too much about the specific hardware it will run on. This asynchronous nature can maximize the freedom for the system to decide how best to execute the work and manage delegated resources during the workflow. Workgraphs may therefore provide a description of the tasks that are to be performed as part of an application pipeline and may include one or more passes or descriptions of other node elements that may form the application processing pipeline. Such a workgraph may be efficient to load and process at runtime.

[0064] This disclosure introduces a new efficient pass dependency schema representation (pass dependency schema) and memory encoding which decouples a workgraph description from its underlying heterogenous task execution schedule and can be consumed efficiently by all system on chip platform devices. The pass dependency schema representation captures three categories of information for both sides of a pass dependency as summarized in Table 1. Table 1 demonstrates that the pass dependency schema includes unique identifying information, event information and information regarding the number of executions.

[0065] In particular, this disclosure provides a device for processing an application pipeline. The device of this disclosure comprises one or more processor that is configured to: receive a workgraph representing one or more pass to be performed in processing an application pipeline. Each of the one or more pass may comprise one or more task formed of one or more command of the application processing pipeline. The workgraph is therefore a representation of the passes that will be performed as part of executing that workgraph to perform one or more tasks. The tasks would be executed as commands stored in a command buffer. The device and the method of the present disclosure is further configured to generate a pass dependency schema representing order in which the one or more pass in the application pipeline is to be executed. In other words, the pass dependency schema may be a template structure for the passes and may be generated using hardware or software to infer the structure of the passes that will be needed to execute the workgraph from analysis of the workgraph itself. As such the pass dependency schema provides a template structure inferred from functions that will be performed by and are described in the workgraph. In generating the pass dependency schema, a template for the pass dependencies is produced, as discussed based on the passes that will be needed in the workgraph, but in addition in some cases, in dependence on an input from a user regarding which passes may be needed to action the workgraph. The pass dependency schema may be inferred from the workgraph based on the ID of the passes comprised within the workgraph, for example if the workgraph describes that a segmentation and a detection functions have a specific IDs and provide the input for post processing provided by a GPU then the pass ID of the post processing performed by the GPU will depend on the resource ID of the segmentation and detection steps and thus the pass dependency schema will provide as part of its structure the ability to perform all three functions. A user may further configure the schema to modify or ignore parts of the inferred pass requirements from the workgraph and may override an inferred relationship between passes in the workgraph by providing a relationship specified by the user. Furthermore, the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema may also comprise event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed. Since the workgraph is separate and decoupled from the pass dependency schema it is possible for the pass dependency list to be adapted without changing the workgraph. In other words, the one or more processor of the device of this disclosure may be configured to: modify the workgraph, independently of the pass dependency schema, to comprise standard domain operations configured to adapt the workgraph to be executable on non-host domain specific accelerators, and / or modify the pass dependency schema list, independently of the workgraph, to represent a further order in which the one or more pass in the application pipeline is to be executed.

[0066] Furthermore, by including the event information in the pass dependency schema it is possible to provide a trigger for a subsequent pass to be executed based on the indication that previous passes have been executed. Using this pass dependency schema, it is possible to express various pass dependency rules using event specifications on both source and destination side of a dependency. Furthermore, any pass dependency complexity canbe expressed using combinations of rules described below. 1. One to One Rule: pass dependency specifies that a single dependent pass execute on target domain specific accelerator when a single independent pass execution has completed on a source domain specific accelerator.

[0067] 2. One to Many Rule: pass dependency specifies that multiple dependent passes execute on multiple target domain specific accelerator when a single independent pass execution has completed on a source domain specific accelerator.

[0068] 3. Many to One Rule: pass dependency specifies that a single dependent pass execute on target domain specific accelerator when multiple independent passes have completed execution on multiple source domain specific accelerator.

[0069] 4. Many to Many Rule: pass dependencies specify that multiple dependent passes execute on multiple target domain specific accelerator when multiple independent passes have completed execution on multiple source domain specific accelerator.

[0070] For efficient pass dependency list generation, source domain specific accelerator identification may be omitted from pass dependency specification, this requires that all source events be unique when expressing a pass dependency. The pass dependency list of this disclosure may be generated based on analysis of the workgraph of this disclosure, for example by software that analyses the functions that will be performed by the workgraph and generates the specific passes and resources for each pass that will be needed to action the workgraph. The pass dependency list is therefore generated by a program that links the passes in the workgraph based on their relationships to one another. As shown in Figure 4 there is an example of the 2 to 1 pass dependency and the events associated with each pass. For example. Figure 4 includes segmentation 402 and detection 403 passes each having an event el or e2 associated with them respectively. The results of these two passes are used as the passes that a subsequent pass is dependent on, in this case, the subsequent pass is a post processing pass 404. As can be seen in Figure 4, the event information e 1 and event information e2 are combined to form event information e3, which dictates that the post processing pass 404 can commence as the information provides an indication that the segmentation and detection passes have been completed. This allows the event information contained within the pass dependency schema to act as a marker for instructing the device to perform the next processing task as shown in Figure 4. Using the pass dependency schema of this disclosure, the pass dependency execution schedule in Figure 4 is expressed using the following notation:

[0071] (NPU(segment)[el], NPU(detection)[e2]) => ([e3]GPU(post-proc))

[0072] This specifies that the GPU post processing pass should only commence execution if and only if both the NPU segmentation and detection passes have completed execution.

[0073] In addition to the above, the present disclosure also provides efficient memory encoding. The pass dependency schema decoupled from the workgraph may be captured in memory by the task scheduler using an efficient encoding format that optimizes all heterogenous task execution workflows from host to domain specific accelerators. The pass dependency list memory format may be in a hardware friendly representation, in the form of a 32-bit hardware descriptor, as shown in Figure

[0074] 5. This fixed-format pass dependency descriptor encoding is lowered down efficiently for consumption by all system on chip platform devices. In other words, the pass dependency schema may be encoded in a 32-bit hardware descriptor, alternatively, the pass dependency schema may be encoded in a 64-bit hardware descriptor. The hardware descriptor format should be understood to be the programmable data structure of the schema, for example the pass dependency schema will describe the structure of the pass dependencies used as part of the workgraph and this description may be provided to the hardware that is to execute the workgraph algorithm in a hardware descriptor format readable by the hardware, which may be in a 32 or 64-bit format. A pass dependency can be expressed using one or many such descriptors as shown as an example in Figure 5, depending on its complexity, so by extension a pass dependency list may map into a description list (DL). A single descriptor specifies how a set of input (source) events emitted by a set of domain specific accelerator devices map to a set of output (destination) events consumed by a set of domain specific accelerators which sequences the next set of passes for the workgraph forward progress execution, e.g., the processing of the commands as dictated by the workgraph. The information space of a pass dependency may be efficiently encoded into this 32-bit hardware descriptor by capturing the pass dependency schema fields. The fields of the pass dependency schema may be thought of as one or more of the following, command code that encodes the type of pass dependency descriptor, events / event information that encodes the domain specific accelerator source and domain specific accelerator destination events that the pass dependency, identifications that encodes a unique identification assigned to domain specific accelerator on the system on chip platform, and topology information that encodes optional information about a pass dependency schedule (schema) such as embedded cyclic paths. In other words, the topology may comprise pass information representing the cyclic properties of the one or more pass. These cyclic properties may relate to the iterative processing of the workgraph algorithm in slices, for example the slicing of the workgraph due to available resources may lead to one or more of the pass of the workgraph being iteratively processed. An example of this is that two passes of a sliced workgraph may be alternately processed in a cyclic manner for a set number of iterations. The number of iterations as well as information describing which passes will be cyclically iterated may be included in the cyclic properties. This memory encoding facilitates the efficient generation of the pass dependency list by the task scheduler because it can stream it to memory during workgraph generation by making use of faster cache aligned memory transactions.

[0075] The present disclosure also provides a dynamic workgraph schedule as part of the method and device of this disclosure.

[0076] The decoupled pass dependency list of this disclosure also provides a means for expressing complex workload processing patterns and facilitates the use of different schedule for executing a workgraph with the selection of the most performant schedule deferred to runtime after system on chip platform discovery of available resources and limits has been performed at start-up. One such schedule for executing workgraph is the use of workload multi-slicing. Workload slicing prior to execution may be performed by multi-slicing wherein a spatial and temporal technique is used in which a workgraph input workload dataset is processed in smaller units called slices such that multiple invocations of the workgraph are used to complete the required processing on the entire input data in order to improve domain specific accelerator memory hierarchy performance. In other words, the one or more processors of the device of this disclosure may be further configured to segment the workgraph into one or more slices, and execute each of the one or more slices of the segmented workgraph. Dynamic schedules may provide sufficiently low-level controls for describing runtime attributes of workgraph scheduling and execution such as resource management and workflow patterns such as cyclic sub workgraphs. Segmentation of workgraph to one or slice can, in some cases, be done without changing the workgraph itself but just changing the pass dependency list with new slice dependencies added to the original pass dependency list while creating data slices. This is only possible because of decoupling of the pass dependency schema and the workgraph. In other cases, the pass dependencies may be fixed and the workgraph may be segmented based on the available resources of the target (endpoint) hardware device e.g., the CPU. In such a case, the pass dependencies will remain fixed while the workgraph itself will be segmented such that reduced amounts of data can be processed at one time and in a more time efficient manner.

[0077] The decoupled pass dependency list descriptors (schema) of this disclosure now provide a method for offloading heterogenous task execution activities away from a host to another system on chip (system on chip) platform device. Such a device that achieves this may be thought of as a Hardware Task Scheduler (HWTS). By parsing and decoding pass dependency list descriptors, the hardware task scheduler has all the information needed to orchestrate a workgraph execution by interoperating with domain specific accelerator devices using events. In other words, the domain specific accelerator does not need to return to the host for each pass of the workgraph when executing. This can reduce the delay in signal transmission. In order for the task scheduler to leverage hardware task scheduler functionality, for each workgraph, the present device and method of this disclosure may generate a pass dependency list and lowers (transmits / passes on) it to the hardware task scheduler for consumption (processing) and embeds event lists (event information) into command buffers assigned to each domain specific accelerator, this event information may comprise the set of input dispatch event information and output completion event information for passes assigned to each domain specific accelerator for its domain specific accelerator adapter to use to interoperate with the hardware task scheduler.

[0078] This approach removes the need to constantly circle back to the host during heterogeneous task scheduling is deferring part of the scheduling and dispatch responsibility to the domain specific accelerator devices. This may be realized using on-board domain specific accelerator solutions like a firmware code or hardware co-processor located in the command frontend subsystem of the domain specific accelerator devices, this may be thought of as a domain specific accelerator Adapter (domain specific accelerator adapter), the task scheduler attaches dependency specification event information into command buffers for processing by domain specific accelerator adapters. Pass dependency list may be provided alongside command buffers.

[0079] Such event configuration may comprise two items of information: (1) Input Dispatch Event (IDE): which specifies that a pass dependency has been met causing domain specific accelerator adapter to dispatch said pass domain operations for execution and (2) Output Completion Event (OCE): which signals that a pass execution on a domain specific accelerator device has completed causing task scheduler to schedule next sets of dependent passes based on the workgraph schedule.

[0080] A hardware task scheduler compatible system on chip platform may include a common event bus fabric linking the host to hardware task scheduler and the hardware task scheduler to all domain specific accelerator devices on the system on chip platform as illustrated in Figure 6. Figure 6 illustrates an example of a hardware task scheduler compatible platform that includes a CPU, GPU, and NPU comprising L2 memory and in communication either directly or indirectly with a host bus (axi). There may also be provided an event bus that communicates with the domain specific accelerator devices and the hardware task scheduler in a two-way manner. Figure 6 also demonstrates that the platform may include a system cache, main memory and I / O devices. Such a platform as shown in Figure 6 may be wholly or at least partially implemented on the device of this disclosure and may be configured to implement the method of this disclosure.

[0081] Figure 7 shows an example of a heterogeneous task scheduling workflow that may be implemented by the platform of Figure 6 and / or the device of the present disclosure. The heterogenous task scheduling workflow may be hardware task scheduler based, which allows the task scheduler overhead to be completely reduced to 0% during workgraph execution which is now handled completely by the hardware task scheduler as the entire heterogenous task scheduling workflow may bypass the host as shown in Figure 7 and summarized in Table 2. Figure 7 demonstrates an example of the workflow (workgraph actions / passes) that may be performed on the platform. For example, Figure 7 demonstrates the workflow steps that are shown in Table 2.

[0082] The efficient schema encoding and memory representation of a pass dependency list descriptor allows for initial pass execution on a domain specific accelerator device to gather relevant system on chip platform information needed to compose and stream out its own autonomous pass dependency list to the hardware task scheduler to initiate device driven workflow independently of the host. This mechanism provides full system autonomy for domain specific accelerator devices to bypass the host to effect background computations and memory operations independently of host supervisor.

[0083] On such example application of an application of the method and device of this disclosure is a common mobile phone application scenario that may perform image filtering, the task graph in Figure 8 illustrates a simplistic heterogenous workflow that is used for filtering an image for display on system on chip platforms. The pass dependency list is used to encode a workgraph which captures all the domain specific accelerator domain operations (rendering 801, conversion 802, filtering 803 and scan out 804), communications and resource management requirements needed for executing the image filtering workflow from a source device (GPU) to a sink device (DPU) when lowered down to the hardware task scheduler for execution. The image filtering workgraph may be configured by the task scheduler for optimum performance by leveraging available system on chip platform System Cache (Fast Memory) resources and workload multi-slicing strategy to accelerate task computations and communications. This is possible as the pass dependency list descriptors allows for the description of complex graph topologies with support for embedding a single or multiple cyclic sub-workgraph to iterate identical chains of task executions to implement workload processing strategies such as multi-slicing as shown in the example of Figure 9.

[0084] Figure 9 is an example of a task dependency graph for image filtering, wherein it can be seen that the workgraph (algorithm) is segmented into one or more, in the case of Figure 9, sliced components. Each of these slices may be performed by a separate domain specific accelerator in the graph for example, the CPU, GPU or NPU. Furthermore, the use of a hardware task scheduler as indicated in Figure 9 allows for event information to be used to inform a multi-slicing subgraph utilising the fast memory to reserve the allocation and deallocation of resources.

[0085] As an example of the system architecture that could be used to execute the optimized image filtering workgraph execution schedule the following components may be used for performing the activities. A host processor that may execute task scheduler which converts a logical workgraph description and pass dependency list into an efficient execution schedule by leveraging system on chip platform heterogenous computing capabilities. Such a host processor may consume two items of information that may comprise domain specific accelerator command buffers comprised of event sequenced task domain operations that are queued up for execution for each domain specific accelerator hardware scheduler pass dependency list that may comprise the list of pass dependencies that captures the workgraph execution schedule dependency information. Furthermore, the host also functions as a domain specific accelerator and consumes transferred domain operation command buffers, wherein the domain specific accelerator executes transfer domain operations. The CPU workload may be configured by the task scheduler in multi-slicing mode from RGB framebuffer into fast memory.

[0086] The device may also include a hardware task scheduler device that may process pass dependency list by loading its descriptors, it builds an internal representation of the workgraph pass dependencies and performs heterogenous task scheduling by sequencing through dependent passes in the workgraph based on events sent / received to / from domain specific accelerator adapters. A domain specific accelerator adapter may also be present as a functional adaptation to domain specific accelerator devices making it interoperate with the system on chip platform hardware task scheduler device by reacting to event list (EL) configurations (i.e. IDE and OCE) attached to domain operations embedded by the task scheduler in submitted command buffers. A GPU device may also be present, which processes event list-sequenced render domain operations command buffers, executes render domain operations when triggered by hardware task scheduler and notifies it when render domain operations have completed. The GPU output is configured by the task scheduler in multi-slicing mode into the RGB framebuffer.

[0087] The device of this disclosure may also in some cases include an NPU device, which may process event list-sequenced machine learning domain operations command buffers, executes machine learning domain operations when triggered by hardware task scheduler and notifies the hardware task scheduler when machine learning domain operations have completed. The NPU workload may be configured by the task scheduler in multi-slicing mode from Fast memory into display framebuffer. A DPU device may also be provided that may process event list-sequenced presentation domain operations command buffers, executes presentation domain operations when triggered by the hardware task scheduler and notifies the hardware task scheduler when presentation domain operations have completed. Finally, a system cache may be provided and may be thought of as a high performance last-level system on chip memory cache configurable both in cache and / or buffer mode and used as a staging area for intermediate workgraph workload dataset. Furthermore, the example workflow in Figure 9 describes that the task scheduler initiates workgraph execution by submitting the pass dependency list descriptors to hardware task scheduler and sending an event (as event information el) to initiate the image filtering heterogeneous task execution workflow. On receiving el, the hardware task scheduler schedules e2 to the GPU and e3 to the CPU which causes both domain operations to commence forward progress until their completion notification events e4 and e5 are sent to the hardware task scheduler respectively - this completes the first phase of heterogenous task execution of the workgraph. In other words, the event information represents whether each of the one or more pass or domain operations has been executed, e.g., by the CPU e5 and the GPU e4 in the case of Figure 9. The workgraph may be executed on a domain specific accelerator. Furthermore, the workgraph may be embedded in one or more command buffer associated with the hardware device (domain specific accelerator) as shown in Figure 9 by the arrows from the sliced segments of the workgraph to the CPU and GPU such that the CPU and GPU may execute the passes. The device of the present disclosure may be configured to generate a pass dependency list representing the dependencies of the one or more pass in the workgraph and provide the pass dependency list and / or the pass dependency schema to a hardware device (domain specific accelerator - CPU / GPU / NPU) for execution. The hardware device may be a dedicated hardware device dedicated to the function performed by the workgraph or specified by the user to execute the workgraph. The commands may then be executed based on these generated pass dependencies, which may include event information.

[0088] The second phase of heterogenous task execution of the workgraph may use workload multi-slicing to progressively composite the final Display framebuffer using a cyclic sub-graph that repeats from the CPU to the NPU. Each sub-graph instance is scheduled by e6 and works on a subset of the RGB framebuffer. The task scheduler configures multiple emits (repeat count) of e6 based on the workload multi-slicing strategy in use which is dependent on workload size and fast memory size. The last phase of heterogenous task execution of the workgraph occurs when the hardware task scheduler schedules e9 and e 10 events respectively which causes the DPU to scan out the full display framebuffer to the screen.

[0089] The mapping of pass dependency list information space items into a descriptor representation may allocate different bit sizes to encode the various workgraph pass dependency information. The hardware descriptor may be expanded to 64-bits or higher to accommodate such information as illustrated in Figure 10. Figure 10 provides an example of the bit allocation from the hardware descriptor. Processing of pass dependency list using the hardware task scheduler may be done in firmware executing on system on chip platform microcontroller or fixed-function co-processor. Furthermore, the hardware task scheduler may function as domain specific accelerator performing domain operations such as but not limited to direct memory access (DMA) operations, command buffer processing and patching. Event processing augmentation using the domain specific accelerator adapter may be implemented using fixed-function logic or as a firmware executing on onboard microcontroller or co-processor on the domain specific accelerator. Furthermore, the event list (event information) may be encoded directly in the domain specific accelerator queues by the task scheduler or separately in a side channel buffer accessible to domain specific accelerator adapter.

[0090] An event bus connecting the hardware task scheduler and domain specific accelerator adapter may be emulated using existing peripheral / memory bus fabric or may be novel fabric. The event bus protocol may be implemented as a protocol on top of existing peripheral bus protocols or may be separate and novel from the existing peripheral bus protocols. The event bus protocol may implement message delivery guarantees to mitigate against event loss, this may be realized using methods such as back pressure controls or message retries based on handshaking acknowledgements.

[0091] Furthermore, the protocol may implement quality of service (QoS) guarantees thus ensuring low latency delivery of events on both sides of the event bus or specifically between high bandwidth domain specific accelerator. Priority controls may be supported by the protocol thereby mitigating against priority inversions when multiple workgraph are making forward progress concurrently on any number of domain specific accelerator, this may be realized on either end of the event bus. A workgraph (e.g., workgraph) that may be received or generated by the device and method of this disclosure may be expanded to include host processors as domain specific accelerators, this allows generic domain operations to be inserted into the workgraph. Generic domain operations may include but not limited to control, compute, transfers, render or machine learning operations that may be used to augment the functionality and / or performance of non-host domain specific accelerators on the system on chip platform. The task scheduler that may be implemented by one or more processors of the device of this disclosure may leverage such host domain operations to perform several workgraph optimizations. One such optimisation is workgraph fusion, wherein multiple dependent workgraphs with intermediate CPU pre / post processing can be merged into a single workgraph. Doing so opens up opportunities for resource pipelining and deferred CPU task execution which improves runtime performance. Another optimisation that the device and method of this disclosure is the use of asynchronous workgraph operations in which fused workgraph execution can be optimized by scheduling some CPU pass tasks for execution just-in- time (i.e. as late as possible). These CPU tasks will execute asynchronously with domain specific accelerator tasks and are used to reduce idle / wait time of subsequent downstream domain specific accelerator pass improving runtime performance. The one or more processors may be further configured to optimise the workgraph by efficient resource management, by providing memory management of intermediate workgraph buffers using asynchronous CPU task execution delivers better efficient management of system memory resources because it carefully manages the (de)allocation lifetime of resources leading to improved performance. A further optimisation that may be achieved by the device and method of this disclosure is workgraph data pipelining, which can improve workgraph performance if a driver can schedule and overlap data transfers using CPU tasks and domain specific accelerator execution to maximize utilization.

[0092] The workgraph topology, which is the way pass nodes and edges are arranged within the graph, may be subject to prior optimizations before being submitted for execution in order to improve workgraph temporal / spatial execution locality of memory, domain operations merging (pass fusing) which keeps intermediate results on the domain specific accelerator / system cache for a subsequent access by downstream passes or workload slicing (or tiling) where the workload is processed one subset a time using the system cache which offers significant bandwidth.

[0093] The pass dependency list may be used to facilitate a distributed heterogeneous task execution (DHTE) where the workgraph passes are distributed for execution on both local and multiple remote computing platforms. The local system on chip platform, referred to as EDGE, may host the task scheduler and any number of domain specific accelerator device. The remote platforms, referred to as CLOUD, may host any number and variety of domain specific accelerator devices. In distributed heterogenous task execution, the hardware task scheduler may be located on the local system on chip platform, referred to as local hardware task scheduler (LHWTS), and may be located on remote computing resource, referred to as remote hardware task scheduler (RHWTS). The task scheduler configures how passes are distributed between edge and cloud platforms or it may defer the latter role to a remote task scheduler (RTS).

[0094] Standard or specialize network transport protocols may be used to connect the task scheduler to the remote task scheduler directly or indirectly by way of interposer software agents on both sides implementing a message passing network transport protocol for distributed heterogenous task execution. The domain specific accelerator devices located on the edge, referred to as local domain specific accelerator (LDSA) and domain specific accelerator devices located on the cloud, referred to as remote domain specific accelerator (RD SA) devices, are uniquely identified in the programming context. Such a context may also include all necessary support functionality needed to remotely address and query the execution state of the remote task scheduler, remote hardware task scheduler and domain specific accelerator devices. Remote domain specific accelerator devices may be used to provide edge platform capabilities missing on local domain specific accelerator devices or supplement it in order to augment its performance profile. In this arrangement, the sharding of workgraph passes to the edge and cloud platforms may take on any permutation and combination explicitly expressed by the application or by runtime. Remote workgraph pass execution may be handled completely by the local hardware task scheduler or by a combination of both local hardware task scheduler and remote hardware task scheduler, or exclusively by the remote hardware task scheduler where each is responsible for pass scheduling for their respective domain specific accelerator devices. Furthermore, there may be multiple instances of cloud platform available to the local hardware task scheduler for distributing passes for execution.

[0095] When scheduling workgraph passes for execution on cloud, the pass dependency list may undergo transformations such as bitpacking and / or compression to improve serialization efficiency before being uploaded to the remove task scheduler. Furthermore, all local domain specific accelerator sub-tasks may be omitted from the pass dependency list before being uploaded to the remote task scheduler to improve memory footprint and upload latency. Workgraph events that are resolved in the cloud may be bundled together with output buffers and sent back to the edge device or signalled independently.

[0096] The device of this disclosure is therefore capable of processing an application pipeline based on a workgraph and event information as described above. This disclosure also provides a method of doing the same as set out herein. The method may comprise receiving a workgraph representing one or more pass to be performed in processing an application pipeline. Each of the one or more pass comprising one or more task formed of one or more command of the application processing pipeline. The method may further include generating a pass dependency schema representing order in which the one or more pass in the application pipeline is to be executed. The pass dependency schema may comprise event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed. Each pass described herein may be thought of as comprised of commands to be executed during processing of the workgraph based on the pass dependency schema, which include the identification information and the event information. It should be understood that the functions described above may be performed as part of a method according to this disclosure, particularly the processes performed by the one or more processors of the device of this disclosure. The pass dependency schema described herein may be heterogeneous in nature and can be adapted to be employed on any domain specific accelerator therefore being applicable for use with multiple domain specific accelerators such as GPUs and CPUs etc. The method of this disclosure may further include the ability to modify the workgraph, independently of the pass dependency schema, to comprise standard domain operations configured to adapt the workgraph to be executable on non-host domain specific accelerators. In addition, the method performed by the device of this disclosure may also modify the pass dependency schema, independently of the workgraph, to represent a further order in which the one or more pass in the application pipeline is to be executed. This provides the advantage that the user can adapt the pass dependency list included associated with the pass dependency schema while not adapting the workgraph in order to change the priority of commands executed by the device.

[0097] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

[0098] Table 1

[0099] Table 2

Claims

CLAIMS1. A device for processing an application pipeline, the device comprising one or more processor configured to : receive a workgraph comprising one or more pass to be performed for processing the application pipeline, each of the one or more pass comprising one or more task formed of one or more command for processing the application pipeline; generate a pass dependency schema representing an order in which the one or more pass in the workgraph is to be processed when the application pipeline is executed; wherein the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema comprises event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed.

2. The device according to claim 1, wherein the pass dependency schema is encoded in a 32-bit or 64-bit hardware descriptor.

3. The device according to any preceding claim, wherein the pass dependency schema further comprises pass information representing the cyclic properties of the one or more pass.

4. The device according to any preceding claim, wherein the one or more processor is further configured to execute the workgraph on a domain specific accelerator.

5. The device according to claim 4, wherein the one or more processor is further configured to: segment the workgraph into one or more slices, and execute each of the one or more slices of the segmented workgraph.

6. The device according to any preceding claim, wherein the one or more processor is further configured to: generate a pass dependency list representing the dependencies of the one or more pass in the workgraph; provide the pass dependency list and / or the pass dependency schema to a hardware device suitable for execution.

7. The device according to claim 6, wherein one or more processors are configured to embed the workgraph in one or more command buffer associated with the hardware device.

8. The device according to claim 6, wherein the one or more processor is further configured to: modify the workgraph to comprise standard domain operations configured to adapt the workgraph to be executable on non-host domain specific accelerators, and / or modify the pass dependency list, independently of the workgraph, to represent a further order in which the one or more pass in the application pipeline is to be executed.

9. A method for processing an application pipeline, the method comprising: receiving a workgraph comprising one or more pass to be performed for processing the application pipeline, each of the one or more pass comprising one or more task formed of one or more command for processing the application pipeline; generating a pass dependency schema representing an order in which the one or more pass in the workgraph is to be processed when the application pipeline is executed;wherein the pass dependency schema is separate and decoupled from the workgraph, and the pass dependency schema comprises event information associated with each of the one or more pass, the event information representing whether each of the one or more pass has been executed.

10. The method according to claim 9, wherein the pass dependency schema is encoded in a 32-bit or 64-bit hardware descriptor.

11. The method according to claim 9 or 10, wherein the pass dependency schema further comprises pass information representing the cyclic properties of the one or more pass.

12. The method according to any one of claims 9 to 11, wherein the method further comprises executing the workgraph on a domain specific accelerator.

13. The method according to claim 12, wherein the method further comprises: segmenting the workgraph into one or more slices, and executing each of the one or more slices of the segmented workgraph.

14. The method according to any one of claims 9 to 13, wherein the method further comprises: generating a pass dependency list representing the dependencies of the one or more pass in the workgraph; providing the pass dependency list to a hardware device suitable for execution.

15. The method according to claim 14, wherein the method further comprises: embedding the pass dependency list and / or workgraph in one or more command buffer associated with the hardware device.

Citation Information

Patent Citations

  • Concurrent workload scheduling with multiple level of dependencies

    EP3869334A1

  • Asynchronous execution mechanism

    US10861126B1

  • Task graph scheduling for workload processing

    US20210373957A1