A new efficient encoding schema of work description for unified hetrogeneous task execution

A data-driven workgraph algorithm addresses inefficiencies in resource and state management in heterogeneous pipeline processing by capturing domain operations and adapting to heterogeneous hardware, resulting in improved performance and efficiency.

WO2025131262A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/086801
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently managing resources, states, and memory in heterogeneous pipeline processing for complex environments and scenes, leading to performance trade-offs and inefficiencies.

Method used

A data-driven approach is introduced, utilizing a workgraph algorithm that captures domain operations and resource descriptions, allowing for efficient processing and adaptation to heterogeneous hardware devices through templating and incremental updates.

Benefits of technology

This approach improves resource and state management, reduces unnecessary host and device work, and enhances heterogeneous pipeline processing, leading to improved performance and efficiency in managing complex environments and scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023086801_26062025_PF_FP_ABST
    Figure EP2023086801_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A device for heterogeneous pipeline processing, the heterogenous pipeline to be executed to process initial data, the device comprising one or more processors, the one or more processors configured to: generate a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node wherein each node is comprised of one or more pass; wherein the template workgraph algorithm is a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on heterogeneous dedicated hardware devices suitable for heterogenous processing and / or execution.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A NEW EFFICIENT ENCODING SCHEMA OF WORK DESCRIPTION FOR UNIFIED HETROGENEOUS TASK EXECUTION

[0002] FIELD OF THE INVENTION

[0003] This present disclosure relates to a device for heterogeneous pipeline processing. The present disclosure also relates to a method of heterogeneous pipeline processing.

[0004] BACKGROUND

[0005] Data-driven rendering is getting increasingly used to manage large and complex environments and scenes. In last few years there has been accelerated development in this direction. For example, Framegraph, Rendergraph, DirectX 12 workgraph to name a few. An example of the resources available in well know API such as Vulkan can be seen in Figure 1.

[0006] One of the most complex issues in trying to efficiently manage large and complex environments and scenes is the management of resources. It includes but is not limited to efficiently managing the resource visibility, resource utilization, resource reuse, resource lifetime, minimizing and batching resource barriers, managing and optimizing resource memory etc. Figure 2 includes an illustration of the resource lifetime for each pass where each pass has associated with it one resource A, B or C and each resource has an associated resource lifetime for that pass.

[0007] Updating resources is a key aspect of development in most game engines these days. Resource update can take a large portion of time spend when rendering. If possible, an application may try to reuse a resource, but resource reuse is not always possible and providing updates are generally expensive. Resources can be accessed in more indirect way. To efficiently manage resources, drivers need more information in advance of processing to provide better optimizations.

[0008] Another complex issue with regards to efficiently managing large and complex environments and scenes is the management of states. States may be the render state, shader state, resource state, dynamic state, compute state, etc. Even fewer state changes can result in full hardware command rebuild cost on a mobile system on chip (SoC). Most state changes are not templatized. That means multiple states structures are redefined for subsequent usages. Few types of state changes affect shader recompilation or pipeline re-creation. Packing (mostly unrelated) states becomes a problem and deriving a hardware state from multiple API state is increasingly difficult. As such, there is a performance trade-off between hardware managed direct states vs software managed indirect states. To efficiently manage states changes and reduce to fewer state changes, drivers need more information in advance in order to improve the current system.

[0009] A further issue that currently exists for efficiently managing large and complex environments and scenes is how the memory management is performed. Memory management which deals with limited device memory on a system, different alignment requirements, binding of memory to different resources and sub resources and handling too many allocations to name few. When comparing explicit vs implicit memory management, there is currently unnecessary double management of memory by both application and driver. As such, there is a need for decoupling memory from its views. Also, deferred allocation mechanism is needed for better work parallelism. To efficiently manage device memory, drivers need more information in advance in order to achieve optimizations.

[0010] Data redundancy and persistency across complex application pipeline is another complex issue. More often data is created and destroyed, without a way to reuse it. It could include data like compiled shaders, render states, shader resources and / or pass resources. Unnecessary resource creation and management can be avoided by knowing if the resource is not going to be used for a set of passes execution.

[0011] Command buffers defer the execution of the commands to be executed after recording those commands in a buffer upfront and submitting at a later stage through a command queue to be executed on a physical device. The reuse of command buffers can be problematic. Primary command buffer can be re-recorded every frame with a separate framebuffer bound to a render pass, and the secondaries could be reused from frame to frame or record one command buffer per swapchain image and reused. By default, command buffers cannot be reset and are effectively immutable once they have been recorded. Updating commands in secondary command buffers, even if there is small delta change between the two command buffers, requires significant host cycles to rebuild hardware commands using those command buffers. Figure 4 shows an example of secondary command buffers being used to account for the inability to reset command buffers.

[0012] Another issue is that applications normally need to circle back to host after getting result of every execution for next device to process. Such a process prevents streaming of data across multiple devices or to same device multiple times or batching and / or slicing work. Slicing of tasks 501 in order to condense the processing of such tasks can be seen as an example in Figure 5, which tasks 501 are sliced into slices. Preparation of a draw call is a resource intensive operation that includes setting up resources and changing internal settings on the hardware before calling draw call. Optimizing execution commands like multiple draw calls can be a hit or a miss. Like merging meshes that use the same material or combining multiple meshes into a single mesh or drawing things that appear multiple times in a scene or grouping large number of env objects. For internal work slicing to work at the correct granularity the slicing requires most resources to be of optimal slice size which is not often the case.

[0013] Application pipelines span more than just graphics and compute operations Sharing resources across multiple devices is challenging. That includes definition and negotiation of exchanged data formats and layouts which is not always optimal. At times a common synchronization primitive and / or work dependency management across devices is not present. An example of the structure of a heterogeneous pipeline can be seen in Figure 6 and includes tasks 601 of the processing pipeline arranged in a pipeline. For implementing an efficient streaming application pipeline with multiple domain operations requires device discovery, a construct to create common resources, a construct to create common synchronization primitives, an ability to slice / tile input / output resources through different domain operators.

[0014] In short, some prior methods can be summarised as follows. FrameGraph includes high level representation of each Graphics Operations (GOps) to render a scene and is built from scratch for every frame, e.g., no templating or caching is used. This approach uses a Code Driven architecture, with deferred rendering and is not device driven. Furthermore, the number of objects is high and adds to complexity to the approach. Such an approach also hides many low-level operations. There is a clear order of execution of graphics operations. This approach also uses automatic and optimized resource creation, sync and memory allocations. There is also control over lifetime and usage of a large portion of render resources. The same is true for a RenderGraph approach with the exception that there is no concept of workgraph frames as no templating used.

[0015] The workgraph DX12 approach is similar to the above approaches in that it includes high level representation of each graphics operations to render a scene and there is no concept of a workgraph frame. A graph of nodes where shader code at each node can request invocations of other nodes, without waiting for them to launch. In addition, the number of objects is high and adds to complexity to the approach. Such an approach also hides many low-level operations. There is a clear order of execution of graphics operations. However, this approach differs in that there is caching. SUMMARY

[0016] According to a first aspect of this disclosure there is provided a device for heterogeneous pipeline processing, the heterogenous pipeline to be executed to process initial data, the device comprising one or more processors, the one or more processors configured to: generate a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node wherein each node is comprised of one or more pass; wherein the template workgraph algorithm is a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on hardware devices suitable for heterogenous processing and / or execution. This provides the advantage that the workgraph algorithm can be customised to be employed on any appropriate hardware by being of generic form capable of being adapted to the constraints of hardware devices suitable for heterogenous processing and / or execution that will be used for execution.

[0017] The device as described above, wherein each of the one or more pass comprises one or more of: operation information comprising domain operations to be performed; resource data representing resources required to execute each of the one or more pass; and task data representing parameters associated with each of the one or more pass. This allows for passes, that may be included in nodes, to include various information and data.

[0018] The device as described above, wherein the one or more processors are further configured to: generate one or more command buffer associated with the hardware devices suitable for heterogenous processing or execution by compiling the template workgraph algorithm, and submit the generated command buffers to a command queue in the hardware devices suitable for heterogenous processing and / or execution for execution. This allows the device to adapt the workgraph algorithm to the specific hardware and as such provides a means for adapting the workgraph algorithm to the heterogeneous hardware devices for execution.

[0019] The device as described above, wherein the one or more processors are further configured to: identify one or more changes in the initial data; identify one or more node of the template workgraph algorithm to be updated based on the identified one or more changes in the initial data; update the one or more identified node based on the changes in the initial data; generate a first incremental workgraph algorithm based on the identified one or more changes in the initial data and / or the identified one or more node, the first incremental workgraph algorithm comprising the identified one or more changes in the initial data and / or one or more updated node; generate a further workgraph algorithm by updating the template workgraph algorithm based on the first incremental workgraph algorithm, wherein the further workgraph algorithm comprises one or more node of the template workgraph algorithm and one or more node of the first incremental workgraph algorithm. This provides the advantage that the template workgraph algorithm does not have to be regenerated from scratch and as a whole, instead this allows the template workgraph to be updated based on delta changes thus providing a more efficient process.

[0020] The device as described above, wherein the further workgraph algorithm is comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm. This allows the further workgraph algorithm to be a composite workgraph algorithm formed of the nodes of both the template and the incremental workgraph algorithms therefore providing an improved and updated workgraph algorithm.

[0021] The device as described above, wherein the one or more processors are further configured to: store in a memory the template workgraph algorithm and / or the first incremental workgraph algorithm; and dynamically load from the memory the template workgraph algorithm and / or the first incremental workgraph algorithm for processing the pipeline. This allows for the workgraph algorithms to be loaded without the need to generate a new workgraph algorithm each time one is required. The device as described above, wherein one or more of: the template workgraph algorithm; the first incremental workgraph algorithm; the further workgraph algorithm; each node of the template workgraph algorithm; each node of the first incremental workgraph algorithm; or each node of the further workgraph algorithm, comprise a unique identifier. This allows each of the workgraph algorithms to be uniquely identifiable and subsequently their nodes to be uniquely identifiable such that they may be identified when multiple workgraph algorithms are employed.

[0022] The device as described above, wherein the one or more processors are further configured to: generate a second incremental workgraph algorithm, wherein the second incremental workgraph algorithm comprises changes in one or more node of the template workgraph algorithm different from the changes in the one or more node of the further workgraph algorithm; update the further workgraph algorithm by combining the second incremental workgraph algorithm, the first incremental workgraph algorithm and the template workgraph algorithm; and generate one or more command buffer associated to the hardware devices suitable for heterogenous processing and / or execution by compiling the updated further workgraph algorithm. This allows for an iterative process to be employed for generating incremental workgraph algorithms and updating the template workgraph algorithm to produce an updated (further) workgraph algorithm.

[0023] According to a further aspect of this disclosure there is provided a method of processing the heterogenous pipeline to be executed to process initial data, the method comprising: generating a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node wherein each node is comprised of one or more pass; wherein the template workgraph algorithm is a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on hardware devices suitable for heterogenous processing and / or execution. This provides the advantage that the workgraph generated can be customised to be employed on any appropriate hardware by being of generic form capable of being adapted to the constraints of heterogenous hardware devices that will be used for execution.

[0024] The method as described above, wherein each of the one or more pass comprises one or more of: operation information comprising domain operations to be performed; resource data representing resources required to execute each of the one or more pass; and task data representing parameters associated with each of the one or more pass. This provides the advantage that the nodes may comprise passes that represent tasks to be performed as part of the workgraph algorithm.

[0025] The method as described above, wherein the method further comprises: generating one or more command buffer associated with the hardware devices suitable for heterogenous processing or execution by compiling the template workgraph algorithm, and submitting the generated command buffers to a command queue in the hardware devices suitable for heterogenous processing and / or execution for execution. This allows the device to adapt the workgraph algorithm to the specific dedicated hardware and as such provides a means for adapting the workgraph algorithm to the heterogeneous devices for execution.

[0026] The method as described above, wherein the method further comprises: identifying one or more changes in the initial data; identifying one or more node of the template workgraph algorithm to be updated based on the identified one or more changes in the initial data; updating the one or more identified node based on the changes in the initial data; generating a first incremental workgraph algorithm based on the identified one or more changes in the initial data and / or the identified one or more node, the first incremental workgraph algorithm comprising the identified one or more changes in the initial data and / or one or more updated node; generating a further workgraph algorithm by updating the template workgraph algorithm based on the first incremental workgraph algorithm, wherein the further workgraph algorithm comprises one or more node of the template workgraph algorithm and one or more node of the first incremental workgraph algorithm. This provides the advantage that the template workgraph algorithm does not have to be regenerated from scratch and as a whole, instead this allows the template workgraph to be updated based on delta changes thus providing a more efficient process. The method as described above, wherein the further workgraph algorithm is comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm. This allows the further workgraph algorithm to be a composite workgraph algorithm formed of the nodes of both the template and the incremental workgraph algorithms therefore providing an improved and updated workgraph algorithm.

[0027] The method as described above, wherein the method further comprises: storing in a memory the template workgraph algorithm and / or the first incremental workgraph algorithm; and dynamically loading from the memory the template workgraph algorithm and / or the first incremental workgraph algorithm for processing the pipeline. This allows for the workgraph algorithms to be loaded without the need to generate a new workgraph algorithm each time one is required.

[0028] The method as described above, wherein one or more of: the template workgraph algorithm; the first incremental workgraph algorithm; the further workgraph algorithm; each node of the template workgraph algorithm; each node of the first incremental workgraph algorithm; or each node of the further workgraph algorithm, comprise a unique identifier. This allows each of the workgraph algorithms to be uniquely identifiable and subsequently their nodes to be uniquely identifiable such that they may be identified when multiple workgraph algorithms are employed.

[0029] BRIEF DESCRIPTION OF THE FIGURES

[0030] The embodiments of the present disclosure will now be described by way of example with reference to the accompanying drawings. In the drawings:

[0031] Figure 1 illustrates an example of resources in Vulkan;

[0032] Figure 2 illustrates an example resource lifetime;

[0033] Figure 3 illustrates an example of the memory binded resources in Vulkan;

[0034] Figure 4 illustrates an example of Vulkan Primary and Secondary Command Buffers;

[0035] Figure 5 illustrates an example of work slicing work scheduling:

[0036] Figure 6 illustrates an example of heterogeneous pipeline processing;

[0037] Figure 7 illustrates an example of a data driven workflow,

[0038] Figure 8 illustrates an example of a data driven workflow for template data;

[0039] Figure 9 illustrates an example of a data driven workflow for incremental data;

[0040] Figure 10 illustrates an example of nodes for a workgraph;

[0041] Figure 11 illustrates an example of node hierarchy;

[0042] Figure 12 illustrates an example of the structure of pass code;

[0043] Figure 13 illustrates an example of the structure of subpipe code;

[0044] Figure 14 illustrates an example of the structure of shader code;

[0045] Figure 15 illustrates an example of the structure of subpipe state code;

[0046] Figure 16 illustrates an example of the structure of subpipe resource type code;

[0047] Figure 17 illustrates an example of the structure of scene node code;

[0048] Figure 18 illustrates an example of the structure of buffer code;

[0049] Figure 19 illustrates an example of the structure of image code;

[0050] Figure 20 illustrates an example of the structure of tensor code;

[0051] Figure 21 illustrates an example of the structure of scene code;

[0052] Figure 22 illustrates an example of the structure of camera perspective code;

[0053] Figure 23 illustrates an example of the structure of mesh code;

[0054] Figure 24 illustrates an example of the structure of material code; Figure 25 illustrates an example of the structure of accessor code;

[0055] Figure 26 illustrates an example of the structure of memory code;

[0056] Figure 27 illustrates an example of the structure of memory view code;

[0057] Figure 28 illustrates an example of a workgraph data used to render a coloured triangle;

[0058] Figure 29 illustrates an example of a deep learning super sample (DLSS) heterogenous processing pipeline expressed in workgraph format;

[0059] Figure 30 illustrates an example of a workgraph with an inbuilt caching mechanism; and

[0060] Figure 31 illustrates an example of a workgraph as a template.

[0061] DETAILED DESCRIPTION OF THE INVENTION

[0062] This disclosure provides methods and techniques to overcome the above problems discussed above. This is achieved by providing herein a data-driven approach, which unlike a code driven approach, decouples logic and resources from runtime API (separate data from systems) to address the above-mentioned problems.

[0063] The data driven approach implemented by the device and method of this disclosure provides an efficient novel way to implement and support newly proposed API extensions targeting a mobile System on Chip for 3D graphics, computing and machine learning applications. In the case of this disclosure a new API extension will be described by way of example only, however the present disclosure could be applied to any customisable API as appropriate. Such an API extension may be applicable to a Vulkan API.

[0064] This disclosure introduces a new way of capturing work description in a new compact, efficient to process, structured and formatted data here by called a workgraph algorithm. A workgraph algorithm may be data that captures a list of graph node elements that may represent representing domain operations and a description of the related resources. Workgraph algorithm capture the user’s algorithmic intent and overall structure, without burdening the developer to know too much about the specific hardware it will run on. Workgraph algorithms may therefore provide a description of the tasks that are to be performed as part of an application pipeline and may include one or more passes or descriptions of other node elements that may form the application processing pipeline.

[0065] By compiling these workgraph algorithms it is possible to generate command buffers that can be submitted using known indirect deferred execution methods for example VkQueueSubmit application programming interface (API) in Vulkan. In this way it is possible to decouple logic and resources from runtime application programming interface (API), thereby separating data from systems. Normally, authoring data is optimized for flexibility while runtime data is optimized for performance. New data driven flow is designed with this core principle in mind i.e., more context one has the better performant and optimal solution one can make.

[0066] The workgraph algorithms of this disclosure use a new templating design to keep the data size low for representing only delta changes across frames / region of interest. This enables device driven flow bypassing host (where device generates work for itself and / or another device without involving host). This also enables exchanging data between edge and cloud or device and end point terminal. It supports work description containing multiple domain operations (Graphics, Compute, Machine Learning, Vision, Image Processing ... ) describing a heterogeneous application pipeline.

[0067] In an example implementation described herein there are three steps that can be used as part of the new data driven workflow. These three steps may be building or loading (generating), compiling and executing. To achieve this new data driven flow, this disclosure introduces a new object, which may be a Vulkan object, that represents a handle for new compact, efficient to process, structured and formatted data containing workgraph algorithms (work description). Using this handle one can load the preinitialized work description data stored in a file, this is known as a workgraph algorithm which is a binary representation of a data structure (work description data structure). Since there is not no concept of frame and support beyond render, such an object can be considered a workgraph algorithm. These workgraph algorithms represent both templated and incremental change data and can, in some cases, be exported in human readable (json) and / or binary file formats. Such workgraph algorithms can be built online or offline and can be compiled to generate command buffers or they can be cached.

[0068] The present disclosure provides a device for heterogeneous pipeline processing, the heterogenous pipeline comprising one or more node to be executed to process initial data, the device comprising one or more processors. The one or more processors may be configured to generate a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node where each node is comprised of one or more pass. Generating in this case may mean that the template workgraph algorithm is generated by the one or more processor or received from outside the device. The term generated herein is also in some cases, intended to include the loading of the workgraph algorithms from a memory that is either provided externally or within the device. In other words, it is possible to dynamically load from a memory the template workgraph algorithm and / or the first incremental workgraph algorithm for processing the pipeline.

[0069] The heterogeneous pipeline may include one or more tasks that are to be performed as part of the pipeline, for example if the pipeline is configured to perform image processing then the tasks that may be performed may be tasks such as sharpening, denoising, bilinear filtering etc. Each of these tasks may be represented by the template workgraph algorithm which provides a work description of these tasks. The tasks performed by the pipeline may each be implemented by one or more node comprised within the template workgraph algorithm. An example of this is that each task of the pipeline may require that specific nodes of the workgraph to be executed in order to perform the task. To generate a workgraph algorithm the device of this disclosure may use an application programming interface (API) to generate data structures that describe the relationship between the one or more node of the heterogeneous pipeline. As such, the one or more task that the heterogeneous pipeline is to perform informs the structure of the nodes comprised in the generated workgraph algorithm.

[0070] The template workgraph algorithm may comprise one or more node, which may in some cases comprise one or more pass and each of the one or more pass comprises one or more of: operation information representing domain operations to be performed. The one or more pass may also comprise resource data representing the resources required to execute each of the passes of the pipeline. The resources that may be comprised in the one or more pass may include sub pipe information that may be made up of constant values and other variables that may be shader programmable. The constant values and customisable variables that may be comprised in the one or more pass may be thought of as the components and inputs / outputs that are required for the pass to be executed as part of the pipeline.

[0071] The one or more pass may also comprise task data representing parameters associated with each of the one or more pass. The parameters represented in the task data may comprise input and output data for each pass that may be well defined and then further parameters that may be adaptable depending on the function to be performed by the pass. As such, the parameters comprised in the task data may vary depending on the function of the pass and may aid in executing the pass.

[0072] The nodes may comprise further components other than passes. Some examples of further components of nodes for a workgraph algorithm of this disclosure can be seen in Figure 10 and Table 1, described later. The template workgraph algorithm may be a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on hardware devices suitable for heterogenous processing and / or execution. The hardware devices suitable for heterogenous execution and / or processing may comprise heterogenous dedicated hardware devices or hardware devices which can be configured or programmed or adapted for heterogenous execution or processing. In other words, the template workgraph algorithm (and the incremental and further workgraph algorithms described below) are data structure files in binary format that describe the configuration of nodes in the heterogenous pipeline needed for the heterogeneous pipeline to perform the processing function. This format of the workgraph algorithm allows the workgraph algorithms of this disclosure to capture the configuration of the processing pipeline in a generic manner that can be provided to any of a number of hardware devices suitable for heterogenous processing and / or execution, for example Central Processing Units (CPU), Graphics Processing Units (GPU), etc. The workgraph algorithms can therefore be applied to any of these hardware devices in order to be executed. In other words, the workgraph algorithms as described herein can be customised by a compiler to be applicable to a domain specific accelerator such a CPU, GPU etc. In this disclosure there are primarily three types of workgraph algorithms described, the first of which is a template workgraph algorithm, which is comprised of a number of nodes and once built or loaded (generated / received) it can be used to process the initial data as part of the heterogeneous pipeline. The second workgraph algorithm is a first incremental workgraph algorithm that is received more frequently and only comprised of a subset of the nodes of the template workgraph algorithm that have changed during the processing of data. The third workgraph algorithm is an updated version of the template workgraph algorithm or of a previously generated updated workgraph algorithm. This further workgraph algorithm represents a workgraph algorithm that has been updated using an incremental workgraph algorithm to represent the changes to the previous workgraph algorithm and thus provide the most up to date workgraph algorithm. The further workgraph algorithm may be comprised of nodes from one or more of the templates workgraph algorithm and / or the first incremental workgraph algorithm, and may also be comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm. Each of these workgraph algorithms (workgraph algorithms) may be importantly of a form that can be adapted to be employed on any dedicated hardware desired by the user. In this way the workgraph algorithms are of a form that is not specific to one dedicated hardware in particular and is of a more general form such that they are adaptable to be heterogeneous dedicated hardware.

[0073] The template workgraph algorithm can be a complete description of the processing to be performed and thus can be generated less frequently, preferably at the initial stage of the pipeline. The incremental workgraph algorithm may only be applied over the top of a template workgraph algorithm. In other words, the incremental workgraph algorithm may represent the identified one or more passes of the template workgraph algorithm to be updated and thus can be used to update the passes of the template workgraph algorithm when applied over it. An incremental workgraph algorithm relative to template workgraph algorithm records only the delta changes. The first incremental workgraph algorithm may be generated based on the identified one or more changes in the initial data and / or the identified one or more node, and the first incremental workgraph algorithm may comprise the identified one or more changes in the initial data and / or one or more updated node. Template workgraph algorithm can be built or loaded less frequently as compared to incremental workgraph algorithm which can be built or loaded more frequently but can only be applied over template. Staging template area internally can be used to patch the incremental delta changes on preloaded template. For each hardware command list instance can be created using a patched template area.

[0074] Such a construction of workgraph algorithms as described herein allows the workgraph construction to be reused between workgraph algorithms and between runs of an application. It reuses the workgraph algorithm by passing the same cache object when creating multiple related workgraph algorithms. This saves building or compiling time. Contents of these caches may be managed by the implementations. The data driven flow implemented by the device and method of this disclosure introduces a workgraph indirect execution method which may be a three-step process that the work graph system may complete, from scratch and every frame or region of interest. This is because an incremental workgraph relative to template workgraph records only the delta changes. Staging template area internally can be used to patch the incremental delta changes on preloaded template. Each hardware command list instance can be created using a patched template area. Figure 7 describes the data driven flows as described in this disclosure. Flow 1 and 2 as shown in Figure 7 are examples of the three-step process that may be performed by the processing apparatus and method of the present disclosure. In particular the template workgraph algorithm may be built 701 or loaded 704 (both may be considered as generating according to this disclosure) in an initial step. The template workgraph algorithm may then be compiled, or in other words adapted, to the domain specific accelerator (hardware devices suitable for heterogenous processing and / or execution) that is desired to be used and then send for execution. The one or more processors may therefore be further configured to: compile 702 the template workgraph algorithm to generate one or more command buffer associated to selected dedicated hardware. The hardware devices suitable for heterogenous processing and / or execution may be configured to execute 703 the generated command buffers by submitting the command buffers to a command queue. In other words, the one or more processors may be configured to submit the generated command buffers to a command queue in the hardware devices suitable for heterogenous processing and / or execution for execution. The method may further include generating further incremental workgraph algorithms representing changes in one or more pass of the template workgraph algorithm different from the changes in the one or more pass of the template workgraph algorithm represented by the first incremental workgraph algorithm; combining the further incremental workgraph algorithms with the first incremental workgraph algorithm and the template workgraph algorithm to form the further workgraph algorithm; and compiling the further workgraph algorithm to generate one or more command buffer associated to selected dedicated hardware (heterogenous dedicated hardware device chosen by the user). This allows for an iterative process to be employed for generating incremental workgraph algorithms and updating the template workgraph algorithm to produce an updated (further) workgraph algorithm.

[0075] The process of Figure 7 is described in more detail in relation to Figure 8. The first step in the process is to build (generate) all the passes in a template work graph(s). This is where one records all the passes including operations, resources and data used by each of these passes. Figure 8 demonstrates a process that may be performed by the one or more processors of the present disclosure or the method. As described above the template workgraph algorithm may be built 801 (generated) or loaded 804 in an initial step. When loading the workgraph algorithm this may be done dynamically from a workgraph algorithm stored in a memory. The template workgraph may then be compiled 802 to make it compliant for the dedicated hardware that will be used the execute 803 the passes. During compiling 802 of the template workgraph algorithm cache data 805 may be requested and provided by the one or more processors. Once the template workgraph algorithm has been compiled 802 it can be submitted to command buffers of the dedicated hardware in order to be executed 803.

[0076] Compiling the workgraph algorithms of this disclosure may include using a hardware or software platform including a driver to read the workgraph algorithm and translate the generic workgraph algorithm from its binary memory representation of a data structure to a format that is specific to the hardware devices suitable for heterogenous processing and / or execution that are intended to execute the workgraph algorithm. In doing this the command buffers are generated which are formatted to be of a format acceptable to the user’s chosen dedicated hardware device.

[0077] Figure 9 provides an alternate example in which a first incremental workgraph algorithm may be generated (901 , 904) and used to update the template workgraph algorithm in the device of this disclosure. The first incremental workgraph algorithms being generated as described here is intended to take the same meaning as the template workgraph algorithm in that this generating may in some cases include both the building and the loading of the workgraph algorithm. The same is true for the further workgraph algorithm described later. In this case the template workgraph algorithm may be loaded 804 or built 801 as in Figure 8, however, prior to compiling the template workgraph algorithm the one or more processors may be further configured to identify one or more changes in the initial data and identify one or more node of the template workgraph algorithm to be updated based on the identified one or more changes in the initial data. The one or more processors may then be configured to update the one or more identified node based on the changes in the initial data and generate (901, 904) a first incremental workgraph algorithm based on the identified one or more changes in the initial data, the first incremental workgraph algorithm comprising the identified one or more changes in the initial data and one and / or more updated node. Once the incremental workgraph algorithm has been generated it can be used by the one or more processors to update the template workgraph algorithm based on the first incremental workgraph algorithm to produce a further workgraph algorithm (which may happen at the compiling of incremental workgraph alongside the template workgraph 906), wherein the further workgraph algorithm comprises one or more node of the template workgraph algorithm and one or more node of the first incremental workgraph algorithm. Incremental workgraph algorithms can be generated to update the template workgraph algorithm or the further algorithm in an iterative manner to ensure the workgraph is constantly updated of any changes. One template workgraph algorithm may have several incremental workgraph algorithms applied to it to generate multiple further workgraph algorithms. The further workgraph algorithm may be comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm. In the iterative updating the one or more processors may be further configured to: generate a second incremental workgraph algorithm wherein the second incremental workgraph algorithm comprises changes in one or more node of the template workgraph algorithm different from the changes in the one or more node of the further workgraph algorithm. The second incremental workgraph algorithm may be the same format as the template workgraph algorithm and the first incremental workgraph algorithm. In some cases, it may represent further changes in the nodes of the template workgraph algorithm and / or first incremental workgraph but in other cases it may represent only the changes in the nodes of the further workgraph algorithm. Once these are generated the one or more processors may combine the second incremental workgraph algorithm with the first incremental workgraph algorithm and the template workgraph algorithm to update the further workgraph algorithm. Finally, the one or more processors may be configured to generate command buffers associated to selected dedicated hardware by compiling the further workgraph algorithm.

[0078] Once the template workgraph algorithm has been updated to produce the further workgraph algorithm, the next step as shown in Figure 9 is to compile 902 the workgraph algorithms to generate command buffers. As in Figure 8 cache data 905 may be used as part of this process. During this step, the workgraph algorithm is traversed, analysed, translated to record and encode hardware specific commands inside command buffers. For each pass, the workgraph algorithm system creates the all pass resources and releases them after the pass is executed if later passes do not use them. Finally, the last step is to execute 903 the command buffers generated from this workgraph compilation step by submitting these command buffers to a command queue. The workgraph algorithm executes all passes in declaration order.

[0079] Each of the workgraph algorithms are comprised of nodes which may be comprised of passes as discussed briefly above. However, the nodes may also be comprised of other components. Some examples of nodes can be seen in Table 1 , which details that nodes may also be comprised of subpipes, shaders, sceneNodes etc. Figure 10 demonstrates a workgraph comprised of nodes 1001 for each process the workgraph will perform. For example, the input / output process includes scene, buffer, image, tensor and accelstructs nodes 1001. Figure 11 demonstrates how each of the nodes 1101 of the example of Figure 10 may be connected to each other in a node hierarchy. Such a hierarchy describes what each of the nodes depend upon.

[0080] Code representing each of the nodes will now be described in relation to Figures 12 to 26.

[0081] Figure 12 demonstrates an example of the structure of pass code. A pass represents a set of domain operations in form of subpipe and its inputs and outputs as pass resources. In the code Pass.src objects represent a list of tuples of input resource type and input resource id for the given Pass. Pass.dst objects represent a list of tuples of output resource type and output resource id for the given pass. Pass. subPipe may be the index of the subPipe used for the given Pass.

[0082] Figure 13 demonstrates an example of the structure of subpipe code. SubPipe. states Subpipe State objects represents a set of state information for the given subpipe. SubPipe.resources Subpipe are resource objects represents a set of resource information for the given subpipe. SubPipe. subPipeType specifies the type of subpipe used. Based on this, either the GRAPHICS, COMPUTE, or ML are defined. A subPipe represents a set of operations for a domain.

[0083] Figure 14 demonstrates an example of the structure of shader code. A shader represents a set of programmable elements within a given subpipe. Shader. shaderType Specifies the type of shader used. Based on this, either the VERTEX_SELADER, FRAGMENT_SHADER, COMPUTE_SHADER or ML_SHADER are defined. Shader.memory View represents the index of the memory View used for the shader. Shader. name The index of the subPipe used for the given pass.

[0084] Figure 15 demonstrates an example of the structure of Subpipe state code. A state represents a set of states that can be set for a given subpipe. SubPipeState. stateType defines the possible types of subpipe states that can be used.

[0085] Figure 16 demonstrates an example of the structure of Subpipe resource code. A resource represents a set of resource that can be set for a given subpipe. SubPipeResource.resourceType defines the possible type of subpipe resources that can be used.

[0086] Figure 17 demonstrates an example of the structure of scene node code. A scene node represents a set of visual objects comprising the scene to render. Nodes may have transformed properties. Nodes are organized in a parent-child hierarchy. Node hierarchy does not contain cycles. Each node may have zero or one parent node. Some examples of scene nodes are SceneNode. children which indexes the subPipe used for the given pass. SceneNode.matrix indexes the subPipe used for the given pass. Other examples are SceneNode. scale, SceneNode.translation, SceneNode.weights.

[0087] Figure 18 demonstrates an example of the structure of buffer code. A buffer represents an arbitrary data stored in linearly in memory and can be used as an input and / or output pass resource type for execution of a pass. Some examples of buffer node codes are Buffer.memory View wherein the index of the memory View used for the buffer. Buffer. accessorType which specifies if the buffer elements are scalars, vectors, or matrices. Buffer.componentType describing the datatype of the buffer components Buffer.byteOffset which describes the offset of components.

[0088] Figure 19 demonstrates an example of the structure of image code for that node. Images referred to by textures are stored in the images array of the asset. Each image contains one of a URI to an external file in one of the supported image formats, or a data URI with embedded data, or a reference to a memory view; in that case mimeType MUST be defined.

[0089] Figure 20 demonstrates an example of the structure of tensor code for that node. Some examples of tensor node codes are Tensor.accessorType which specifies if the tensor elements are scalars, vectors, or matrices. Tensor.componentType that describes the datatype of the tensor components and Tensor, shape which describes the shape of the tensor.

[0090] Figure 21 demonstrates an example of the structure of scene code for that node. This node object may comprise a scene to render.

[0091] Figure 22 demonstrates an example of the structure of camera code for that node. A camera’s projection to apply a transform to place the camera in the scene, which may have transformed properties. Figure 23 demonstrates an example of the structure of mesh code for that node. Mesh code may describe geometry as arrays of primitives. Primitives correspond to the data required for GPU draw calls Primitives specify one or more attributes, corresponding to the vertex attributes used in the draw calls. Indexed primitives also define an indices property. Attributes and indices are defined as references to accessors containing corresponding data. Each primitive may also specify a material and a mode that corresponds to the GPU topology type (e.g., triangle set). Any scene node may contain one mesh Splitting one mesh into several primitives can be useful to limit the number of indices per draw call or to assign different materials to different parts of the mesh.

[0092] Figure 24 demonstrates an example of the structure of material code for that node. Figure 25 demonstrates an example of the structure of accessor code for that node. Figure 26 demonstrates an example of the structure of memory code for that node. A memory is arbitrary data stored as a binary blob. The memory may contain any combination of data. The URI property is the URI to the memory data.

[0093] Figure 27 demonstrates an example of the structure of memory view code for that node. A view into a memory generally represents a subset of the memory. A memory view represents a contiguous segment of data in a memory, defined by a byte offset into the memory specified in the byteOffset property and a total byte length specified by the byteLength property of the memory view. All memory data should use little endian byte order.

[0094] Figure 28 demonstrates a workgraph that may be generated or stored to be loaded in order to generate a rendered coloured triangle, once compiled and executed. As can be seen in the code above and the workgraph of Figure 28 one or more of the templates workgraph algorithm, the incremental algorithm; the further algorithm; each of the nodes of the template workgraph algorithm; each of the nodes of the incremental workgraph algorithm; or each of the nodes of the template further algorithm may comprise a unique identifier. The unique identifier allows each of the nodes to be attributed to one of the workgraph algorithms processed by the one or more processors and method of this disclosure. Figure 28 shows each of the nodes that would be required to render such a coloured triangle as in the example and the processes required by each of the nodes similarly to that shown in Figure 10. In some cases, each of the workgraph algorithms described in this disclosure may be assigned a unique identifier. That is one or more of: the template workgraph algorithm, the first incremental workgraph algorithm; the further workgraph algorithm; each node of the template workgraph algorithm; each node of the first incremental workgraph algorithm; or each node of the further workgraph algorithm, comprise a unique identifier.

[0095] Figure 29 illustrates an example of a deep learning super sample (DLSS) heterogenous processing pipeline expressed in workgraph format. Due to the use of the workgraph algorithm format of this disclosure it is possible to utilise the unique identifiers in each pass to determine the processes in the workgraph algorithm that will be required for subsequent passes. For example, as shown in Figure 29 a sceneid 0 and buffered 0 are input to a first pass graphics subpipe to produce an imageid 0. In a second pass a tensor having id 0 is modified to produce a tensor with id 1. Then subsequently, a third pass of the workgraph algorithm takes both the image id 0 and the buffer id 1 (associated with tensor id 1) and computes a subpipe 2 to produce a image with id 1. This is also demonstrated by the flow chart of Figure 29, wherein the image generated in the first pass bypasses the second pass and proceeds to be used in the third pass Gaussian filter.

[0096] Figure 30 illustrates an example workgraph cache that allows such a workgraph algorithm of Figure 29 to be performed. As such, WorkGraphCache allows the result of workgraph construction to be reused between workgraph algorithms and between runs of an application. The workgraph algorithm cache can be reused by passing same cache object when creating multiple related workgraph algorithms. This saves time because the workgraph need not be rebuilt or recompiled. Contents of these caches are expected to be managed by the implementations. In the workgraph cache the workgraph algorithm may be associated with a hardware descriptor such that it is ready to be used. Figure 31 provides an example of a workgraph algorithm template being updated based on delta changes in a staging area to produce commands that may be stored in the memory for execution. An incremental workgraph algorithm may produce / store delta changes in the workgraph algorithm as shown in Figure 31 and the incremental graph algorithm is always relative to template workgraph records only the delta changes. Template workgraph algorithm can be built or loaded less frequently as compared to incremental workgraph algorithm which can be built or loaded more frequently but can only be applied over template. Staging template area internally can be used to patch the incremental delta changes on preloaded template. In other words, can be used to update the template workgraph algorithm using the incremental workgraph algorithm to produce the further workgraph algorithm representing the updated template workgraph algorithm. For each hardware command list instance can be created using a patched template area.

[0097] The data driven approach as described herein provides a number of advantages over a code driven approach. Some of these advantages are that it presents a simplified method of solving complex problems, improves resource management {visibility, utilization, reuse, lifetime, caching, etc}, improves state management {reduces draw calls and state changes where possible}. Further advantages are that a data driven approach improves scene management {passes, materials, meshes, camera, etc}, data management {command buffer, state, resource, scene} - reuse. Hardware utilization is improved by internal work slicing, and performance is also improved by eliminating unnecessary host and device work. A data driven approach also improves heterogeneous pipeline processing and data sharing and provides a better solution for parallel stream processing with ability to distribute work across heterogeneous hardware devices. In addition, lower processing execution overhead, lower hardware and memory resource utilization and lower-level synchronization for lower processing latency and hardware scheduling can be achieved.

[0098] A data drive approach also supports heterogeneous multi passes, a templating design, supports filesystem interaction, unified work representation and management, allows reuse of pipeline states and resources table, allows cross- vendor export for later use possible with a standard data exchange and fixed data format, changes after recording are possible, implicit efficient object, memory and synchronization management with more context can also be achieved. In addition, it supports 32-bit ids as compared to 64-bit handles, fully self-contained data, high-level representation for better work description, possibility of post optimization once built and recorded, provides enough context for high level optimizations, reduces validation errors by simplifying command buffer recording and generation, device driven feature built into the design and provides more support for indirect methods.

[0099] The workgraph Vulkan Extension described herein includes high level representation of each domain operation to render, compute, machine learn, etc. or describe a stream heterogenous application pipeline. There is also no concept of frame, built only when the graph changes, and a templating design is used. Caching is used. This is a data driven arch, including deferred work operations and device driven supported. It is also simplified with fewer set of objects and has the capability to hide many low-level operations. Furthermore, there is clear order of execution of domain operation, with lower overhead. This approach also includes automatic and optimized resource creation, sync, and memory allocations, automatic work batching, work slicing and work sync. Finally, this approach provides full control over lifetime and usage of a large portion of pass or shader resources.

[0100] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

[0101]

[0102] Table 1

Claims

CLAIMS1. A device for heterogeneous pipeline processing, the heterogenous pipeline to be executed to process initial data, the device comprising one or more processors, the one or more processors configured to: generate a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node wherein each node is comprised of one or more pass; wherein the template workgraph algorithm is a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on hardware devices suitable for heterogenous processing and / or execution.

2. The device according to claim 1, wherein each of the one or more pass comprises one or more of: operation information comprising domain operations to be performed; resource data representing resources required to execute each of the one or more pass; and task data representing parameters associated with each of the one or more pass.

3. The device according to any preceding claim, wherein the one or more processors are further configured to : generate one or more command buffer associated with the hardware devices suitable for heterogenous processing or execution by compiling the template workgraph algorithm, and submit the generated command buffers to a command queue in the hardware devices suitable for heterogenous processing and / or execution for execution.

4. The device according to any preceding claim, wherein the one or more processors are further configured to: identify one or more changes in the initial data; identify one or more node of the template workgraph algorithm to be updated based on the identified one or more changes in the initial data; update the one or more identified node based on the changes in the initial data; generate a first incremental workgraph algorithm based on the identified one or more changes in the initial data and / or the identified one or more node, the first incremental workgraph algorithm comprising the identified one or more changes in the initial data and / or one or more updated node; generate a further workgraph algorithm by updating the template workgraph algorithm based on the first incremental workgraph algorithm, wherein the further workgraph algorithm comprises one or more node of the template workgraph algorithm and one or more node of the first incremental workgraph algorithm.

5. The device according to claim 4, wherein the further workgraph algorithm is comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm.

6. The device according to any preceding claim, wherein the one or more processors are further configured to: store in a memory the template workgraph algorithm and / or the first incremental workgraph algorithm; and dynamically load from the memory the template workgraph algorithm and / or the first incremental workgraph algorithm for processing the pipeline.

7. The device according to claim 4, wherein one or more of: the template workgraph algorithm, the first incremental workgraph algorithm;the further workgraph algorithm; each node of the template workgraph algorithm; each node of the first incremental workgraph algorithm; or each node of the further workgraph algorithm, comprise a unique identifier.

8. The device according to claim 4, wherein the one or more processors are further configured to: generate a second incremental workgraph algorithm, wherein the second incremental workgraph algorithm comprises changes in one or more node of the template workgraph algorithm different from the changes in the one or more node of the further workgraph algorithm; update the further workgraph algorithm by combining the second incremental workgraph algorithm, the first incremental workgraph algorithm and the template workgraph algorithm; and generate one or more command buffer associated to the hardware devices suitable for heterogenous processing and / or execution by compiling the updated further workgraph algorithm.

9. A method of processing the heterogenous pipeline to be executed to process initial data, the method comprising: generating a template workgraph algorithm, for initially processing the initial data, based on one or more task performed by the heterogeneous pipeline, the template workgraph algorithm comprising one or more node wherein each node is comprised of one or more pass; wherein the template workgraph algorithm is a binary memory representation of a data structure formed of the one or more node, where the template workgraph algorithm is adaptable for execution on hardware devices suitable for heterogenous processing and / or execution.

10. The method according to claim 9, wherein each of the one or more pass comprises one or more of: operation information comprising domain operations to be performed; resource data representing resources required to execute each of the one or more pass; and task data representing parameters associated with each of the one or more pass.

11. The method according to any one of claims 9 or 10, wherein the method further comprises: generating one or more command buffer associated with the hardware devices suitable for heterogenous processing or execution by compiling the template workgraph algorithm, and submitting the generated command buffers to a command queue in the hardware devices suitable for heterogenous processing and / or execution for execution.

12. The method according to any one of claims 9 to 11, wherein the method further comprises: identifying one or more changes in the initial data; identifying one or more node of the template workgraph algorithm to be updated based on the identified one or more changes in the initial data; updating the one or more identified node based on the changes in the initial data; generating a first incremental workgraph algorithm based on the identified one or more changes in the initial data and / or the identified one or more node, the first incremental workgraph algorithm comprising the identified one or more changes in the initial data and / or one or more updated node; generating a further workgraph algorithm by updating the template workgraph algorithm based on the first incremental workgraph algorithm, wherein the further workgraph algorithm comprises one or more node of the template workgraph algorithm and one or more node of the first incremental workgraph algorithm.

13. The method according to claim 12, wherein the further workgraph algorithm is comprised of the one or more pass from the template workgraph algorithm and one or more pass from the first incremental workgraph algorithm.

14. The method according to any one of claims 9 to 13, wherein the method further comprises: storing in a memory the template workgraph algorithm and / or the first incremental workgraph algorithm; and dynamically loading from the memory, the template workgraph algorithm and / or the first incremental workgraph algorithm for processing the pipeline.

15. The method according to claim 12, wherein one or more of: the template workgraph algorithm, the first incremental workgraph algorithm; the further workgraph algorithm; each node of the template workgraph algorithm; each node of the first incremental workgraph algorithm; or each node of the further workgraph algorithm, comprise a unique identifier.

Citation Information

Patent Citations

  • Task graph generation for workload processing

    US20210373892A1