Graphics processor
By tracking the execution status of the initial ‘boot’ processing job and controlling the task processing of the ‘main’ processing job, the strict processing barrier between the initial ‘boot’ processing job and the ‘main’ processing job in the graphics processor is removed, and the problem of inefficiency is solved, improving the utilization rate of the processing core and the security of the graphics processing operations.
Patent Information
- Application Number
- CN202411716875.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-30
AI Technical Summary
When the graphics processor executes a processing job sequence including the initial ‘boot’ processing job, there is a problem of inefficiency in the prior art, especially when there is a strict processing barrier between the initial ‘boot’ processing job and the subsequent ‘main’ processing job, resulting in a decrease in utilization of the processing core.
By tracking the execution status of the initial ‘boot’ processing job and controlling the task processing of the ‘main’ processing job if necessary, ensuring that the ‘main’ shader program only starts execution after the initial ‘boot’ processing job is completed, removing the strict processing barrier, allowing the processing core to run and execute dependent ‘main’ processing jobs before the current command stream is executed.
This method improves the operation efficiency of the graphics processor, reduces the waiting time during the initial ‘boot’ processing job execution, improves the utilization rate of the processing core, and ensures the security of the graphics processing operations by forcing data dependence in the hardware.
Smart Images

Figure CN120070148A_ABST
Abstract
Description
[0001] The present invention relates to performing data processing using a graphics processor, and more particularly to the operation of a graphics processor when executing a sequence of processing jobs that includes one or more initial "bootstrap" processing jobs (wherein, as will be further explained below, the initial "bootstrap" processing jobs execute corresponding "bootstrap" shader programs that are to be executed before the corresponding "main" shader program, which will execute for individual "main" processing jobs in the sequence of processing jobs).
[0002] Modern graphics processors typically include one or more processing (shader) cores that execute the programmable processing stages of the graphics processing pipeline implemented by the graphics processor, commonly referred to as "shaders".
[0003] Thus, the graphics processor processing (shader) cores are processing units that perform processing by running (usually small) programs for each "work item" in the output to be generated. In the case of generating a graphics output (such as a render target, such as a frame to be displayed), in this regard, a "work item" can be a sampling location (e.g., in the case of a fragment shader), but can also be a vertex or a ray, for example, depending on the graphics processing (shading) operation being discussed. In the case of a compute shading operation, each "work item" in the output being generated will be a data instance (item) in the work "space" where the compute shading operation is being performed.
[0004] The shader programs to be executed by a given "shader" of the graphics processing pipeline will typically be provided by an application that requires graphics processing using a high-level shader programming language (such as GLSL, HLSL, OpenCL, etc.). This shader program will typically consist of "expressions" that indicate the desired programming steps defined in the relevant language standard (specification). The high-level shader program is then translated by a shader language compiler into binary code for the target graphics processing pipeline. This binary code will consist of "instructions" specified in the instruction set specification for the given target graphics processing pipeline. The compilation process for converting shader language expressions into binary code instructions can proceed via multiple intermediate representations of the program within the compiler. Thus, a program written in the high-level shader language can be translated into a compiler-specific intermediate representation (and there can be several successive intermediate representations within the compiler), where the final intermediate representation is translated into binary code instructions for the target graphics processing pipeline.
[0005] Accordingly, unless the context otherwise requires, references herein to "expressions" refer to shader language constructs to be compiled into target graphics processor binary code (i.e., to be expressed in hardware microinstructions). Depending on the shader language under discussion, such shader language constructs may be referred to as "expressions", "statements", etc. (for convenience, the term "expression" will be used herein, but this is intended to cover all equivalent shader language constructs, such as "statements" in GLSL). "Instructions" correspondingly refer to the actual hardware instructions (code) that are issued to execute the "expressions".
[0006] To execute a shader program, the graphics processor will include one or more suitable execution units (or circuits) for that purpose. The execution units will include programmable processing circuits for executing the shader programs (the "shaders" of the graphics processing pipeline).
[0007] The actual data processing operations performed by the execution units when executing the shader program are typically performed by the corresponding functional units of the execution units, and the corresponding functional units may include (but are not limited to) texture mapping units configured to perform certain texture operations. Thus, the functional units will perform appropriate data processing operations in response to the instructions in the (shader) program being executed and according to its requirements.
[0008] In addition to the programmable execution units (using their associated functional units) that execute the shader programs, the graphics processor processing (shader) core typically may also include one or more essentially fixed-function (hardware) stages for implementing certain stages of the graphics processing (rendering) pipeline. These fixed-function stages may be used to handle certain "front-end" processing operations to be performed before the shader program. For example, in the case where a fragment shader is to be executed for a rendering job, the "front-end" processing operations may include, for example, operations such as primitive list reading, resource allocation, vertex extraction, rasterization, early depth / stencil testing, etc., which set the required state for executing the fragment shader program that performs the actual rendering operations to produce the required graphics processing output (render target) (such as a frame for display). Substantially fixed-function (hardware) stages may also be used to handle certain post-shader actions (such as late depth / stencil testing or tile writing). However, in this regard, various arrangements will be possible, for example, depending on the specific configuration of the graphics processing pipeline and the type of processing job being executed.
[0009] Typically, there may be many parallel processing (shader) cores within a graphics processor, enabling the graphics processor to simultaneously process multiple different tasks in parallel. Thus, in a tile-based rendering system, for example, where the rendering output is subdivided into multiple rendering tiles for processing, the corresponding tasks for rendering different tiles (tasks) can be issued to different processing (shader) cores, such that the tiles can be rendered in parallel. Accordingly, each graphics processor processing (shader) core is operable and configured to implement an instance of a graphics processing pipeline for processing a given rendering task assigned to it. Thus, this can provide more efficient graphics processor operation.
[0010] Accordingly, when the graphics processor receives a set of commands from a host processor (e.g., a CPU) that is executing an application that requires graphics processing to perform a particular processing "pass" (the processing pass may typically include a certain sequence of one or more processing "jobs", each processing "job" producing a particular output of the processing pass, and where each processing "job" may typically include a corresponding set of one or more tasks to be processed for that particular output associated with the processing job), then the commands are processed within a suitable command processing unit (e.g., a command stream front end / job manager) of the graphics processor to identify the processing jobs to be executed, and then a suitable task iterator schedules the processing of the corresponding tasks to be executed for the processing jobs of the processing pass, where these tasks are assigned to available processing (shader) cores for processing. When tasks are assigned to processing (shader) cores for processing, the processing (shader) cores are thus operable and configured to load the required data for setting up the shader program via a "front end" processing stage and then execute the required shader program to produce the required output. The output of the rendering task can then be written out accordingly (and written out accordingly).
[0011] Although the graphics processor operation has been described above with respect to a single processing pass, it will be appreciated that a graphics processor is typically operable to perform multiple processing passes, e.g., for producing a set of frames, e.g., for display. In some cases, different processing passes in a sequence of processing passes being executed may be independent of each other (e.g., where different processing passes are associated with different frames, or are specifically written to different (data buffers)). However, typically, at least some of the different processing passes in a sequence of processing passes being executed are related to each other. There are also cases where different processing jobs within a particular processing pass may be dependent on each other, such that in a given sequence of processing jobs to be executed as part of a particular processing pass, there may be certain data (processing) dependencies between the processing jobs, and these data (processing) dependencies, if not enforced, can result in certain artifacts in the final output.
[0012] A specific example of this scenario would be when the graphics processor is executing a processing pass in which a sequence of processing jobs is to be executed, including one or more initial "bootstrap" processing jobs, where the initial "bootstrap" processing jobs within the sequence of processing jobs execute corresponding "bootstrap" shader programs that are to be executed before a corresponding "main" shader program, and the corresponding "main" shader program is to be executed later (i.e., within the same processing pass) by a separate "main" processing job within the sequence of processing jobs.
[0013] For example, the presence of expressions in shader programs to be executed that operate on constant inputs at runtime (i.e., when the application is being executed) can be identified, and then those expressions are actually extracted and executed in a separate initial "bootstrap" shader program before the "main" shader program. The initial "bootstrap" shader program can thus execute certain constant program expressions of the original shader program in an initial "bootstrap" shader program for the original shader program before a separate subsequent "main" shader program. The subsequent "main" shader program can then use the results of the initial bootstrap shader program. In this case, the processing pass can thus include one or more initial "bootstrap" processing jobs followed by a set of subsequent "main" processing jobs that are dependent on the one or more initial "bootstrap" processing jobs (since the "main" shader program should not be executed by any processing (shader) core for any task of any "main" processing job until the corresponding initial "bootstrap" shader program that is to be executed before the "main" shader program has been executed, such that there is a strict processing dependency between the execution of the "main" shader program and the execution of the "bootstrap" shader program within the (same) processing pass).
[0014] For example, the use of such "bootstrap" shader programs is described in U.S. Patent No. 9,189,881, assigned to Arm Limited, the entire content of which is incorporated herein by reference.
[0015] However, the applicant believes that there is still scope for improvement in the operation of the graphics processor when executing a sequence of processing jobs that includes one or more initial "bootstrap" processing jobs (where the initial "bootstrap" processing jobs execute corresponding "bootstrap" shader programs that are to be executed before a corresponding "main" shader program, and the corresponding "main" shader program is to be executed for a separate "main" processing job within the sequence of processing jobs).
[0016] According to a first aspect of the present invention, there is provided a method of operating a graphics processor, the graphics processor including a set of one or more processing cores, the method comprising:
[0017] When performing a pass of processing that includes one or more initial "bootstrap" processing jobs, where the initial "bootstrap" processing jobs execute corresponding initial "bootstrap" shader programs that are to be executed before the corresponding "main" shader programs to be executed for separate "main" processing jobs within the same pass of processing, such that the "main" shader programs have a dependency on the initial "bootstrap" shader programs of the initial "bootstrap" processing jobs, and where the "main" processing jobs include corresponding sets of one or more tasks to be processed for the "main" processing jobs, each task being operable to execute a corresponding instance of the "main" shader program:
[0018] Track whether any initial "bootstrap" processing job is currently being processed by the set of one or more processing cores; and
[0019] When a corresponding task to be processed as part of a "main" processing job is posted to a corresponding processing core for processing, while the set of one or more processing cores is simultaneously performing processing for at least one initial "bootstrap" processing job:
[0020] Based on the tracking of whether any initial "bootstrap" processing job is currently being processed by the set of one or more processing cores, control the processing of the task of the "main" processing job within the processing core such that the "main" shader program of the "main" processing job is not executed with respect to the task, at least until any initial "bootstrap" processing job that is being processed simultaneously with the task and on which the "main" shader program has a dependency has completed its processing.
[0021] According to a second aspect of the present invention, there is provided an image processor comprising:
[0022] A set of one or more processing cores;
[0023] A task posting circuit that is operable and configured to post tasks to the set of one or more processing cores for processing; and
[0024] A control circuit,
[0025] Wherein the control circuit is configured to:
[0026] When the graphics processor is executing a processing pass that includes one or more initial "boot" processing jobs, where the initial "boot" processing jobs execute respective initial "boot" shader programs that are to be executed before corresponding "main" shader programs to be executed for separate "main" processing jobs within the same processing pass, such that the "main" shader programs have a dependency on the initial "boot" shader programs of the initial "boot" processing jobs, and where the "main" processing jobs include respective sets of one or more tasks to be processed for the "main" processing jobs, each task being operable to execute a respective instance of the "main" shader program:
[0027] Tracking whether any initial "boot" processing job is currently being processed by the set of one or more processing cores; and
[0028] Based on tracking whether any initial "boot" processing job is currently being processed by the set of one or more processing cores, controlling the processing of tasks of the "main" processing jobs within the processing cores such that when a respective task to be processed as part of a "main" processing job is issued to a respective processing core for processing, while the set of one or more processing cores is simultaneously executing the processing of at least one initial "boot" processing job:
[0029] The "main" shader program of the "main" processing job is not executed with respect to the task, at least until any initial "boot" processing job that is being processed simultaneously with the task and on which the "main" shader program has a dependency has completed its processing.
[0030] The present invention relates to graphics processor operation in which certain (e.g., constant) program expressions of a primitive shader program can be executed in an initial "boot" shader program for the primitive shader program before a separate subsequent "main" shader program. The subsequent "main" shader program can then use the results of the initial "boot" shader program. In this way, for example, instructions emitted for a constant program expression can be executed only once in the initial boot shader program, rather than having to be executed multiple times in the corresponding primitive shader program whenever the result of the expression in question is needed. Thus, this can remove repetitive computations from the shading process.
[0031] As will be further explained below, an initial "bootstrap" shader program is thus extracted from the original shader program and associated with a collection of one or more subsequent "main" shader programs that have also been extracted from the (same) original shader program. Thus, the original shader program (e.g., as provided by an application) is effectively split into an initial "bootstrap" shader program and a corresponding collection of one or more subsequent "main" shader programs, which one or more subsequent "main" shader programs will thus have a data (processing) dependency on the initial "bootstrap" shader program (or programs) associated with them.
[0032] This type of "initial" shader program (or "pre-shader" program) will be referred to herein as a "bootstrap" shader program (it should be understood, however, that the term "bootstrap" shader program is intended to cover any and all equivalent such arrangements in which some portion of the original shader program may be executed by the initial shader program, which executes separately and prior to the corresponding "main" shader program that uses the results of the initial shader program (such that the "main" shader program has a data (processing) dependency on the initial shader program).
[0033] The initial "bootstrap" shader program can thus (and typically) be executed as part of a corresponding initial "bootstrap" processing job that is used to execute the required initial "bootstrap" shader program and appropriately write out the results of the initial "bootstrap" shader program such that the results of the initial "bootstrap" shader program execution can then be used, when needed, by the corresponding "main" shader program that will be executed as part of a subsequent "main" processing job (but the initial "bootstrap" processing job is preferably only used to produce such intermediate results, i.e., by executing the initial "bootstrap" shader program, and not, for example, to produce part of a final rendered output (e.g., for display), which will instead be produced by one or more subsequent "main" processing jobs).
[0034] Accordingly, the present invention generally relates to the operation of a graphics processor when performing a pass including a sequence of processing operations, the sequence of processing operations including one or more such initial "bootstrap" processing operations, wherein, as explained above, the initial "bootstrap" processing operations execute corresponding "bootstrap" shader programs, the corresponding "bootstrap" shader programs to be executed before corresponding "main" shader programs to be executed for individual "main" processing operations in the (same) sequence of processing operations (pass). In this case, the execution of the "main" shader programs associated with the "main" processing operations is dependent on the results of the execution of the corresponding initial "bootstrap" shader programs for one or more initial processing operations (such that the "main" shader programs should not be executed until their initial "bootstrap" shader programs have been executed), and the "main" processing operations thus have a data (processing) dependency on the one or more initial "bootstrap" processing operations.
[0035] Typically (and in the preferred embodiment), the same graphics processor command stream is responsible for initiating the initial "bootstrap" processing operations and the corresponding, dependent "main" processing operations.
[0036] In some more traditional graphics processor arrangements, a strict (hard) processing barrier is enforced between processing operations within the graphics processor command stream such that the graphics processor explicitly waits until an earlier processing operation has completed its processing before releasing a later processing operation for processing. In the above-described scenario where the command stream includes one or more initial "bootstrap" processing operations, this processing barrier thus ensures that any initial "bootstrap" shader program will be executed before the corresponding dependent "main" shader program, since the processing operation including the "main" shader program will not be released from the command stream for processing until the one or more initial "bootstrap" processing operations have completed their processing.
[0037] However, the present applicant has recognized that the execution of such initial "bootstrap" shader programs can often take thousands of cycles during which the set of processing cores of the graphics processor is typically utilized relatively low. For example, modern graphics processors typically include a relatively large number of processing cores, and the present applicant has found that in more traditional arrangements where a strict (hard) processing barrier is enforced between different processing operations, since the processing cores cannot start executing any processing for the next processing operation until the current processing operation has ended, it is often the case that most processing cores are idle for many processing cycles while waiting for the execution of the "bootstrap" shader program on a single processing core to end.
[0038] Accordingly, the present applicant has recognized that it would be beneficial to remove such a strict (hard) processing barrier between an initial "bootstrap" processing operation and a subsequent dependent "main" processing operation, i.e., in order to allow the processing core to run and start executing at least some of the processing of the dependent "main" processing operation concurrently with the processing of the corresponding initial "bootstrap" processing operation on which the dependent "main" processing operation depends, while still ensuring safe (artifact-free) graphics processing operations by ensuring that any dependencies between the "bootstrap" shader program execution and the "main" shader program execution within a particular processing pass are properly enforced within the processing core. The present invention provides a particularly efficient mechanism for doing so, as will be further explained below.
[0039] Accordingly, in accordance with the present invention, the method includes a (task issuing circuit (task iterator)) issuing tasks for a "main" processing operation to corresponding processing (shader) cores without waiting for any initial "bootstrap" processing operation currently being processed by a set of one or more processing (shader) cores to finish its processing.
[0040] This can then help reduce the latency associated with the execution of any initial "bootstrap" shader program, since in accordance with the present invention, "main" processing operations that may require the results of an initial "bootstrap" shader program (such that the "main" processing operations have a dependency on the corresponding initial "bootstrap" processing operation) can be issued for processing relatively early, e.g., and in particular, such that at least some of the processing of the "main" processing operation can be executed concurrently with (and with) the execution of the initial "bootstrap" shader program on which the "main" processing operation depends, provided that when it is necessary to enforce the dependency between the "main" shader program and its corresponding initial "bootstrap" shader program, the execution of the "main" shader program for any tasks that are part of the processing being done as the "main" processing operation (for which the "main" shader program will execute) can be appropriately controlled (e.g., stopped) within the processing (shader) core.
[0041] In this way, at least some tasks to be processed for a particular "main" processing operation can effectively begin to be "pre-loaded" into the processing (shader) cores for processing, while a set of processing (shader) cores is currently performing processing for an initial "bootstrap" processing operation on which the "main" processing operation depends. For example, and in particular, such that at least some "front-end" processing of those tasks can be performed concurrently with the execution of a "bootstrap" shader program on which the "main" shader program associated with the "main" processing operation depends, i.e., processing (or any other processing that may depend on the execution of the "bootstrap" shader program) up to and excluding the execution of the "main" shader program associated with the "main" processing operation. This then means that when the initial "bootstrap" processing operation for a given "main" processing operation has been completed (i.e., one or more initial "bootstrap" shader programs have been fully executed), the subsequent (next) "main" processing operation in the sequence of processing operations being performed is already being processed within the processing (shader) cores, such that the "main" shader program can execute relatively faster (e.g., and preferably, immediately) with respect to any tasks for the "main" processing operation that have been pre-loaded into the processing (shader) cores in this way after the corresponding "bootstrap" shader program that is to be executed before the "main" shader program has ended.
[0042] Accordingly, by pre-loading the "main" processing operation into the processing cores in this way, where a dependency on the initial "bootstrap" processing operation is enforced within the processing cores, the present invention can reduce the latency associated with performing such initial "bootstrap" processing operations, and thus also improve processing (shader) core utilization. Removing a strict (hard) processing barrier in the command stream between the initial "bootstrap" processing operation and the subsequent dependent "main" processing operation also allows the task issue circuitry (task iterator) to run ahead (whereas in more traditional graphics processors that enforce such strict (hard) processing barriers in the command stream, the task issue circuitry (task iterator) may stop to wait for the command processing unit (command stream front end / job manager (circuitry)) to create further jobs for it to process).
[0043] In this regard, the applicant has recognized that as long as the "main" shader program is not executed for any task to be processed for the "main" processing job before any corresponding initial "bootstrap" shader program on which the executed "main" shader program depends, then it is generally safe to start issuing tasks for the "main" processing job to a set of processing (shader) cores for processing and to perform at least some front-end processing of those tasks for the "main" processing job within the processing cores simultaneously with the corresponding initial "bootstrap" processing job (or multiple corresponding initial "bootstrap" processing jobs) of the "main" processing job (i.e., processing up to but not including the execution of the "main" shader program or any processing that may depend on the results of the initial "bootstrap" shader program).
[0044] Accordingly, the present invention allows at least some processing of the "main" processing job to overlap with the processing of the corresponding initial "bootstrap" processing job on which the "main" processing job depends across a set of processing (shader) cores of a graphics processor, particularly such that at least some front-end processing of the tasks for the "main" processing job can be performed simultaneously with the processing of the corresponding initial "bootstrap" processing job (i.e., with the execution of the "bootstrap" shader program), wherein the data (processing) dependencies between the "main" shader program and the initial "bootstrap" shader program associated with the processing job under discussion are enforced within the processing (shader) cores (in hardware) (e.g., rather than being enforced within the graphics processor command stream as in the more traditional methods mentioned above).
[0045] Thus, by allowing the task issuing circuitry (task iterator) to issue tasks for the "main" processing job to the corresponding processing (shader) cores for processing without having to wait for any initial "bootstrap" processing job, for which there may be data (processing) dependencies on it, to finish its processing, it can be (and at least sometimes will be) the case that tasks for the "main" processing job are issued to the corresponding processing (shader) cores for processing while a set of one or more processing (shader) cores is currently (still) executing the processing for the initial "bootstrap" processing job, at least the "main" shader program for the "main" processing job has a processing (data) dependency on the initial "bootstrap" processing job (in more traditional arrangements, this situation is avoided due to the strict (hard) processing barriers between processing jobs). In the case where tasks for the "main" processing job are issued to the corresponding processing (shader) cores for processing while a set of one or more processing (shader) cores is currently (still) executing the processing for the initial "bootstrap" processing job (for which there is a processing (data) dependency), according to the present invention, additional control is performed on the processing of the task such that the "main" shader program for the "main" processing job is not executed on the task, at least until any initial "bootstrap" processing job that is being processed concurrently with the task and on which the "main" shader program has a dependency has finished its processing. This then allows any required data (processing) dependencies between the "main" shader program and the initial "bootstrap" shader program within a particular processing pass to be enforced within the processing (shader) cores as needed.
[0046] In particular, to enforce the data (processing) dependencies between the "main" shader program and the initial "bootstrap" shader program within a particular processing pass, the present invention (when executing a processing pass that includes one or more initial "bootstrap" processing jobs) tracks whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing cores. Thus, preferably, a suitable record is maintained as to whether the set of the processing (shader) cores is currently executing any initial "bootstrap" processing job.
[0047] Tracking whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing (shader) cores can generally be performed in any suitable and desired manner (using any suitable record or data structure) according to the specific requirements of the present invention. In this regard, various arrangements will be possible.
[0048] For example, and preferably, tracking is performed by maintaining a reference counter that stores a count indicating how many initial "bootstrap" processing jobs, if any, are currently being executed. Thus, in a preferred embodiment, when a task to be processed as part of an initial "bootstrap" processing job is posted to a set of processing cores for processing, the reference counter preferably increments accordingly. Correspondingly, when a task for an initial "bootstrap" processing job ends, the reference counter preferably then decrements. The reference counter thus preferably tracks whether any task for any initial "bootstrap" processing job is currently being executed by the set of processing cores of (any one) graphics processor.
[0049] To facilitate such tracking, the initial "bootstrap" processing job is preferably marked in the graphics processor command stream as being separate from regular (e.g., "main") processing jobs. For example, as will be explained below, the initial "bootstrap" processing job is preferably created by a shader compiler that executes as part of a software driver for the graphics processor, and the software driver prepares the graphics processor command stream. Thus, when such a command stream is prepared and submitted to the graphics processor for execution, the driver can and preferably does mark the initial "bootstrap" processing job accordingly within the command stream.
[0050] For example, the command stream received by the graphics processor is preferably processed by an appropriate command processing unit of the graphics processor (e.g., a command stream front end / job manager (circuit)) that is operable and configured to identify the presence of such initial "bootstrap" processing jobs, and such identification is preferably done based on such marking of the initial "bootstrap" processing job within the graphics processor command stream. This then provides a very simple and effective mechanism for identifying any such initial "bootstrap" processing jobs within the graphics processor command stream (and thus triggering the specific operations of the present invention). However, various other arrangements will be possible, and in general, the identification of the initial "bootstrap" processing jobs can be more or less complex as needed.
[0051] In a preferred embodiment, the task posting circuit (task iterator) of the graphics processor is then responsible for dividing the processing job into corresponding tasks and controlling the scheduling and assignment of the corresponding tasks to the graphics processor processing (shader) cores for processing. In this regard, it should be understood that a particular processing job can and in a preferred embodiment does typically include multiple individual processing tasks that can then be assigned to the processing cores for processing. For example, in a typical graphics processor, there will be multiple processing cores. Thus, it is possible and preferably common to perform the assignment of tasks to the processing cores to attempt to balance the processing among the available processing (shader) cores of the graphics processor and ensure a higher utilization of the available processing (shader) cores.
[0052] (Each task can in turn be broken down into smaller units of work for processing by the processing (shader) cores. For example, each task can spawn multiple threads, each of which executes a separate instance of the shader program.)
[0053] Accordingly, the task issue circuitry (task iterator) of the graphics processor is preferably operable and configured to break processing jobs into their respective tasks and then issue those tasks to the respective processing (shader) cores of the graphics processor for processing (and such task scheduling is preferably done in a normal manner for scheduling tasks onto the processing (shader) cores of the graphics processor). These tasks can then be executed separately (e.g., in parallel) across different processing (shader) cores of the graphics processor. For example, in the context of a rendering job, the rendering output can be subdivided into a plurality of render tiles (or "metatiles"), where different tiles ("metatiles") are assigned to different processing (shader) cores for processing. In this regard, various arrangements will be possible.
[0054] Accordingly, any reference herein to issuing a processing job to a processing (shader) core (or generally, to a processing (shader) core) for processing should be understood accordingly as preferably referring to issuing the respective task from among a plurality of tasks to be processed for the processing job in question to the processing (shader) core for processing. In this regard, it will be appreciated that different tasks within a processing can generally and typically be assigned to different processing (shader) cores. Thus, a processing job can generally be executed across a plurality of different processing (shader) cores.
[0055] In a preferred embodiment, the task issue circuitry (task iterator) of the graphics processor performs tracking of whether any initial "boot" processing job is currently being processed by a set of one or more processing (shader) cores. Preferably, this is done as part of the normal task scheduling operation.
[0056] For example, in a preferred embodiment where tracing is performed using a reference counter, as described above, the reference counter is preferably maintained by a task issuing circuit (task iterator) that controls the allocation and scheduling of tasks for different processing jobs to processing cores. Thus, in the preferred embodiment, when a task for an initial "bootstrap" processing job is issued to a processing (shader) core for processing, the task issuing circuit (task iterator) is operable and configured to increment the reference counter accordingly. Thus, the task issuing circuit (task iterator) can easily keep track of which (if any) processing (shader) cores are currently executing the initial "bootstrap" processing job. Correspondingly, when a processing (shader) core finishes the task for the initial "bootstrap" processing job, this is preferably signaled to the task issuing circuit (task iterator) (e.g., as part of normal job completion signaling to indicate that all tasks associated with the job have ended), so that the reference counter is decremented.
[0057] Thus, the reference counter preferably keeps track of how many tasks associated with the initial "bootstrap" processing job are currently being executed. In this regard, it will be understood that the initial "bootstrap" processing job can and typically will only contain a single task. Thus, in practice, there may be no difference between tracing at the work level or the task level, and both methods can work effectively. However, in principle, the initial "bootstrap" processing job can contain any number of tasks, and in such cases, the tracing can be performed with respect to individual tasks associated with the initial "bootstrap" processing job or for the initial "bootstrap" processing job as a whole, and in this regard, various arrangements will be possible. For example, in the case where the initial "bootstrap" processing job can contain multiple tasks, the task issuing circuit (task iterator) can increment the reference counter only when the first task for the initial "bootstrap" processing job is issued, and decrement the reference counter only when all tasks associated with the initial "bootstrap" processing job have ended, and in this way, the reference counter can be used to keep track of whether any initial "bootstrap" processing job is currently being executed at the job level (if desired).
[0058] Thus, tracking whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing (shader) cores is preferably performed globally across all processing (shader) cores within the set of one or more processing cores. That is, the tracking preferably tracks whether any processing (shader) core among the processing (shader) cores is currently executing any task of the initial "bootstrap" processing job (in which case, no processing (shader) core should execute the "main" shader program associated with the corresponding subsequent "main" processing job until the initial "bootstrap" processing job has ended). For example, as described above, it is often the case that the initial "bootstrap" processing job is being executed by only a single processing (shader) core, but this means that all other processing (shader) cores must wait for the initial "bootstrap" processing job to end before executing any dependent "main" shader program (and thus the present invention provides a mechanism for handling such a situation).
[0059] Tracking whether the set of processing (shader) cores is currently executing any initial "bootstrap" processing job can then be (and is) used accordingly to control the processing of (individual) tasks to be processed as part of a subsequent "main" processing job, e.g., and preferably, to enforce (when needed) any data (processing) dependencies between the "main" shader program to be executed as part of the "main" processing job and the initial "bootstrap" shader program to be executed before the "main" shader program as part of a separate initial "bootstrap" processing job.
[0060] Thus, when tasks to be processed for a particular "main" processing job are issued to a particular processing (shader) core for processing, it can be (and preferably is) determined, based on such tracking (e.g., and preferably, by checking a reference counter when providing a reference counter), whether any initial "bootstrap" processing job is currently being executed by any processing (shader) core among the processing (shader) cores within the set of processing (shader) cores of the graphics processor, on which the "main" processing job potentially depends.
[0061] If it is determined that no initial "bootstrapping" processing job is currently being executed (i.e., there is no task associated with the currently ongoing initial "bootstrapping" processing job), then no further control or checking regarding that task is required. On the other hand, when it is determined that at least one task associated with the initial "bootstrapping" processing job is currently being executed by a set (one or more processing cores among them) of the processing cores of the graphics processor such that at least some of the tasks to be processed regarding a particular "main" processing job will be issued for processing simultaneously with the initial "bootstrapping" processing job on which the "main" processing job potentially depends, then appropriate control can then (and) be performed for the tasks for the subsequent "main" processing job, such as and preferably to stop the execution of the "main" shader program until it can be determined that the initial "bootstrapping" processing job has completed its processing.
[0062] Thus, when the corresponding tasks to be processed as part of the "main" processing job are issued to the corresponding processing cores for processing, while the set of one or more processing cores is simultaneously performing the processing for at least one initial "bootstrapping" processing job within the same processing pass as the "main" processing job targeted by the task being executed, the present invention includes controlling the processing of the tasks for the "main" processing job within the processing cores based on tracking whether any initial "bootstrapping" processing job is currently being processed by the set of one or more processing cores, such that the "main" shader program for the "main" processing job is not executed regarding that task, at least until any initial "bootstrapping" processing job for the current processing pass has ended.
[0063] The control to enforce data (processing) dependencies is thus performed within the processing (shader) cores by controlling the processing within individual tasks (i.e., the entities issued to the corresponding processing (shader) cores for processing), while the tracking is preferably performed at the level of the processing job as a whole.
[0064] In this regard, it should be understood that control will thus and typically be required to be performed across different processing (shader) cores, since different tasks to be processed for a particular "main" processing job can and typically will be assigned to different processing (shader) cores (and these will typically be different processing (shader) cores from the processing (shader) cores executing any initial "bootstrapping" processing job).
[0065] According to the specific requirements of the present invention, such control can be performed in various suitable ways as needed.
[0066] For example, in some embodiments, this can be accomplished by the processing (shader) core explicitly checking, prior to the execution of any "main" shader program, whether any initial "bootstrap" processing job is currently being executed by any of the processing (shader) cores while processing a particular task of the "main" processing job. In this case, the processing (shader) core may need to send a message to the task issue circuit (task iterator) at this moment to perform a dependency check. Then the execution of the "main" shader program can be gated, waiting for an appropriate response from the task issue circuit (task iterator). For example, once the task issue circuit (task iterator) confirms that no initial "bootstrap" processing job is currently being executed, this can be signaled back to the processing (shader) core that requested the dependency check, and that processing (shader) core can then execute the "main" shader program for the processing job it is currently executing.
[0067] However, this method may require increased signaling between the processing (shader) core and the task issue circuit (task iterator), as separate dependency checks may need to be performed for each task of each processing job.
[0068] Thus, as another example, and in a preferred embodiment, the task issue circuit (task iterator) is operable and configured to check whether any initial "bootstrap" processing job is currently being executed by any of the processing (shader) cores when issuing the tasks of a "main" processing job to the processing (shader) cores for processing. That is, in a preferred embodiment, when tasks are assigned to the processing (shader) cores for processing, the task issue circuit (task iterator) performs a check as to whether any initial "bootstrap" processing job is currently being executed by any of the processing (shader) cores. In this case, each task associated with the "main" processing job being assigned to the processing (shader) cores for processing can thus (and in this embodiment) be associated with an appropriate indicator (such as a flag) indicating whether any initial "bootstrap" processing job is currently being executed by any of the processing (shader) cores at the moment the task is issued for processing.
[0069] Thus, in a preferred embodiment, if no initial "bootstrap" processing job is being executed when a particular task of a subsequent "main" processing job is issued for processing, this is correspondingly indicated for the task in question, and in this case no additional control (such as stopping) needs to be performed. In this case, the task can be processed, for example, in the normal manner, where the "main" shader program to be executed as part of the "main" processing job involved in the task is executed immediately after any front-end processing of the task in question has ended.
[0070] In this regard, it should be understood that processing jobs should and preferably are issued for processing in a strict order such that any initial "bootstrap" processing job is issued for processing before any corresponding "main" processing job. Thus, if at the time a particular task of a particular "main" processing job is issued for processing, no initial "bootstrap" processing job is currently being executed by any processing (shader) core in the processing (shader) core, then this means that any initial "bootstrap" processing job for that particular "main" processing job must have ended, such that the results of any initial "bootstrap" shader program executed as part of those initial "bootstrap" processing jobs must therefore be available for use by the "main" shader program associated with that particular "main" processing job.
[0071] It should also be understood that "bootstrap" processing jobs can generally end relatively quickly. Thus, performing a check at the moment an individual task within a "main" processing job is issued for processing can be particularly effective because for most tasks of most processing jobs, there will be no "bootstrap" processing job currently in progress at the moment the task is issued for processing (and thus no further checking or control will be required once this has been determined).
[0072] On the other hand, in the case where the task issuing circuit (task iterator) determines that at least one initial "bootstrap" processing job is currently being executed by one (or more) processing (shader) cores in the processing (shader) core at the moment a particular task of a particular "main" processing job is issued for processing, this situation is accordingly indicated when the task is issued to the processing (shader) core for processing, and additional control can then (and then) be performed for the task in question such that any associated "main" shader program to be executed for the task that is part of the "main" processing job is not executed until it can be determined that it is safe to do so, i.e., until any initial "bootstrap" processing job has ended its processing (which, as will be further explained below, can be appropriately signaled by the processing (shader) core, e.g., as part of normal job status signaling).
[0073] Thus, when a task is issued for processing, if any initial "bootstrap" processing job is currently being executed by the set of processing cores, then preferably a suitable indication is set to ensure that any thread for any task of the "main" processing job does not execute a dependent "main" shader program to be executed as part of a subsequent "main" processing job until the initial "bootstrap" processing job has ended (and this control is preferably executed across all processing (shader) cores because the task in question being controlled can be assigned to a processing (shader) core different from the processing (shader) core executing the initial "bootstrap" processing job).
[0074] Thus, in a preferred embodiment, if at the moment a particular task of a particular "main" processing job is issued for processing, there is at least one initial "bootstrap" processing job currently being executed by one (or more) of the processing (shader) cores, then the task is issued to the processing (shader) core for processing, and the processing (shader) core can then (and preferably does) begin to execute some processing for preloading a certain state, etc., as long as the processing is not dependent on the result of the initial "bootstrap" shader program. However, any processing that is dependent on the result of the initial "bootstrap" shader program (such as, the loading of any "main" shader program state created by the initial "bootstrap" shader program, and the execution of the "main" shader program) is stopped until it can be determined that any initial "bootstrap" processing job that is currently being executed (and thus the execution of the "main" shader program can be dependent on) has finished its processing.
[0075] In an embodiment, therefore, the execution of the "main" shader program that is part of the "main" processing job is appropriately gated (controlled) based on determining that there is at least one initial "bootstrap" processing job currently being executed by one (or more) of the processing (shader) cores. Thus, any task associated with the "main" processing job can be (only) processed until the moment of the execution of the "main" shader program, but preferably then the execution of the "main" shader program is gated based on determining that there is at least one initial "bootstrap" processing job currently being executed by one (or more) of the processing (shader) cores, such that further processing of the task is effectively stopped.
[0076] For example, a suitable barrier operable to gate the execution of the "main" shader program may be included within the graphics processing pipeline before the "main" shader program execution. This barrier may be included within the graphics processing pipeline at any suitable and desired moment before the "main" shader program execution. When the task of the "main" processing job is issued for processing while at least one initial "bootstrap" processing job is currently being executed, this may then be indicated accordingly, for example by setting an appropriate flag regarding the task. Thus this indication (flag) causes the processing of the task to stop at the barrier when set. On the other hand, if the indication (flag) is not set (i.e., or cleared), then the task may (and preferably does) ignore the barrier, and the further processing of the task including the "main" shader program execution continues accordingly. Thus, as described above, when at least one initial "bootstrap" processing job ends, the task issuing circuit (task iterator) may then signal to the processing (shader) core that no initial "bootstrap" processing job is currently being executed, and this may cause the indication (flag) to be appropriately cleared so that the task may continue past the barrier. In this regard, various arrangements will be possible.
[0077] If the execution of the "main" shader program of the task is gated in this manner, then when it is determined that one or more initial "bootstrap" processing jobs that caused the execution of the "main" shader program to be gated have ended their processing, the gating of the "main" shader program execution may and should be lifted accordingly. This may be done in various suitable ways as needed, but in the preferred embodiment, this is done using the reference counter described above. That is, when the (relevant) reference counter decrements to zero, the task issuing circuit (task iterator) preferably signals to each processing (shader) core that is executing a dependent processing job that no initial "bootstrap" processing job is being executed any longer, and this signaling then triggers the processing (shader) core to lift the gating of the "main" shader program execution and allow the processing of the task to end.
[0078] In this regard, it should be noted that the task issuing circuit (task iterator) is generally able to determine which processing (shader) cores are executing "main" processing jobs that are dependent on the initial "bootstrap" processing jobs, and thus this signaling may and preferably selectively (only) be done to those processing (shader) cores. In other embodiments, the signaling may be broadcast to all processing (shader) cores. In this regard, various arrangements will be possible.
[0079] In the described embodiments, the graphics processor (task issue circuit (task iterator)) is thus operable and configured to track whether any processing (shader) core is currently executing an initial "bootstrap" processing job, and use this tracking to appropriately gate the execution of any "main" shader program that may be dependent on the initial "bootstrap" processing job, such that any tasks associated with the "main" processing job are stopped prior to the execution of the "main" shader program. The graphics processor (task issue circuit (task iterator)) can then signal to the processing (shader) core to ungate the execution of the "main" shader program when it is safe to execute any "main" shader program. In this way, the "main" processing job can be preloaded into the processing (shader) core, as described above, while still allowing dependencies on the initial "bootstrap" processing job to be enforced across different processing (shader) cores. In this regard, it will be understood that in typical cases, the "main" processing job will be assigned to a different processing (shader) core than the processing (shader) core that is currently finishing the execution of the initial "bootstrap" processing job. However, the embodiments provide an effective mechanism for handling these dependencies.
[0080] Note that in this regard, the tracking in the preferred embodiment does not specifically track whether the "main" processing job whose tasks are currently being issued for processing actually requires the results of any initial "bootstrap" processing job that is currently in progress, and preferably enforces control (stops) whenever there is any initial "bootstrap" processing job in progress. Thus, in the preferred embodiment, it is assumed that the execution of the "main" shader program associated with a subsequent "main" processing job should always be stopped before any initial "bootstrap" processing job that is currently in progress is completed (regardless of whether there is an actual processing (data) dependency). This assumption is generally acceptable (and correct), since in typical cases, because processing jobs are issued for processing in sequence, where the workload within a typical processing pass involves a set of one or more initial "bootstrap" processing jobs, followed by a set of one or more "main" processing jobs that use the results of the (one or more) previous initial "bootstrap" processing jobs (i.e., as part of the same processing pass), it will naturally be the case that when an initial "bootstrap" processing job is currently in progress at the moment when a "main" processing job is being issued for processing, the "main" processing job will be dependent on the initial "bootstrap" processing job that is currently in progress.
[0081] However, at some point, when executing a "main" processing job, the graphics processor may encounter another initial "bootstrap" processing job within the graphics processor command stream (which is to be executed before another set of subsequent "main" processing jobs, i.e., for a subsequent processing pass).
[0082] In this regard, it should be understood that a processing pass is any suitably defined sequence of processing operations. In the context of the present invention, a given processing pass may thus typically include zero or more initial "bootstrap" processing operations, followed by one or more subsequent "main" processing operations, where the one or more subsequent "main" processing operations are dependent on corresponding initial "bootstrap" processing operations within the same processing pass. Typically, a graphics processor will not execute a single processing pass independently, but will instead be operated to execute a sequence of processing passes. Thus, once a particular (first) processing pass has ended, the graphics processor may begin to execute the next (second) processing pass, which may include its own set of initial "bootstrap" processing operations and dependent "main" processing operations. The second processing pass may or may not be dependent on the first processing pass.
[0083] If the second processing pass is dependent on the first processing pass, then it may be necessary to enforce a relatively stricter (harder) processing barrier between the different processing passes. For example, if the initial "bootstrap" processing operation of the second (later) processing pass is dependent on the result of the first (earlier) processing pass, then the initial "bootstrap" processing operation of the second (later) processing pass should not be executed until the first (earlier) processing pass has ended, and thus a stricter (harder) processing barrier can and preferably should be enforced before the initial "bootstrap" processing operation of the second (later) processing pass is executed.
[0084] On the other hand, as long as the second (later) processing pass is not dependent on the first (earlier) processing pass, it is generally safe to execute the initial "bootstrap" processing task of the second (later) processing pass concurrently with any "main" processing task that ends the first (earlier) processing pass (since any "main" processing task of the first (earlier) processing pass will generally only be dependent on the initial "bootstrap" processing tasks within that first (earlier) processing pass and will generally not be dependent on the results of any later-occurring processing passes). However, using the above simple tracking mechanism, for example, in the case of using a reference counter to count any initial "bootstrap" processing operations currently in progress, if a later-occurring initial "bootstrap" processing operation is issued for processing concurrently with an earlier-occurring "main" processing operation, then even if there is no possible dependency, the later-occurring initial "bootstrap" processing operation of the later processing pass may cause the "main" processing operation of the earlier processing pass to stop. The initial "bootstrap" processing operation of the later processing pass may also be dependent on the result of the earlier processing pass. This may potentially lead to a possible deadlock situation.
[0085] Thus, in a preferred embodiment, again, preferably a stricter (harder) processing barrier is enforced between any "main" processing jobs of the first (earlier) processing pass and any initial "bootstrap" processing jobs of the second (later) processing pass. To facilitate this, preferably an additional mechanism is provided that marks "main" processing jobs that are dependent on an earlier initial "bootstrap" processing job as dependent. If there is at least one dependent "main" processing job being executed, no new initial "bootstrap" processing job (i.e., for a subsequent processing pass) is allowed to be started. Thus, such a mechanism prevents subsequent initial "bootstrap" processing jobs from causing deadlocks. Additionally, enforcing such a barrier before subsequent initial "bootstrap" processing jobs generally has little impact on performance.
[0086] Thus, in a preferred embodiment, a stricter (harder) barrier is still enforced within the command stream between "main" processing jobs and any initial "bootstrap" processing jobs that occur later in the sequence of processing jobs. However, in this regard, various arrangements will be possible.
[0087] For example, in a preferred embodiment, as mentioned above, an execution trace is performed to simply track whether any initial "bootstrap" shader program is currently being executed, without attempting to explicitly track which processing pass the initial "bootstrap" shader program pertains to (and in this case, a harder (stricter) barrier can be enforced before any further initial "bootstrap" shader program (i.e., for the next processing pass) to ensure that any initial "bootstrap" shader program implicitly pertains to the current processing pass).
[0088] However, the trace can also be performed for individual processing passes. That is, the trace can explicitly track whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing cores for the current processing pass (where potentially separate traces are performed for different processing passes, e.g., using appropriate processing pass (e.g., age) identifiers indicating which processing pass a particular set of initial "bootstrap" processing jobs and subsequent "main" processing jobs pertain to). Thus, in an embodiment, tracking whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing cores is performed on a per-processing-pass basis, such that for an individual processing pass, it is tracked whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing cores for that (specific) processing pass.
[0089] In this case, it is possible to track accordingly which initial "bootstrap" processing job a particular "main" processing job depends on, and control the processing of the "main" processing job based only on the tracking of the particular initial "bootstrap" processing job (or particular initial "bootstrap" processing jobs) on which the "main" processing job depends. In this case, for example, as long as a later "bootstrap" processing job itself does not depend on an earlier processing pass (in which case a strict (hard) processing barrier should preferably be enforced), the initial "bootstrap" processing job of a second, later processing pass can be executed, for example, concurrently with the processing of the "main" processing job of the first, earlier processing pass. Thus, in an embodiment, the processing of the task that controls the "main" processing job can include controlling the processing of the task based on tracking whether any initial "bootstrap" processing job is currently being processed by a set of one or more processing cores for the same processing pass that includes the "main" processing job (however, the initial "bootstrap" processing jobs and the "main" processing jobs of different processing passes can be executed concurrently).
[0090] An initial "bootstrap" processing job can also depend on a previous initial "bootstrap" processing job within the same or a previous processing pass. In this case, it may be necessary to enforce a more strict (harder) processing barrier between different initial "bootstrap" processing jobs in the command stream. Thus, in a preferred embodiment, a more strict (harder) barrier is enforced before any new initial "bootstrap" processing job is issued for processing (such that preferably there is a barrier between the "main" processing job of an earlier processing pass and the initial "bootstrap" processing job of a later processing pass, and preferably also between different initial "bootstrap" processing jobs within the same processing pass, since these different initial "bootstrap" processing jobs can also, in principle, depend on each other) (in which case the number of in-progress initial "bootstrap" processing jobs can only be zero or one). Again, this does not introduce significant latency, since the initial "bootstrap" processing jobs typically end relatively quickly (compared to subsequent "main" processing jobs).
[0091] However, various arrangements will be possible in this regard, and different initial "bootstrap" processing jobs (for different processing passes) can also be executed concurrently, for example, as long as an appropriate mechanism is provided to enforce any potential processing (data) dependencies that may exist.
[0092] Accordingly, the present invention avoids a strict (hard) processing barrier between an initial "bootstrap" processing job and a dependent "main" processing job (although as described above, a more strict (harder) processing barrier is preferably still enforced before any new initial "bootstrap" processing job is issued). The graphics processor then tracks whether any initial "bootstrap" processing job is currently being executed by a set of one or more of its processing (shader) cores, and controls the processing of the dependent "main" processing job based on this tracking to enforce any data (processing) dependencies on the initial "bootstrap" processing job as needed.
[0093] Accordingly, the present invention can generally include issuing a "main" processing job to a set of one or more processing cores for processing concurrently with at least one initial "bootstrap" processing job; and controlling the processing of the "main" processing job within the set of one or more processing cores such that the "main" shader program for the "main" processing job is not executed until all of its corresponding initial "bootstrap" shader programs have been executed.
[0094] Accordingly, the effect and benefit of all of this is to allow at least some of the processing of subsequent "main" processing jobs to be executed concurrently with their corresponding initial "bootstrap" processing jobs, thus moving the processing barrier between the "main" processing job and the initial "bootstrap" processing job from the command stream into the graphics processor processing (shader) cores, where any data (processing) dependencies between the "main" processing job and the initial "bootstrap" processing job are then managed by the graphics processor (hardware) to ensure safe (artifact-free) graphics processor operation. As explained above, this can then increase processing (shader) core utilization and also reduce the latency associated with such initial "bootstrap" processing jobs.
[0095] In this regard, additional effects and benefits of the present invention are that, in many cases, since the initial "bootstrap" processing operation typically ends relatively quickly, any possible data (processing) dependencies between the initial "bootstrap" processing operation and subsequent "main" processing operations will have been naturally resolved by the time the "main" shader program is to be executed. That is, in many cases, when the front-end processing of the tasks of the "main" processing operation that is executed in parallel with the execution of the "bootstrap" shader program that is part of the initial "bootstrap" processing operation has been completed, any "bootstrap" shader program that is executed concurrently with this front-end processing will also have completed its execution, such that the "main" shader program can then be executed immediately. That is, in many cases, it will not be necessary to stop the execution of the "main" shader program because, by the time it is ready to be executed, the initial "bootstrap" shader program may and typically will have been completed. Thus, being able to pre-load the "main" processing operation can provide a significant performance improvement because, in many typical cases, the "main" shader program can be executed immediately with respect to the tasks to be processed as part of the "main" processing operation, without having to stop (control) the processing of the "main" shader program, as would typically be the case where there is no significant dependency on the "bootstrap" shader program by the time the front-end processing of the task has been completed.
[0096] (Similarly, it is possible that by performing front-end processing to pick certain tasks without having to execute the "main" shader program at all and by allowing this front-end processing to be executed concurrently with the initial "bootstrap" shader program, this can be determined relatively early on).
[0097] Accordingly, the present invention provides a mechanism to ensure that the execution of the "main" shader program can be stopped when it is necessary to do so to allow the corresponding initial "bootstrap" shader program to finish executing, thereby ensuring safe (correct) graphics processing operations. However, the applicant has recognized that in typical graphics processing applications, in many cases it may not actually be necessary to stop the execution of the "main" shader program (and in such cases, it may be quite simply necessary to check whether any initial "bootstrap" shader program is currently executing, but if not, then the "main" shader program can then be executed as normal, without further control (e.g., stopping) its processing).
[0098] Therefore, the present invention can provide various benefits compared to other methods.
[0099] As described above, the present invention relates to cases where an initial "bootstrap" shader program is used to execute certain expressions within an original shader program. For example, the use of such a "bootstrap" shader program is described in U.S. Patent No. 9,189,881 assigned to ARM Holdings plc, the entire contents of which are incorporated herein by reference.
[0100] Thus, the original shader program is effectively split into an initial "bootstrap" shader program and a corresponding subsequent "main" shader program that uses the results of the initial "bootstrap" shader program. The "main" shader program can thus contain load instructions that point to output values that have been generated and stored by executing the initial bootstrap shader program.
[0101] Thus, for an initial "bootstrap" shader program that is to be executed as part of a corresponding initial "bootstrap" processing job, an embodiment can further include subsequently executing a subsequent "main" shader program corresponding to the initial bootstrap shader program, the subsequent "main" shader program containing load instructions that point to output values generated and stored for a constant program expression by executing the initial bootstrap shader program. Subsequently executing the subsequent shader program can include: in response to the load instructions of the subsequent shader program that point to output values generated and stored for the constant program expression by executing the initial bootstrap shader program, loading the output values generated and stored for the constant program expression by executing the initial bootstrap shader program for processing by the subsequent shader program.
[0102] Subsequently executing a subsequent shader program corresponding to an initial bootstrap shader program in a set of initial bootstrap shader programs can be performed in any desired and suitable manner. For example, subsequently executing a subsequent shader program corresponding to an initial bootstrap shader program in a set of initial bootstrap shader programs can include executing the subsequent shader program on the (same) graphics processing pipeline.
[0103] Generally, any number (including zero) of initial "bootstrap" processing jobs can be defined for a particular processing pass. That is, a processing pass can typically include zero or more initial "bootstrap" processing jobs and any number of "main" processing jobs. The case where there are zero initial "bootstrap" processing jobs can still be effectively handled by the tracing of the present invention (because in this case, the tracing will always determine that no initial "bootstrap" processing job is being executed, such that no additional control (e.g., stopping) is required). Alternatively, if it can be identified earlier that there are no initial "bootstrap" processing jobs in a particular sequence of processing jobs to be executed, some or all of the present invention can be selectively disabled. That is, in some embodiments, the operation according to the present invention can be selectively triggered by a graphics processor (command processing unit (command stream front end / job manager (circuit))) identifying that the sequence of processing jobs includes one or more initial "bootstrap" processing jobs. In this regard, various arrangements are possible.)
[0104] The initial "bootstrap" shader program can be generated and configured in any appropriate and desired manner according to the specific requirements of the present invention.
[0105] For example, creating an initial bootstrap shader program in a set of initial bootstrap shader programs may include identifying constant program expressions in an original shader program. An implementation may then include creating an initial bootstrap shader program in the set of initial bootstrap shader programs, where the initial bootstrap shader program contains instructions for executing the constant program expressions.
[0106] The implementation may further include creating a subsequent shader program corresponding to an initial bootstrap shader program in the set of initial bootstrap shader programs, where the subsequent shader program contains a load instruction that points to an output value to be generated and stored for the constant program expression by executing the initial bootstrap shader program. Creating a subsequent shader program corresponding to an initial bootstrap shader program in the set of initial bootstrap shader programs may include: removing instructions for executing the constant program expression from the original shader program, and replacing the instructions for executing the constant program expression with a load instruction that points to an output value to be generated and stored for the constant program expression by executing the initial bootstrap shader program.
[0107] Identification of constant program expressions in the original shader program, creation of corresponding initial bootstrap shader programs, and creation of corresponding subsequent shader programs may be performed as needed. For example, identification of constant program expressions in the original shader program may identify those expressions in any suitable form during, for example, a compilation process, such as identifying them as "expressions" in a high-level shader language, or as a corresponding instruction set in object code for a graphics processing pipeline, or as a set of appropriate "operations" in an intermediate representation of the shader program.
[0108] Accordingly, identification of constant program expressions may be performed on or using an intermediate representation of the original shader program. Similarly, an initial bootstrap shader program containing instructions for executing constant program expressions may be created in the form of a higher-level shader language program, which is then converted to the necessary instructions for execution, for example, on a graphics processing pipeline, or may be created directly as an instruction set for execution, for example, on a graphics processing pipeline, or may be created in the form of an intermediate representation, which is then converted to instructions for execution, for example, on a graphics processing pipeline. Thus, creation of the initial bootstrap shader program may create the initial bootstrap shader program in an intermediate representation form, which intermediate representation form of the initial bootstrap shader program will then be translated into binary code "instructions" for execution, for example, on a graphics processing pipeline.
[0109] The identification of constant program expressions in the original shader program, the corresponding creation of an initial bootstrap shader program, and the corresponding creation of subsequent shader programs can be performed and are performed by any suitable stage or component of the graphics processing system. For example, a compiler for one or more of the shaders being discussed can perform such operations.
[0110] As discussed above, the original shader program to be executed by a given programmable shading stage will typically be provided by an application that requires the use of a high-level shader programming language such as GLSL, HLSL, OpenCL, etc. for graphics processing. This shader program is then translated by a shader language compiler into binary code for the target graphics processing pipeline. Thus, in an embodiment, the shader compiler can identify constant expressions in the original shader program being discussed, prevent instructions for executing those constant program expressions from being emitted into the target graphics processing pipeline binary code, alternatively create a separate binary code containing only the hardware instructions for the constant program expressions, and then provide the relevant binary code to the graphics processing pipeline for execution.
[0111] Constant program expressions can include any desired and suitable constant program expressions, such as those that operate on constant inputs. Constant inputs can include any desired and suitable inputs, such as inputs that do not change and / or will not change between draw calls. Constant inputs can also or alternatively include inputs that can change between draw calls but are determined to be constant at runtime for one or more specific draw calls (i.e., "runtime constant inputs").
[0112] The output values of the initial bootstrap shader program can be stored as needed. The output values can be stored in such a way that when the replaced load instructions in a subsequent shader program are executed, those values can be loaded by the subsequent shader program and treated as input values. Thus, the output of the initial bootstrap shader program can be a memory region that stores the input values for subsequent shader programs. This memory region can be any storage device accessible via graphics processing pipeline instructions in the shader program (such as main memory, stack memory, tile buffers, uniform memory, etc.). Such a memory region can be directly addressed, remapped to a color buffer (render target), or remapped to the stack region of the initial bootstrap shader program.
[0113] The output of the initial bootstrap shader program can be mapped (written to) the color buffer (render target), and the output color buffer of the initial bootstrap shader program is then mapped to the input uniform of the corresponding subsequent shader program (i.e., the load instructions in the subsequent shader program can point to and indicate a load from the output color buffer to be generated by the initial bootstrap shader program).
[0114] In this regard, various arrangements will be possible.
[0115] The present invention can generally be applied to any suitable graphics processing system.
[0116] According to specific requirements of the present invention, when operating in accordance with the present invention, the graphics processor can be used for all forms of output that the graphics processing pipeline can be used to generate, such as frames for display, output rendered to a texture, etc. The present invention can generally be used for both graphics and non-graphics (e.g., "compute") workloads as well as hybrid workloads.
[0117] Thus, the processing job being executed can generally include any suitable and desired processing job. For example, this can include rendering jobs in which the shader program includes a fragment shader program. The present invention can find a particular utility in this context because the fragment front-end processing stage in the graphics processing pipeline can generally be relatively important when executing fragment rendering jobs (e.g., compared to "compute" jobs where there may be relatively little front-end processing). Thus, the ability to pre-load such fragment jobs into a set of processing (shader) cores can provide a significant reduction in latency. However, generally speaking, the present invention can be used for any type of processing job (including "compute" jobs) that can be executed by the graphics processor and in which an initial "bootstrap" shader program can be utilized.
[0118] In a preferred embodiment, the present invention relates to a tile-based rendering system in which the rendering output (e.g., a frame) is subdivided into a plurality of rendering tiles for rendering purposes. In this case, each rendering tile can and preferably does correspond to a respective sub-region of the overall rendering output (e.g., frame) being generated. For example, the rendering tiles can correspond to rectangular (e.g., square) sub-regions of the overall rendering output.
[0119] In a preferred embodiment, rasterization is used to perform the rendering. However, it should be understood that the present invention is not necessarily limited to rasterization-based rendering and can generally be used for other types of rendering, including ray tracing or hybrid ray tracing arrangements.
[0120] In some embodiments, a graphics processing system includes and / or communicates with one or more memories and / or memory devices that store the data described herein and / or store software for performing the processes described herein. The graphics processing system may also communicate with a host microprocessor and / or with a display for displaying an image based on data generated by the graphics processing system.
[0121] In a particularly preferred embodiment, the various functions of the present invention are performed on a single graphics processing platform that generates and outputs rendered data, which is, for example, written to a frame buffer for a display device.
[0122] The present invention may be implemented in any suitable system such as a suitably configured microprocessor-based system. In a preferred embodiment, the present invention is implemented in a computer- and / or microprocessor-based system.
[0123] The various functions of the present invention may be performed in any desired and suitable manner. For example, the functions of the present invention may be implemented in hardware or software as needed. Thus, for example, the various functional elements, stages, and pipelines of the present invention may include one or more suitable processors, one or more controllers, functional units, circuits, processing logic, microprocessor arrangements, etc. that are capable of operating to perform the various functions, etc., such as suitably configured dedicated hardware elements or processing circuits and / or programmable hardware elements or processing circuits that may be programmed to operate in a desired manner.
[0124] It should also be noted here that, as will be appreciated by those skilled in the art, the various functions of the present invention, etc., may be repeated and / or performed in parallel on a given processor. Similarly, if desired, the various processing stages may share processing circuitry.
[0125] Accordingly, the present invention extends to graphics processors and graphics processing platforms that include means for operating in accordance with any one or more aspects of the present invention described herein. Such a graphics processor may originally include any one or more or all of the usual functional units, etc. included in a graphics processor, depending on the hardware required to perform the specific functions discussed above.
[0126] Those skilled in the art should also understand that all aspects and embodiments described in the present invention can and preferably do suitably include any one or more or all of the preferred and optional features described herein.
[0127] The method according to the present invention can be implemented at least in part using software, such as a computer program. Thus, it can be seen that, when considered from other aspects, the present invention provides: computer software which, when installed on a data processing device, is particularly adapted to execute the method described herein; a computer program element which includes a portion of computer software code for executing the method described herein when the program element runs on a data processing device; and a computer program which includes code means adapted to execute all steps of the method described herein when the program runs on a data processing system. The data processor may be a microprocessor system, a programmable FPGA (Field Programmable Gate Array), etc.
[0128] The present invention also extends to a computer software carrier comprising such software which, when used to operate a graphics processor, a renderer or a microprocessor system comprising a data processing device, causes the steps of the method of the present invention to be carried out in conjunction with the said graphics processing device, processor, renderer or system. Such a computer software carrier may be a physical storage medium, such as a ROM chip, RAM, flash memory, CD ROM or disk, or it may be a signal, such as an electrical signal via a wire, an optical signal or a radio signal, such as a signal to a satellite, etc.
[0129] It should also be understood that not all steps of the method of the present invention need to be carried out by computer software, and thus, in a broader aspect, the present invention provides computer software and such software installed on a computer software carrier for carrying out at least one of the steps of the method stated herein.
[0130] The present invention can thus suitably be embodied as a computer program product for use with a computer system. Such an embodiment may include a series of computer-readable instructions fixed on a tangible medium, such as a non-transitory computer-readable medium, for example a disk, CD ROM, ROM, RAM, flash memory or hard disk. It may also include a series of computer-readable instructions which can be transmitted to the computer system via a modem or other interface device through a tangible medium (including but not limited to an optical communication line or an analog communication line) or passively using wireless technology (including but not limited to microwave, infrared or other transmission technologies). The series of computer-readable instructions embody all or part of the functions described previously herein.
[0131] Those skilled in the art will appreciate that such computer-readable instructions can be written in a variety of programming languages for use with many computer architectures or operating systems. Additionally, such instructions can be stored using any current or future memory technology (including but not limited to semiconductor, magnetic, or optical technologies), or transmitted using any current or future communication technology (including but not limited to optical, infrared, or microwave technologies). It is contemplated that such computer program products can be distributed as a removable medium with an accompanying printed or electronic document (e.g., shrink-wrapped software), pre-loaded with a computer system, such as on a system ROM or fixed disk, or distributed from a server or electronic bulletin board via a network (e.g., the Internet or World Wide Web).
[0132] The preferred embodiments of the present invention will now be described solely by way of example and with reference to the accompanying drawings, in which:
[0133] Figure 1 An exemplary computer graphics processing system is shown;
[0134] Figure 2 A graphics processing pipeline capable of operating in the manner of the techniques described herein is schematically shown;
[0135] Figure 3 The use of a "bootstrap" shader program within the techniques described herein is shown;
[0136] Figure 4 The scheduling of a graphics processor working to a graphics processor shader core according to an embodiment is schematically shown;
[0137] Figure 5A A job submission process according to an embodiment is shown;
[0138] Figure 5B A job completion process according to an embodiment is shown;
[0139] Figure 6A A task submission process according to an embodiment is shown;
[0140] Figure 6B A task completion process according to an embodiment is shown;
[0141] Figure 7 A task processing process according to an embodiment is shown; and
[0142] Figure 8 The scheduling of a graphics processor working to a graphics processor shader core according to another embodiment is schematically shown.
[0143] Figure 1Illustrates a typical computer graphics processing system. An application 2 such as a game executing on a host processor (CPU) 1 will require graphics processing operations to be performed by an associated graphics processing unit (GPU) (graphics processor) 3, which executes a graphics processing pipeline. To this end, the application will generate API (application programming interface) calls, which are interpreted by a driver 4 for the graphics processor 3 running on the host processor 1 to generate appropriate commands for the graphics processor 3 to generate the graphics output required by the application 2. To facilitate this, a set of "commands" will be provided to the graphics processor 3 in response to commands from the application 2 running on the host system 1 for graphics output (e.g., to generate a frame to be displayed).
[0144] As Figure 1 shown, the graphics processing system will also include an appropriate memory system 5 for use by the host CPU 1 and the graphics processor 3.
[0145] When a computer graphics image is to be displayed, it is typically first defined as a series of primitives (polygons), which are then divided (rasterized) into graphics fragments, which are in turn used for graphics rendering. During normal graphics rendering operations, the renderer will modify the (e.g.) color (red, green, and blue, RGB) and transparency (α, a) data associated with each fragment so that the fragment can be correctly displayed. Once these fragments have fully traversed the renderer, their associated data values are stored in memory ready for output, e.g., for display.
[0146] In an embodiment of the present invention, graphics processing is performed in a pipelined manner, where one or more pipeline stages operate on data to generate a final output (e.g., the displayed frame).
[0147] Figure 2 Illustrates an exemplary graphics processing pipeline 10 that can be executed by the graphics processor 3 according to one embodiment. Figure 2 The shown graphics processing pipeline 10 is a "tile-based" rendering system and will thus produce tiles of an output data array, such as the output frame to be generated. Thus, an example will now be described in the context of "tile-based" rendering. In Figure 2 this, rasterization is used to perform the rendering, as will be further explained below. However, it should be understood that the present invention is not necessarily limited to rasterization-based rendering and can generally be used for other types of rendering, including ray tracing or hybrid ray tracing arrangements. Similarly, the present invention is not necessarily limited to tile-based rendering and can also be used for other types of rendering, including immediate mode rendering arrangements.
[0148] The output data array can generally be an output frame intended to be displayed on a display device such as a screen or a printer, but can also include, for example, the "render-to-texture" output of a graphics processor or other suitable arrangements.
[0149] Figure 2 The main elements and pipeline stages of a graphics processing pipeline 10 according to an embodiment of the present invention are shown. As will be understood by those skilled in the art, there may be other elements of the graphics processing pipeline that are Figure 2 not shown herein.
[0150] It should also be noted here that Figure 2 is only schematic, and in practice, for example, even though the functional units and pipeline stages shown are schematically shown as individual stages in Figure 2 herein, they may also share important hardware circuits. Similarly, some of the elements depicted in Figure 2 herein need not be provided, and Figure 2 only shows an example of the graphics processing pipeline 10. It should also be understood that each of the stages, elements, units, etc. of the graphics processing pipeline as Figure 2 shown can be implemented as needed and will accordingly include, for example, appropriate circuits and / or processing logic components, etc. for performing the required operations and functions.
[0151] As Figure 2 shown, the graphics processing pipeline will be executed on and implemented by a graphics processing unit (GPU) (graphics processor) 3, and thus, the graphics processing unit will include functional units, processing circuits, etc. capable of operating to perform the functions required for the graphics processing pipeline stages.
[0152] (It should be understood that the graphics processing unit (GPU) (graphics processor) 3 can and typically will include Figure 2 various other functional units, processing circuits, etc. not shown herein. This can include various functional units, processing circuits, etc. for performing non-graphics processing work. For example, in addition to graphics processing work, the graphics processing unit (GPU) (graphics processor) 3 may also be capable of operating to perform general "computing" operations and thus may also include various functional units, processing circuits, etc. capable of operating to perform such non-graphics processing work. Thus, although Figure 2Not shown, but in addition to the fragment shader endpoint 21 described below, the shader core 38 may also include a suitable "compute" shader endpoint that is operable and configured to issue compute tasks to the execution engine 31 for processing. For example, the shader core 38 may also include other suitable endpoints as needed, which are operable and configured to issue other types of tasks to the execution engine 31 for processing. In this regard, various arrangements will be possible.)
[0153] Figure 2 Shows the stages of the graphics processing pipeline after the assembler (not shown) of the graphics processor has prepared the primitive list (since the graphics processing pipeline 10 is a tile-based graphics processing pipeline).
[0154] (In fact, the assembler determines which primitives need to be processed for different regions of the output. In an embodiment of the present invention, these regions may be represented, for example, as tiles into which the overall output is divided for processing purposes, or a collection of multiple such tiles. To this end, the assembler compares the position of each primitive to be processed with the position of the regions, and adds the primitive to the corresponding primitive list of each region that the assembler determines the primitive can (possibly) fall into. Any suitable and desired technique for classifying and binning primitives into tile lists (such as exact binning or bounding box binning or any technique in between) may be used for the assembly process.)
[0155] Once the assembler has completed the preparation of the primitive lists (the primitive lists to be processed for each region), then each assembler can be rendered according to its associated primitive list.
[0156] To this end, each tile is processed by Figure 2 the graphics processing pipeline stage shown.
[0157] Thus, a fragment task iterator 20 is provided for scheduling processing work into the graphics processing pipeline 10.
[0158] Thus, the fragment task iterator 20 can schedule the graphics processing pipeline to generate a first output, which can be, for example, a frame to be displayed. In an embodiment of the present invention, where the graphics processing pipeline 10 is a tile-based system where the output has been divided into multiple rendering tiles, the graphics processing pipeline 10 iterates over the set of tiles of the first output, thereby rendering each tile in turn.
[0159] As Figure 2As shown, the graphics processor 3 includes a master controller in the form of a job manager circuit (command stream front end circuit) 35 that is operable to receive tasks for processing by the graphics processor 3 from the host processor 1, and the job manager 35 can then transfer the relevant jobs to the corresponding elements of the graphics processor and the graphics processing pipeline 10 via an appropriate bus / interconnect.
[0160] Thus, as Figure 2 shown, the job manager 35 will in particular issue fragment processing tasks to the fragment task iterator 20 for the fragment task iterator 20 to then schedule and dispatch appropriate fragment shading tasks onto the graphics processing pipeline 10.
[0161] In an embodiment of the present invention, the graphics processing pipeline 10 is implemented by means of appropriate processing (“shader”) cores. Specifically, as Figure 2 shown, the graphics processor 3 includes a plurality of “shader” cores, each of which is configured to implement a corresponding parallel instance of the graphics processing pipeline 10. Thus, the fragment task iterator 20 is operable and configured to issue tasks to different shader cores among the shader cores 38, for example to attempt to balance the processing jobs between different shader cores.
[0162] (Although Figure 2 not shown in the figure, there may be various other task iterators that control the issuance of “compute” or other tasks, etc.)
[0163] As will be further explained below, each “shader” core includes a fragment “front end” 30 and a programmable stage (execution engine 31), the fragment “front end” of which can generally be implemented in essentially fixed-function hardware and perform the setup for the fragment shader program, and the programmable stage executes the fragment shader program to perform actual rendering.
[0164] When a rendering task (i.e., a tile) is assigned to a given shader core 38 for processing, then the tile is processed (rendered) accordingly (i.e., by the graphics processing pipeline 10).
[0165] For a given tile being processed, the primitive list reader (or “polygon list reader”) 22 thus identifies a series of primitives to be processed for the tile (the primitives listed in the primitive list of the tile), and then issues the ordered primitive sequence of the tile into the graphics processing pipeline 10 for processing.
[0166] The resource allocator 23 then configures and manages the memory space allocation for depth (Z), color, etc. and the buffer 33 for the tile of the output being generated. For example, these buffers can be provided as part of the RAM located on the graphics processing pipeline (chip) (locally).
[0167] The vertex loader 24 then loads the vertices of the primitive, which are then passed into a primitive setup unit (or "triangle setup unit") 25 that operates to determine edge information representing the edges of the primitive based on the vertices of the primitive.
[0168] The edge information of the reordered primitive is then passed to a rasterizer 27 that rasterizes the primitive into a set of one or more sample positions and generates individual graphics fragments with appropriate positions (representing the appropriate sample positions) according to the primitive for rendering the primitive.
[0169] The fragments generated by the rasterizer 27 are then sent forward to the remainder of the pipeline for processing.
[0170] For example, in an embodiment of the present invention, the fragments generated by the rasterizer 27 are subject to an (early) depth (Z) / stencil test 29 to see if any fragments can be discarded (culled) at this stage. To this end, the Z / stencil test stage 29 compares the depth values of the fragments (associated therewith) issued from the rasterizer 27 with the depth values of the fragments that have been rendered (these depth values are stored in a depth (Z) buffer that is part of a tile buffer 33) to determine whether the new fragments will be blocked by the rendered fragments. Meanwhile, an early stencil test is performed.
[0171] Then, the fragments passing through the fragment early Z and stencil test stage 29 are subject to additional culling operations, such as a "forward pixel kill" test, for example as described in U.S. Patent Application Publication No. 2019 / 0088009 (Arm Holdings d), and the remaining fragments are then passed to a fragment shading stage (in the form of an execution engine 31) for rendering.
[0172] The processing stages including the primitive list reader (or "polygon list reader") 22 up to the (early) depth (Z) / stencil test 29 together constitute a fragment "front end" 30 that is used to set up the data required to perform the fragment processing operations executed by the execution engine 31.
[0173] The execution engine 31 then performs appropriate fragment processing operations on the fragments that have passed the early Z and stencil tests in order to process the fragments to generate appropriate rendered fragment data.
[0174] The fragment processing can include any suitable and desired fragment shading processes, such as executing a fragment shader program for the fragment, applying a texture to the fragment, applying fog or other operations to the fragment, etc., to generate appropriate rendered fragment data.
[0175] Thus, as Figure 2As shown, in an embodiment of the present invention, the execution engine 31 includes a programmable execution unit (engine) 32 that is operable to execute a fragment shader program for a corresponding execution thread (where each thread corresponds to one work item for the output being generated, such as an individual fragment) in order to perform the required fragment shading operations, thereby generating the rendered fragment data. In this regard, the execution unit 32 operates in any suitable and desired manner and includes any suitable and desired processing circuitry, etc.
[0176] In an embodiment of the present invention, the execution threads may be arranged into thread "groups" or "warps", where the threads in a group run in lockstep, one instruction at a time, i.e., each thread in the group executes the same single instruction before moving on to the next instruction. In this way, instruction fetch and scheduling resources can be shared among all the threads in the group. Such thread groups may also be referred to as "subgroups", "warps", and "wavefronts". For convenience, the term "thread group" will be used herein, but this is intended to cover all equivalent terms and arrangements unless otherwise specified.
[0177] Accordingly, Figure 2 Also shown is a thread group controller in the form of a warp manager 34 that is configured to control the assignment of work items (e.g., fragments) to corresponding thread groups for the programmable execution unit 32 to perform fragment shading operations, and to issue the thread groups to the programmable execution unit 32 for the corresponding thread groups to execute the fragment shader program.
[0178] As Figure 2 shown, the programmable execution unit 32 also communicates with the memory 5.
[0179] Once the fragment shading is complete, the output rendered (colored) fragment data is written to the tile buffer 33, from which the output rendered (colored) fragment data can be output to a frame buffer (e.g., in the memory 5) for display. The depth value of the output fragment is also appropriately written to the Z-buffer within the tile buffer 33. (The tile buffer stores a color buffer and a depth buffer that store the appropriate color, etc. or Z value respectively for each sampling position represented by the buffer (essentially for each sampling position of the rendered tile being processed).) These buffers store an array of fragment data representing parts (tiles) of the overall output (e.g., the image to be displayed), where the corresponding set of sampling values in the buffer corresponds to the corresponding pixels of the overall output (e.g., each 2×2 set of sampling values may correspond to an output pixel, where 4× multisampling is used).
[0180] As described above, the tile buffer 33 is typically provided as part of the RAM located on the graphics processor (locally).
[0181] Once the tiles for output have been processed, the data in the tile buffer can be written back to an external memory output buffer, such as a frame buffer of a display device (not shown), e.g., in memory 5. (The display device may include, for example, a display including a pixel array, such as a computer monitor or a printer.)
[0182] Then the next tile is processed, and so on, until enough tiles have been processed to generate the entire output (e.g., a frame (image) to be displayed). Then the process is repeated for the next output (e.g., frame), and so on.
[0183] As described above with respect to Figure 2 the graphics processing pipeline 10 includes a programmable processing or "shader" stage in the form of a fragment shading stage (but generally there may also be various other "shader" stages, such as vertex shaders, hull shaders, domain shaders, geometry shaders, etc.) for executing respective shader programs, these respective shader programs having one or more input variables and producing a set of output variables and being provided by an application program. To this end, the application program 2 provides shader programs implemented using a high-level shader programming language (such as GLSL, HLSL, OpenCL, etc.). These shader programs are then translated by a shader language compiler into binary code for the target graphics processing pipeline 10. This may include, for example, creating one or more intermediate representations of the program within the compiler. (The compiler may be, for example, a part of the driver 4, where there are special API calls to cause the compiler to run. Thus, the execution of the compiler can be regarded as part of the draw call preparation done by the driver in response to API calls generated by the application program).
[0184] Shader programs typically contain constant expressions (constructs expressed in the shader language with constant inputs). These constant expressions can be classified into two types: compile-time constant expressions (defined in the language specification, such as literal values, arithmetic operators with constant variables, etc.); and run-time constant expressions. Run-time constant expressions are not defined anywhere, but can be regarded as global variables that are known to be constant for a particular draw call (i.e., for all pipeline stages used within that draw call) and for all operations within the shader program that depend only on the global variable in question. In the case of run-time constant expressions, the compiler does not know the value of the variable at compile time. An example of such a variable is an expression identified as "uniform" in a GLSL shader program.
[0185] Examples of expressions that operate on runtime constants include: global variables that are known to be constant for a particular draw call; constant expressions as defined in the shader language specification; shader language expressions formed by operators on operands that are all runtime constant expressions; and shader language constructs that are defined as constant expressions in the language specification and all of whose operands are runtime constants.
[0186] This implementation particularly relates to a situation where a graphics processor is executing a pass that includes one or more initial processing jobs, where the initial processing jobs execute corresponding initial shader programs to be executed before a corresponding "main" shader program that will be executed for a separate "main" processing job within the same pass.
[0187] For example, such runtime constant expressions in a shader program can be identified and extracted from the original shader program such that the runtime constant expressions are alternatively executed in an initial shader program ("bootstrap" shader program).
[0188] To this end, a shader compiler identifies such runtime constant expressions in a given shader program to be executed, removes such expressions from the original shader program (preventing such expressions from being emitted into the target GPU code), and creates a separate shader program (binary code) containing hardware instructions only for the identified expressions and metadata for those expressions to thereby create a "bootstrap" shader program that can be executed before the main shader program. The metadata contains information necessary to execute the "bootstrap" shader program in the graphics processing pipeline and later be able to fetch the results of the bootstrap shader program. Thus, the metadata can include, for example, one or more of a memory layout for the input, a memory layout for the output, and / or a description of where the output is written. The metadata can be different for different architectures / implementations.
[0189] The compiler also replaces the original runtime constant expressions in the main shader program with appropriate load instructions that point to where the output results from the bootstrap shader program will be stored.
[0190] This is done for some shader programs and in one implementation all shader programs in a shader program to be executed for a given desired graphics processing output.
[0191] Figure 3 This process is illustrated. As Figure 3As shown, the shader compiler will receive a shader program in a high-level programming language to be compiled (step 40), and first identify any runtime constant expressions in the shader program (step 41). Then, the shader compiler will remove the instructions emitted for such expressions from the original shader program, and replace those instructions in the original main shader program with appropriate load instructions pointing to where the output results from the boot shader program will be stored (step 42). The shader compiler then creates a separate shader program (binary code) containing only the hardware instructions for the identified runtime constant expressions and any necessary metadata for those instructions, thereby creating a "boot" shader program that can be executed before the main shader program (step 43).
[0192] In this embodiment, the compiler configures the boot shader programs such that they output to the color buffer (render target) in the tile buffer 33. The corresponding load instructions replaced into the main shader program then map to this color buffer, such that the main shader program will use the results of the boot shader program as its input (when needed). Other arrangements for the output of the boot shader can be used if desired, such as remapping the stack area of the boot shader. Generally, any storage device (such as main memory, stack memory, tile buffer, uniform memory, etc.) that can be accessed via graphics processing pipeline instructions in the shader program can be used for the output of the boot shader program.
[0193] Once the boot shader program and the main shader program for execution have been compiled, the boot shader program is executed on the graphics processing pipeline 10 (step 44), followed by the execution of the modified main shader program (step 45). To this end, the driver 4 on the host processor 1 of the graphics processing unit 3 initializes the data required for the draw call, creates a dependency chain of the necessary jobs for the draw call phase and the boot shader, and then sends these jobs to the graphics processing pipeline 10 for execution.
[0194] In some more traditional graphics processor operations, the driver 4 then ensures that the created boot shader program is executed after the data for the relevant draw call has been initialized but before the corresponding draw call phase is activated (i.e., before the main shader program is executed). This then ensures that the boot shader is executed on the graphics processing unit 3 such that all computations of the boot shader are completed before the main shader program is called. Thus, in some more traditional graphics processor operations, these dependencies are managed by including explicit strict (hard) barriers in the graphics processor command stream, such as by inserting an appropriate "wait" command between the "run_fragment boot" and the main "run_fragment" commands, such as the following:
[0195] run_fragment bootstrapping
[0196] Wait for the bootstrapping to complete
[0197] run_fragment
[0198] ……
[0199] This more traditional graphics processor operation ensures that all bootstrapping shader programs finish before any main shader program that can use the results of the bootstrapping shader programs is executed. However, the bootstrapping shader programs can often take thousands of cycles to execute and typically do not fully utilize the shader cores. Thus, due to the strict (hard) barriers within the graphics processor command stream, the job manager circuit (command stream front-end circuit) 35 cannot start issuing render tasks for the next "run_fragment" rendering job for processing until the bootstrapping shader program on which the rendering job potentially depends has finished execution. Typically, the bootstrapping shader program can execute on only a single core and thus all other shader cores can be idle for several cycles, waiting for the single core to finish executing the bootstrapping shader program.
[0200] Accordingly, in accordance with the present embodiment, the task manager circuit (command stream front-end circuit) 35 is permitted to issue dependent main processing tasks to the respective shader cores 38 for processing concurrently with the respective bootstrapping tasks of these main processing tasks, thereby removing the strict (hard) barriers between the bootstrapping tasks and their dependent main processing tasks. This then effectively allows the graphics processor to run ahead of its current command stream execution and start performing at least some processing of the dependent main processing jobs without having to wait for any associated bootstrapping shader program to have finished its execution. This can thus increase the average shader core utilization and also help reduce the latency by effectively preloading jobs into the task iterator 20 and the shader cores. In contrast, in the more traditional graphics processor operation described above, the need to enforce strict (hard) barriers within the graphics processor command stream can mean that the execution of the task iterator 20 and / or the shader cores 38 stops, waiting for the bootstrapping jobs to complete.
[0201] Any dependencies between the bootstrapping processing jobs and the main processing jobs within the same processing pass can then be (and in the present embodiment are) enforced within the shader cores 38 as needed, rather than enforcing strict (hard) processing barriers within the graphics processor stream. In particular, to facilitate the management of these dependencies, as Figure 4 shown, the task iterator 20 can be operative and configured to maintain a reference counter 420 (or a set of reference counters 820, as will be described below with respect to Figure 8(which is further explained), the reference counter keeps track of whether any boot job is currently in progress. This tracking can then be used to control the execution of any potentially dependent main processing jobs within the shader core 38, for example to ensure that any main shader program is not executed until the corresponding boot shader program for that main shader program has ended.
[0202] The overall control operation according to the present embodiment will now be described.
[0203] Figure 5A is a flowchart showing the operation of the job manager circuit (command stream front-end circuit) 35 when a processing job is issued to the task iterator 20 according to the present embodiment. As Figure 5A shown, as long as there are still jobs to be issued (step 51 - Yes), the job manager circuit (command stream front-end circuit) 35 attempts to issue these jobs to the task iterator 20, for example, as described above. If the job is not strictly dependent on any previous job such that strict (hard) dependencies do not need to be enforced (step 52 - No), then the job can be issued to the task iterator 20 accordingly, which controls the issuance of tasks to the shader core 38 without having to wait for any previous job to complete. Once all jobs have been issued (step 51 - No), the job manager issuance process is complete.
[0204] As described above, according to the present embodiment, main processing jobs are no longer considered to be strictly dependent on their corresponding boot jobs (i.e., the boot jobs within the same processing pass). Thus, in the present embodiment, the job manager circuit (command stream front-end circuit) 35 is operable to issue main processing jobs to the task iterator (i.e., at step 54) without waiting for any boot jobs on which these main processing jobs depend to have ended. That is, according to the present embodiment, there is no need for an explicit "ait" command to be inserted between the "run_fragment boot" command and any dependent main "run_fragment" commands within the same processing pass, and thus the command stream associated with that processing pass can be, for example, as follows:
[0205] run_fragment boot
[0206] run_fragment
[0207] ……
[0208] On the other hand, there may be cases where strict (hard) barriers should still be enforced within the graphics processor command stream and the job manager circuit (command stream front-end circuit) 35 is accordingly operable and configured to do so (i.e., at step 52).
[0209] An example of this situation is before any new boot job is issued, for example for use in subsequent processing passes, because in this case, it may still be necessary to enforce a strict (hard) barrier before issuing the new boot job. For example, in Figure 4 In the example shown, a single reference counter 420 is used to track how many boot jobs are currently in progress, without attempting to track which processing pass these boot jobs pertain to. In this case, it may actually be safe to issue the boot job for a later processing pass for processing simultaneously with the main processing job of the previous processing pass that has not yet ended (because the main processing job typically only has dependencies on the corresponding boot shader programs within the same processing pass). However, Figure 4 The rough tracking shown in
[0210] run_fragment boot 1 for processing pass 1
[0211] run_fragment main 1 for processing pass 1
[0212] Wait for all jobs in processing pass 1 to end
[0213] run_fragment boot 1 for processing pass 2
[0214] run_fragment main 1 for processing pass 2
[0215] ……
[0216] In this case, when it is determined that strict dependencies should be enforced (step 52 - yes), the job manager circuit (command stream front-end circuit) 35 then should and does wait for any previous jobs to complete (step 53), and then issues the dependent jobs to the task iterator for processing (step 54).
[0217] In the case where there are two processing passes and the later processing pass has dependencies on the previous processing pass, it may also be determined whether the boot shader program for the later processing pass has any dependencies on the earlier processing pass. If not, then the "wait" command can be further shifted back, for example as follows:
[0218] run_fragment boot 1 for processing pass 1
[0219] run_fragment main 1 for pass 1
[0220] run_fragment leading 1 for pass 2
[0221] Wait for all jobs in pass 1 to finish
[0222] run_fragment main 1 for pass 2
[0223] …
[0224] This then allows the leading shader jobs for later passes to be issued earlier, which will also avoid the problem of low shader core utilization.
[0225] Accordingly, there are various examples of possible dependencies where it may still be desirable to enforce stricter (harder) processing barriers between processing jobs within the graphics processor command stream, and this can be done, for example, as is typically the case, by inserting appropriate "wait" commands into the graphics processor command stream. These "wait" commands then cause the job manager circuit (command stream front end circuit) 35 to wait for any previous jobs that cause possible dependencies to finish before issuing the next job to the task iterator 20.
[0226] Figure 5B is a flowchart showing the corresponding completion process of the job manager circuit (command stream front end circuit) 35. Accordingly, as Figure 5B shown, the job manager circuit (command stream front end circuit) 35 is operable to receive a response (step 55) when a processing job is completed (from the task iterator 20). When a processing job that causes the task manager circuit (command stream front end circuit) 35 to issue a stream stop (i.e., at step 53) is completed, the completion status of that job can be notified accordingly (step 56). The wait condition (i.e., at step 53) is thus removed, and dependent jobs can proceed.
[0227] In this regard, various other arrangements would be possible. For example, in other embodiments, as Figure 8 shown, the task iterator 20 is operable and configured to maintain a set 820 of reference counters that track on a per-pass basis whether any leading jobs are currently in progress. Accordingly, as Figure 8 shown, each reference counter in the set 820 of reference counters is associated with a corresponding job identifier. In this case, dependencies can be tracked and enforced within individual passes. In that case, then it can be allowed to issue leading jobs for later passes to be processed concurrently with the main processing jobs for earlier passes.
[0228] As described above, the job manager circuit (command stream front-end circuit) 35 issues jobs to the task iterator 20, and then the task iterator decomposes these jobs into corresponding processing tasks for scheduling to the corresponding shader cores 38.
[0229] Figure 6A is a flowchart showing the operation of the task iterator 20 when processing tasks are issued to the corresponding shader cores 38 according to the present embodiment. The task iterator 20 issue process is thus triggered by the task iterator 20 receiving a job from the job manager circuit (command stream front-end circuit) 35 (i.e., triggered by Figure 5A step 54 in).
[0230] As Figure 6A shown, as long as there are tasks to be issued (step 61 - Yes), the task iterator 20 attempts to schedule the tasks to the corresponding shader cores 38 for processing. As part of this, it is checked (at step 62) whether the task involves a boot job. If the task is a boot job, then the reference counter 420 (or the corresponding reference counter from the set of reference counters 820) can be incremented accordingly (step 63) before issuing the task to the available shader core 38 for processing (step 64). On the other hand, if the task does not involve a boot job, the task is simply issued to the available shader core 38 for processing (step 64) without updating the reference counter.
[0231] Figure 6B is a flowchart showing the corresponding completion process of the task iterator 20. As Figure 6B shown, whenever a task is completed, the task iterator 20 receives a corresponding response from the shader core 38 (step 65). If the completed task is not a boot task (step 66 - No), then no further notification needs to be performed and the completion process is complete. However, if the completed task is a boot task (step 66 - Yes), then the reference counter 420 (or the corresponding reference counter from the set of reference counters 820) is decremented accordingly (step 67).
[0232] The reference counter thus keeps track of how many tasks associated with boot jobs are currently being processed. (It will be recognized here that a boot job can and usually will only contain a single task, and thus keeping track of the number of tasks associated with the currently ongoing boot job is generally equivalent to keeping track of the number of currently ongoing boot jobs. However, in principle, it is also possible to explicitly keep track of how many boot jobs are currently ongoing at the job level. In this regard, various arrangements will be possible.)
[0233] This tracking can accordingly be used to enforce dependencies between boot jobs and dependent main processing jobs within the shader core 38 as needed. For example,Figure 7 is a flowchart showing the processing control for tasks within shader core 38. In this example, and generally, tasks can be divided into a set of front-end operations that can be safely processed independently of any earlier boot job and a set of dependent operations that require the results of an earlier boot job. For example, independent operations can include operations performed in the fragment front-end 30 of shader core 38, while dependent operations can include the execution of the main shader program within the execution engine 31 of shader core 38.
[0234] When a task is issued to shader core 38, the independent part of the task can (and does) thus be processed immediately (step 72). If the task has no valid dependencies (step 73 - no), then it is safe to continue processing the dependent part of the task (and thus do so) (step 75). Once the dependent part of the task has been processed (i.e., at step 75), the task is complete and this is signaled back to task iterator 20 (i.e., at step 65). However, in the case where the task has valid dependencies (step 73 - yes), then the processing of the task should be stopped (step 74) until the dependencies are released, such that the processing of the dependent part of the task (i.e., at step 75) is gated, waiting for the dependencies to be released.
[0235] Thus, in the present embodiment, when task iterator 20 issues a task to shader core 38 for processing, if the task has valid dependencies (e.g., because there is currently an in-progress boot job that the task depends on), then preferably the task is annotated accordingly as having dependencies. For example, this can be done by setting a suitable dependency bit, where if such a dependency bit is set, then the processing of the task within shader core 38 stops (i.e., at step 74) until the dependencies are released. Conversely, if the dependency bit is cleared, then the processing of the task does not stop (i.e., it is determined at step 73 that there are no dependencies). Thus the dependency bit can be used to indicate whether a task has valid dependencies, which means that the processing of the dependent part of the task should be stopped within shader core 38.
[0236] Thus, if the reference counter maintained by task iterator 20 indicates that there is a currently ongoing boot job at the time of issuing a task to shader core 38 for processing such that the task has a valid dependency (i.e., the associated reference counter is non-zero), then the processing of the task is stopped within shader core 38 (i.e., at step 74) until the dependency is released. At some point during the graphics processing operation, the boot task of the boot job that caused the dependency will complete, and this will be signaled to the task iterator to cause the reference counter to decrement (i.e., at step 67). Once all boot tasks have completed and the reference counter has thus been decremented to zero (step 68), this means that the dependency can be released, and this is accordingly signaled to any waiting shader cores 38 (step 69) to release the dependency (i.e., at step 74).
[0237] Once all tasks in the job have been completed (step 70 - yes), then the job is completed, and this can be signaled back accordingly to job manager circuitry (command stream front end circuitry) 35 (step 71) to release any waiting jobs (i.e., by triggering the Figure 5B job manager circuitry (command stream front end circuitry) 35 completion flow shown in
[0238] Thus, in accordance with the present invention, it is possible to remove the strict (hard) barrier within the graphics processor command stream between a boot job and a subsequent main processing job within the same processing pass, where any dependencies between the main job and the boot job are then enforced within the shader core, e.g., as described above. This then allows the graphics processor to effectively run and start executing at least some processing of the dependent main processing jobs prior to the execution of its current command stream, without having to wait for any associated boot shader program to have finished its execution.
[0239] The foregoing detailed description is presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the technology described herein to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. The embodiments were chosen and described in order to best explain the principles of the technology and its practical application, to thereby enable others skilled in the art to best utilize the technology in various embodiments and with various modifications as are suited to the particular use contemplated. The scope of the invention is intended to be defined by the appended claims.
Claims
1. A method of operating a graphics processor, the graphics processor comprising a set of one or more processing cores, the method comprising: When executing a processing pass comprising one or more initial processing jobs, wherein the initial processing jobs execute respective initial shader programs to be executed prior to corresponding “main” shader programs to be executed for separate “main” processing jobs within the same processing pass, such that the “main” shader programs have a dependency on the initial shader programs of the initial processing jobs, and wherein the “main” processing jobs comprise respective sets of one or more tasks to be processed for the “main” processing jobs, each task being operable to execute respective instances of the “main” shader programs: tracking whether any initial processing job is currently being processed by the set of one or more processing cores; and When a respective task to be processed as part of a "primary" processing job is issued to a respective processing core for processing while the set of one or more processing cores is concurrently performing processing for at least one initial processing job: Based on the tracking of whether any initial processing job is currently being processed by the set of one or more processing cores, the processing of the task of the "main" processing job within the processing core is controlled so that the "main" shader program of the "main" processing job is not executed with respect to the task, at least until any initial processing job that is being processed simultaneously with the task and on which the "main" shader program has a dependency has completed its processing.
2. The method according to claim 1, comprising: When a task for a "main" processing job is issued to the corresponding processing core for processing: determining, based on the tracking, whether any initial processing jobs upon which the task has a potential dependency are currently being executed by the set of one or more processing cores; as well as When it is determined that at least one initial processing job on which the task has a potential dependency is currently being processed by the set of one or more processing cores at the time when the task is issued to the respective processing cores for processing: issuing the task to the corresponding processing core for processing concurrently with the at least one initial processing job being processed by the set of one or more processing cores; as well as performing at least some processing of said task up to but not including execution of said "main" shader program, but at least waiting until said at least one initial processing job being executed by said set of one or more processing cores concurrently with said task of said "main" processing job has completed its processing prior to executing said "main" shader program.
3. A method according to claim 2, wherein when it is determined that at the time when the task is issued to the corresponding processing core for processing, at least one initial processing job on which the task has a potential dependency is currently being processed by the set of one or more processing cores, the method includes indicating that execution of the "main" shader program of the task should be stopped until the at least one initial processing job completes its processing, the indication causing the processing core to stop execution of the "main" shader program until a signal indicating that the at least one initial processing job has completed its processing is received.
4. The method according to claim 3, further comprising: The processing core then receives a signal indicating that the at least one initial processing job has completed its processing; And in response to such a signal, the processing core continues processing the task, including executing the "main" shader program.
5. A method according to claim 1, 2 or 3, wherein the tracking whether any initial processing job is currently being executed by the set of one or more processing cores includes maintaining a reference counter indicating how many initial processing jobs are currently being executed by the set of one or more processing cores.
6. A method according to claim 1, 2 or 3, wherein the tracking of whether any initial processing job is currently being executed by the set of one or more processing cores is performed by a task issuance circuit that controls the issuance of tasks to the set of one or more processing cores.
7. A method according to claim 1, 2 or 3, wherein the tracking of whether any initial processing job is currently being processed by the set of one or more processing cores is performed on a per-processing pass basis, so that for individual processing passes, tracking is performed on whether any initial processing job is currently being processed by the set of one or more processing cores for the processing pass, and wherein the processing of the task that controls the "main" processing job includes controlling the processing of the task based on the tracking of whether any initial processing job is currently being processed by the set of one or more processing cores for the same processing pass including the "main" processing job.
8. A method according to claim 1, 2 or 3, wherein the sequence of processing jobs includes a first processing pass including one or more initial processing jobs and one or more subsequent "main" processing jobs, and a second processing pass including one or more initial processing jobs and one or more subsequent "main" processing jobs, and wherein a processing barrier is enforced between the first processing pass and the second processing pass so that the initial processing job of the second processing pass is not released for processing at the same time as any processing job of the first processing pass.
9. A method according to claim 1, 2 or 3, wherein a processing barrier is enforced in the processing job sequence before any initial processing job in the processing job sequence so that the initial processing job is not released for processing at the same time as any earlier processing job in the processing job sequence.
10. The method of claim 1 , 2 or 3, wherein the “main” processing job is a fragment rendering job in which a fragment shader is to be executed.
11. A graphics processor, comprising: A collection of one or more processing cores; task issuing circuitry operable and configured to issue tasks to the set of one or more processing cores for processing; and Control circuit, The control circuit is configured as follows: When the graphics processor is executing a processing pass comprising one or more initial processing jobs, wherein the initial processing jobs execute respective initial shader programs to be executed prior to corresponding “main” shader programs to be executed for separate “main” processing jobs within the same processing pass, such that the “main” shader programs have a dependency on the initial shader programs of the initial processing jobs, and wherein the “main” processing jobs comprise respective sets of one or more tasks to be processed for the “main” processing jobs, each task being operable to execute respective instances of the “main” shader programs: tracking whether any initial processing job of the current processing pass is currently being processed by the set of one or more processing cores; as well as Based on said tracking of whether any initial processing job of said current processing pass is currently being processed by said set of one or more processing cores, controlling the processing of tasks of a “main” processing job within said processing cores, such that when a corresponding task to be processed as part of a “main” processing job is issued to a corresponding processing core for processing while said set of one or more processing cores is simultaneously executing processing of at least one initial processing job within the same processing pass as said “main” processing job for which said task is executed: The “main” shader program of the “main” processing job is not executed with respect to the task, at least until any initial processing job for the current processing pass and upon which the “main” shader program has a dependency has ended.
12. The graphics processor according to claim 11, comprising: When a task for a "main" processing job is issued to the corresponding processing core for processing: determining, based on the tracking, whether any initial processing jobs upon which the task has a potential dependency are currently being executed by the set of one or more processing cores; as well as When it is determined that at least one initial processing job on which the task has a potential dependency is currently being processed by the set of one or more processing cores at the time when the task is issued to the respective processing cores for processing: issuing the task to the corresponding processing core for processing concurrently with the at least one initial processing job being processed by the set of one or more processing cores; as well as performing at least some processing of said task up to but not including execution of said "main" shader program, but at least waiting until said at least one initial processing job being executed by said set of one or more processing cores concurrently with said task of said "main" processing job has completed its processing prior to executing said "main" shader program.
13. A graphics processor according to claim 12, wherein when it is determined that at the time when the task is issued to the corresponding processing core for processing, at least one initial processing job on which the task has a potential dependency is currently being processed by the set of one or more processing cores, the method includes indicating that execution of the "main" shader program of the task should be stopped until the at least one initial processing job completes its processing, the indication causing the processing core to stop execution of the "main" shader program until a signal is received indicating that the at least one initial processing job has completed its processing.
14. The graphics processor of claim 13, the method further comprising: The processing core then receives a signal indicating that the at least one initial processing job has completed its processing; And in response to such a signal, the processing core continues processing the task, including executing the "main" shader program.
15. The graphics processor of claim 11, wherein said tracking whether any initial processing jobs are currently being executed by said set of one or more processing cores comprises maintaining a reference counter indicating how many initial processing jobs are currently being executed by said set of one or more processing cores.
16. The graphics processor of claim 11, wherein the tracking of whether any initial processing job is currently being executed by the set of one or more processing cores is performed by a task issuing circuit that controls issuing tasks to the set of one or more processing cores.
17. A graphics processor according to claim 11, wherein the control circuit is configured to track whether any initial processing job is being processed by the set of one or more processing cores on a per-processing pass basis, so that for individual processing passes, tracking whether any initial processing job is currently being processed by the set of one or more processing cores for the processing pass, and the control circuit is then configured to control the processing of tasks of the "main" processing job based on the tracking of whether any initial processing job is currently being processed by the set of one or more processing cores for the same processing pass including the "main" processing job.
18. A graphics processor according to claim 11, wherein the sequence of processing jobs includes a first processing pass including one or more initial processing jobs and one or more subsequent "main" processing jobs, and a second processing pass including one or more initial processing jobs and one or more subsequent "main" processing jobs, and wherein a processing barrier is enforced between the first processing pass and the second processing pass so that the initial processing job of the second processing pass is not issued for processing at the same time as any processing job of the first processing pass.
19. The graphics processor of claim 11, wherein a processing barrier is enforced in the sequence of processing jobs prior to any initial processing job in the sequence of processing jobs such that the initial processing job is not issued for processing concurrently with any earlier processing job in the sequence of processing jobs.
20. A computer readable medium storing computer software code which, when executed on one or more data processors, causes the data processors to perform a method according to claim 1, 2 or 3.
Citation Information
Patent Citations
Forward killing of threads corresponding to graphics fragments obscured by later graphics fragments
US20190088009A1
Graphics processing
US9189881B2