Graphics processing systems
The integration of a memory region access permission checking circuit in graphics processing systems addresses the challenge of secure memory access by ensuring that only authorized threads within the same trust domain can access work group local storage, thereby preventing data leakage and enhancing system security.
Patent Information
- Application Number
- PCT/GB2024/052921
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-12
AI Technical Summary
Existing graphics processing systems face challenges in securely managing memory access permissions for work groups, particularly in preventing data leakage across different trust domains when using work group local storage.
The implementation of a memory region access permission checking circuit that determines whether an execution thread is permitted to access a specific region of the work group local storage based on trust domain affiliation, ensuring that only authorized threads within the same trust domain can access allocated memory regions.
This solution enhances the security of work group local storage by preventing unauthorized access and data leakage between work groups from different trust domains, thereby ensuring data integrity and system security.
Smart Images

Figure GB2024052921_12062025_PF_FP_ABST
Abstract
Description
[0001] Graphics Processing Systems
[0002] BACKGROUND
[0003] The technology described herein relates to graphics processing systems (and graphics processors), and in particular to the operation of graphics processors / systems when using local storage to (temporarily) store data for a group of execution threads.
[0004] Many graphics processors include one or more processing (shader) cores, that execute, inter alia, programmable processing stages, commonly referred to as “shaders”, of a graphics processing pipeline that the graphics processor implements. For example, a graphics processing pipeline may include one or more of, and typically all of: a geometry shader, a vertex shader and a fragment (pixel) shader. These shaders are programmable processing stages that execute shader programs on input data values to generate a desired set of output data, such as appropriately shaded and rendered fragment data in the case of a fragment shader, for processing by the rest of the graphics processing pipeline and / or for output.
[0005] It is also known to use graphics processors and graphics processing pipelines, and in particular the shader operation of a graphics processor and graphics processing pipeline, to perform more general computing tasks, e.g. in the case where a similar operation needs to be performed in respect of a large volume of plural different input data values. These operations are commonly referred to as “compute shading” operations, and a number of specific compute APIs, such as OpenCL and Vulkan, have been developed for use when it is desired to use a graphics processor and a graphics processing pipeline to perform general computing operations. Compute shading is used for computing arbitrary information. It can be used to process graphics-related data, if desired, but is also used for tasks not directly related to performing graphics processing.
[0006] When performing “shader” processing, a graphics processor shader core will execute a (typically small) program for each “work item” in an output to be generated. In the case of generating a graphics output, such as a render target, such as a frame to be displayed, a “work item” in this regard is usually a vertex or a sampling position (e.g. in the case of a fragment shader). In the case of compute shading operations, each “work item” in the output being generated will be, for example, the data instance (item) in the work “space” that the compute shading operation is being performed on. In graphics processor shader operation, including in compute shading, each “work item” will be processed by means of an execution thread which will execute the instructions in the shader program in question for the work item in question.
[0007] In such arrangements, the work load of the graphics processor is commonly subdivided into respective groups of work items (and correspondingly execution threads), which are correspondingly referred to as “work groups”. A work group is typically a collection of a few dozen to a few hundred execution threads (corresponding to respective work items), that are all guaranteed to exist at the same time (i.e. having the same lifetime) and are able to perform communication and synchronisation with each other (for example using work group-wide barriers). These execution threads are normally but not necessarily, short-lived, with typical lifetimes being dozens to thousands of instructions.
[0008] In order to facilitate data sharing between the threads within a work group, a given (and each) work group will typically be allocated a memory region known as “work group local storage” that the threads in the work group can read from and write to. This memory region may be allocated from either normal system memory or from an on-chip resource, and remains available to all threads of the work group for the lifetime of the work group as a whole. When a work group reaches the end of its lifetime, the work group’s “local storage” memory region is normally deallocated, making it available for use by another work group.
[0009] The Applicants believe that there remains scope for improved graphics processor / system operation when performing processing for work groups in the above manner.
[0010] BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments of the technology described herein will now be described by way of example only and with reference to the accompanying drawings, in which:
[0012] Figure 1 shows an exemplary graphics processing system;
[0013] Figure 2 shows schematically an embodiment of a graphics processor that can be operated in the manner of the technology described herein;
[0014] Figure 3 shows the work group local storage shared memory in an embodiment of the technology described herein in more detail;
[0015] Figure 4 is a flow chart illustrating the operation of the warp manager when implementing a “trust domain”-based bounds check an embodiment;
[0016] Figure 5 is a flow chart illustrating the corresponding “trust domain”-based bounds check operation that is performed according to such an embodiment;
[0017] Figure 6 is a memory diagram showing the memory layout according to such an embodiment; Figure 7 is a flow chart illustrating the operation of the warp manager when implementing a “range”-based bounds check another embodiment;
[0018] Figure 8 is a flow chart illustrating the corresponding “range”-based bounds check operation that is performed according to such an embodiment;
[0019] Figure 9 is a memory diagram showing the memory layout according to such an embodiment;
[0020] Figure 10 shows the work group local storage shared memory in another embodiment of the technology described herein;
[0021] Figure 11 shows the work group local storage shared memory in another embodiment of the technology described herein;
[0022] Figure 12 shows the work group local storage shared memory in another embodiment of the technology described herein;
[0023] Figure 13 shows the work group local storage shared memory in another embodiment of the technology described herein;
[0024] Figure 14 shows the work group local storage shared memory in another embodiment of the technology described herein; and
[0025] Figure 15 shows the work group local storage shared memory in another embodiment of the technology described herein.
[0026] Like reference numerals are used for like elements in the Figures where appropriate.
[0027] DETAILED DESCRIPTION
[0028] A first embodiment of the technology described herein comprises a graphics processing system, the graphics processing system comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processing system; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the graphics processing system further comprising one or more memory region access permission checking circuits that control memory accesses to the storage, the one or more memory region access permission checking circuits configured to: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage.
[0029] A second embodiment of the technology described herein comprises a method of operating a graphics processing system, the graphics processing system comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processing system; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed; and one or more memory region access permission checking circuits that control memory accesses to the storage, the method comprising: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits: determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and preventing the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage.
[0030] The technology described herein relates to graphics processing systems (and graphics processors) that include a programmable processing unit (a “shader core”), that has access to storage that can be allocated for use by “work groups” executing on the programmable processing unit (as discussed above) (“work group local storage”).
[0031] The graphics processing system in the technology described herein further comprises a memory region “access permission checking” circuit that controls memory accesses to the storage (or plural such circuits that control (different) memory accesses to the storage) and which memory region “access permission checking” circuit(s), as will be discussed in more detail below, is operable and configured to perform a suitable memory region “access permission checking” operation such that respective regions of the (work group local) storage can only be (and are only) accessed by execution threads of respective work groups being executed by the programmable processing unit that are permitted to access the region(s) of the (work group local) storage in question (with the memory region “access permission checking” circuit(s) further operable and configured to prevent an execution thread from accessing a respective region of the storage that the execution thread is requesting access to when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage).
[0032] For instance, it will be appreciated that different work groups may generally reside in the same “trust domain” (e.g. when they come from the same application / process and / or virtual machine, etc.) but may also reside in different trust domains (e.g. when they come from different applications / processes and / or virtual machines, etc.). When different work groups reside in different trust domains, data should not be able to leak across the trust domains via the work group local storage.
[0033] Thus, in embodiments, an execution thread is determined to be other than (i.e. not) permitted to access a respective region of the (work group local) storage (at least) when that region of the storage is currently allocated for use by a work group that resides in a different trust domain to the work group including the execution that is requesting access to the region in question. This determination can be made in various suitable ways, as desired.
[0034] For instance, in an embodiment, this determination can be (and is) made on a per work group basis.
[0035] Thus, in embodiments, when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits are configured to (and perform as part of the method steps to): determine whether the respective region of the storage that the execution thread is requesting access to has been allocated for use by the group of work items including that execution thread; and prevent access to the respective region of the storage (that the execution thread is requesting access to) when it is determined that the request is being made for an execution thread other than an execution thread of the work group for which the respective region of the storage has been allocated for use by.
[0036] This then means that a respective work group is only permitted to (and can only) access its own respective work group local storage and that data cannot leak between different work groups that are being processed at the same time (regardless of whether those work groups come from the same or different trust domains).
[0037] The present Applicants also recognise however that it may be acceptable for data to leak between different work groups so long as it is known that those work groups come from the same trust domain. In that case, if different work groups within the same trust domain attempt to access each other’s data, this would merely represent an “undefined” access (e.g., and so could be handled using the normal procedure for handling such “undefined” accesses for the graphics processor / system in question), but does not present a security risk.
[0038] Thus, in other embodiments, the determination as to whether an execution thread is permitted to access a particular region of the (work group local) storage that it is attempting to access can be (and is) made on a per trust domain basis (with access being prevented if the execution thread is attempting to access a region of storage that has been allocated for use by a work group or work groups coming from a different trust domain).
[0039] Thus, in general, the storage is storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, wherein different groups of work items may reside in the same or in different trust domains. In that case, in embodiments, when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits are configured to (and perform as part of the method steps to): determine whether the respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or another group of work items within the same trust domain as the group of work items including that execution thread; and prevent access to the respective region of the storage when it is determined that the request is being made for an execution thread other than an execution thread of the group of work items for which the respective region of the storage has been allocated for use by or an execution thread of another group of work items within the same trust domain.
[0040] Thus, according to embodiments, the memory region “access permission checking” that is operation performed by the one or more memory region “access permission checking” circuit(s) according to the technology described herein is operable and configured such that respective regions of the (work group local) storage can only be (and are only) accessed by execution threads of respective work groups being executed by the programmable processing unit within the same trust domain as the work group for which that region of the (work group local) storage has been allocated for use as the respective work group local storage for that work group (and in some embodiments to ensure that respective regions of the (work group local) storage can only be (and are only) accessed by execution threads of the respective work group being executed by the programmable processing unit as the work group for which that region of the (work group local) storage has been allocated for use).
[0041] In this respect, as will be described further below, it will be appreciated that the access permissions in the technology described herein are (in an embodiment) defined per work group (or at least per “trust domain”) (e.g. rather than per- application), and are dynamically changed as work groups are started and finish, i.e. based on, and as part of, the temporary allocation of the (work group local) storage to a work group.
[0042] Thus, by providing such a memory region “access permission checking” circuit (or circuits), the technology described herein provides a (hardware-based) mechanism for prohibiting memory accesses to a respective region of the (work group local) storage that has been allocated for use by a particular work group for any requesters other than execution threads for the particular work group for which the respective region of the (work group local) storage has been allocated for use by. This may then enhance the security of the use of the work group local storage, for example by ensuring that any “out of bounds” memory accesses (e.g., and in particular, across different trust domains) can be (and are) prevented.
[0043] Furthermore, such memory region “access permission checking” circuit or circuits can be, and in an embodiment is, suitably located along the memory access path to the (work group local) storage such that this check can be (and in an embodiment is) performed for all memory accesses that are made to the (work group local) storage (based on the (dynamic) memory access permissions that are defined for the work groups), and this check is in an embodiment performed automatically via such memory region “access permission checking” circuit or circuit (in hardware) (e.g. rather than relying on the programmer doing this).
[0044] Thus, according to the technology described herein, when (and whenever) a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage via such a memory region “access permission checking” circuit, the memory region “access permission checking” circuit is configured to determine or not this memory access should be permitted, in particular, and in an embodiment, by determining whether the respective region of the storage that the execution thread is requesting access to has been allocated for use by the group of work items including that execution thread, or at least has been allocated for a group of work items within the same trust domain as the group of work items including that execution thread.
[0045] So long as the request is from an execution thread of the group of work items that is permitted to access the region in question (e.g. because the request is from an execution thread within the same trust domain as the group of work items for which the region of the storage has been allocated for use (or indeed is from that particular group of work items)), the memory region access can (and should) be permitted, and the memory region “access permission checking” circuit thus permits the memory access to proceed further along the memory access path to the (work group local) storage. For example, after the memory region “access permission checking” operation is complete, if the request passes the memory region “access permission checking” operation, the request is then passed further along the memory access path to the (work group local) storage, e.g., and in an embodiment, so that the memory access request can then be serviced, e.g. by performing the requested read and / or write transactions, as normal.
[0046] (It will be appreciated here that various additional checks could also be performed along the memory access path to the (work group local) storage after the memory region “access permission checking” operation is complete and so even if the request passes through the memory region “access permission checking” circuit, the memory access may still be prohibited at a different point along the memory access path to the (work group local) storage).
[0047] On the other hand, when it is determined that the request is not permitted, e.g., and in an embodiment, because the request is not being made for an execution thread of the work group for which the respective region of the storage has been allocated for use by (or at least for another work group within the same trust domain), the memory region “access permission checking” operation is failed, and the memory region “access permission checking” circuit is thus operable and configured to prevent access to the (work group local) storage. Thus, the request is prevented from accessing a respective region of the storage that has been allocated for a different work group, and the memory region “access permission checking” circuit may instead return in response to the request a suitable null response (or other response that is not (is other than) dependent upon what was already stored in the entry (region) in question).
[0048] This can then help prevent different work groups (which may reside in different trust domains) from being able to access each other’s data.
[0049] The Applicants correspondingly believe therefore that the technology described herein provides an improved arrangement and operation for graphics processors (and graphics processing systems), in particular when using “work group local storage” as shared storage for execution threads in a respective work group.
[0050] The technology described herein also extends to a graphics processor for use within a graphics processing system as described above.
[0051] Another embodiment of the technology described herein, therefore comprises a graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; the graphics processor further comprising one or more memory region access permission checking circuits that control memory accesses to storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the one or more memory region access permission checking circuits configured to: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage. Correspondingly, another embodiment of the technology described herein comprises a method of operating a graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; and one or more memory region access permission checking circuits that control memory accesses to the storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the method comprising: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and the one or more memory region access permission checking circuits preventing the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage.
[0052] As will be appreciated by those skilled in the art, these embodiments of the technology described herein can and in an embodiment do include any one or more or all of the features of the technology described herein described herein, as appropriate. Thus, the determination as to whether an execution thread is permitted to access a particular region of the storage is in an embodiment performed either on a per work group or per trust domain basis, e.g. as explained above.
[0053] The graphics processing system can be any suitable and desired graphics processing system that includes a programmable processing unit that can execute program instructions.
[0054] The programmable processing unit of the graphics processing system can be any suitable and desired programmable processing unit (“core”) that is operable to execute (shader) programs.
[0055] The programmable processing unit is in an embodiment part of a graphics processor of the graphics processing system. Thus, the system in an embodiment comprises a graphics processor comprising the programmable processing unit. The graphics processor can be any suitable and desired graphics processor that includes a programmable processing unit that can execute program instructions.
[0056] The graphics processor / system may comprise a single programmable processing unit, or may have plural such units. Where there are a plural programmable processing units, each processing unit can, and in an embodiment does, operate in the manner of the technology described herein.
[0057] Where there are plural programmable processing units, each unit may be provided as a separate circuit to other programmable processing units of the graphics processor / system, or the programmable processing units may share some or all of their circuits (circuit elements).
[0058] The (and each) programmable processing unit should, and in an embodiment does, comprise appropriate circuits (processing circuits / logic) for performing the operations required of the programmable processing unit.
[0059] Thus, the (and each) programmable processing unit will, for example, and in an embodiment does, comprise an instruction execution circuit (execution engine) operable to, and configured to, execute program instructions for execution threads. This instruction execution circuit (engine) should, and in an embodiment does, comprise a set of at least one functional unit (circuit) operable to perform data processing operations for an instruction being executed by an execution thread. An execution unit (engine) may comprise only a single functional unit, or could comprise plural functional units, depending on the operations the execution unit (engine) is to perform.
[0060] In an embodiment, the graphics processor / system and the programmable processing unit is operable to execute (shader) programs for sets (“warps”) of plural execution threads together, e.g. in lockstep, one instruction at a time.
[0061] (It will be appreciated here that a “work group” thus typically, and in an embodiment, will contain a plurality of such execution thread groups (“warps”). Thus, a “work group” in an embodiment comprises a plurality of such execution thread groups (“warps”) that in an embodiment all have access to the work group (local) storage that is allocated for that work group.)
[0062] In this case the functional units, etc., of a given execution unit are in an embodiment configured and operable so as to facilitate such thread warp arrangements. Thus, for example, the functional units are in an embodiment arranged as respective execution lanes, one for each thread that a thread warp may contain. The (programmable processing unit of the) graphics processor in an embodiment also comprises any other appropriate and desired units and circuits required for the operation of the programmable processing unit(s), such as appropriate control circuits (control logic) for controlling the execution unit(s) (engine) to cause and to perform the desired and appropriate processing operations.
[0063] Thus the (programmable processing unit of the) graphics processor / system in an embodiment also comprises an appropriate thread (warp (set)) manager (controller) circuit (a “warp manager”) that is operable to issue sets (warps) of threads to the execution unit (engine) for execution and to control the scheduling of sets (warps) of threads on / to the execution unit (engine) for execution.
[0064] The thread (warp) manager in an embodiment comprises an execution thread generator (spawner) circuit that generates (spawns) (warps of) threads for execution; and an execution thread scheduler circuit that schedules (warps of) threads for execution (this may be part of the thread generator).
[0065] The graphics processor / system in an embodiment also comprises and / or has access to appropriate (local) storage, such as registers / register file, caches, etc., and an appropriate interface to, and communication with memory (a memory system) of or accessible to the graphics processor / system (e.g., and in an embodiment, via an appropriate cache hierarchy), together with appropriate load / store units and communication paths for transferring data between the local storage and memory system of or accessible to the graphics processor / system.
[0066] The memory and memory system is in an embodiment a main memory of or available to the graphics processor / system, such as a memory that is dedicated to the graphics processor / system, or a main memory of a data processing system that the graphics processor / system is part of.
[0067] The storage that is used to (temporarily) store data for execution threads for work groups in the technology described herein (the storage that is used for “work group local storage”) can be any suitable and desired storage of or available to execution threads when executing on the programmable processing unit (of the graphics processor / system).
[0068] This storage is intended to be storage (memory) via which execution threads for respective work groups can communicate values to other threads in the work group, but which is to be allocated for use by threads of a work group temporarily (and so will be reused from one work group and / or process to another).
[0069] It is different therefore, for example, to storage (e.g. system memory) that is to be used to “transfer” data values between different work groups, processes, and / or components of the overall data processing system. Rather, it is intended to be, and is, local and temporary storage (memory) that is used for work groups and that can be allocated for use by work groups, but which then will be de-allocated from a work group in question (and available for use by another work group) once a work group has terminated.
[0070] Thus the storage that is used for “work group local storage” in the technology described herein is intended to be, and to act and to be used as, a “scratch pad memory”, for use by work groups executing on the programmable processing unit.
[0071] The allocation of storage regions within the “work group local storage” to respective work groups being executed by the programmable processing unit can be performed in any suitable and desired manner. In an embodiment, this is done in the normal manner for the allocation of work group local storage for the graphics processor and graphics processing system in question.
[0072] In an embodiment, respective regions of the storage may be (and are) allocated to plural different work groups at the same time, i.e. such that there will be plural different “work group local storage” regions allocated to different work groups in the storage at the same time (and these plural different work groups may reside in different trust domains, e.g. where the different work groups belong to different applications / processes and / or virtual machines, etc. for which graphics processing is being performed at the same time). It could also, e.g., be the case that any given work group’s work group local storage in the storage comprises plural distinct (separate) regions in the storage, if desired.
[0073] The storage that is used for the “work group local storage” may comprise storage that is dedicated to and specifically set aside for the purposes of providing “work group local storage” (and in one embodiment that is the case), or it may comprise storage that as well as being intended to be used as “work group local storage”, is also available to be used for other purposes, such as by other processes that the graphics processor / system may perform (and in another embodiment this is the case).
[0074] In the latter case, the storage that is to provide the “work group local storage” will accordingly have plural “requesters” (“masters”) able to access it, comprising at least execution threads and the corresponding processes and circuits that are using the “shared” storage for “work group local” storage, and other processes and circuits that are using the storage for other purposes.
[0075] The storage that provides the work group local storage may be configured as desired, e.g. as a single bank of storage. In an embodiment, the storage is configured as multiple, independently accessible (memory) banks (e.g. to allow more than one memory access per clock cycle). In the case where storage (whether dedicated to work group local storage or otherwise) is configured across multiple (memory) banks, then in an embodiment, each storage (memory) bank has its own accessor circuit that controls memory accesses for execution threads of work groups that are currently being executed by the programmable processing unit, with each accessor circuit (in an embodiment) operable and configured to operate independently for its respective storage bank (once started). For instance, each accessor circuit may maintain a respective queue of memory accesses (transactions) to be serviced for execution threads of work groups that are currently being executed by the programmable processing unit relating to its respective storage (memory) bank.
[0076] The storage that provides the work group local storage can be provided as desired in the graphics processing system.
[0077] In an embodiment, it is storage that is local to (on-chip with) the graphics processor, and in an embodiment local to and on-chip with the programmable processing unit of the graphics processor / system. Thus, the storage is in an embodiment storage that provides a faster, more efficient and higher bandwidth path (from the instruction execution circuit of) the programmable processing unit, than, for example, memory of the (main) memory system or available to the graphics processor / system.
[0078] Accordingly, the storage in an embodiment does not form (is not) part of a cache hierarchy or (main) memory (system), and in an embodiment does not require communication over any communication buses external to the graphics processor and / or programmable processing unit to be accessed by / for a work group executing on the programmable processing unit.
[0079] Correspondingly, in an embodiment, the system comprises a graphics processor that comprises both the programmable processing unit and the storage that is used for (that provides) the “work group local storage”. Thus, the memory region “access permission checking” circuit (or circuits, wherein plural such circuits are provided) is in an embodiment also provided local to and on-chip with the programmable processing unit of the graphics processor / system, as part of the same graphics processor.
[0080] Thus, a further embodiment of the technology described herein comprises a graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the graphics processor further comprising one or more memory region access permission checking circuits that control memory accesses to the storage, the one or more memory region access permission checking circuits configured to: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage.
[0081] Correspondingly, a yet further embodiment of the technology described herein comprises a method of operating a graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed; and one or more memory region access permission checking circuits that control memory accesses to the storage, the method comprising: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits configured to: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage (that the execution thread is requesting access to) when it is determined that the execution thread is other than (i.e. is not) permitted to access that region of the storage .
[0082] As will be appreciated by those skilled in the art, these embodiments of the technology described herein can and in an embodiment do include any one or more or all of the features of the technology described herein described herein, as appropriate.
[0083] For instance, as discussed above, in the technology described herein, a memory region access “permission checking” operation is performed such that when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: it is determined whether the execution thread is permitted to access that region of storage, and this determination is in an embodiment made by determining whether the respective region of the storage that the execution thread is requesting access to has been allocated for use by the group of work items including that execution thread (with access to the respective region of the storage then being controlled based on such determination).
[0084] Correspondingly, the graphics processing system (graphics processor) includes one or more memory region “access permission” checking circuits, with each memory region “access permission” checking circuit being operable and configured to perform such memory region access permission checking operations on respective memory accesses that proceed via that memory region “access permission” checking circuit.
[0085] Subject to the particular requirements of the technology described herein, the memory region access “permission checking” operation that is performed by the memory region “access permission” checking circuit (or circuits) according to the technology described herein may comprise any suitable and desired memory region access “permission checking” operation (and the memory region “access permission” checking circuit (or circuits) may correspondingly be suitably configured to perform any such memory region access “permission checking” operation, as desired).
[0086] For example, in embodiments, the memory region access “permission checking” operation involves a suitable “bounds check”. The “bounds” that are checked should, and in an embodiment do, come from the graphics processor / system (hardware) (and conversely should not come from an untrusted source such as the instruction stream). Subject to the particular requirements of the technology described herein, such (hardware-based) memory region bounds check could however be implemented in any suitable and desired manner. Two main examples are contemplated.
[0087] In a first main example, a “address range”-based bounds check is performed, in which when a region of the (work group local) storage is allocated for use by a particular work group, or for a set of work groups coming from the same trust domain, an indication of the address range of the region of the (work group local) storage that has been allocated for that work group / trust domain is stored in such a manner that it can be (and is) associated with all execution threads generated for that work group / trust domain.
[0088] This indication may and in an embodiment does comprise an indication of a base (start) address and size of the region of the storage that has been allocated, but could also for example comprise an indication of the base (start) address and the top (end) address defining the region. In some cases, e.g. if non-contiguous regions of the storage are allocated, the indications may comprise a set of indicators defining the base (start) and size and / or top (end) address defining each non-contiguous region. Various arrangements would be possible in this regard.
[0089] Thus, according to this first main example, the “address range”-based bounds check may be performed either on a per work group basis or on a per trust domain basis. In the case where the bounds check is to be performed on a per work group basis, each work group for which (work group local) storage has been allocated will thus be associated with a respective indication of the address range of the region of the (work group local) storage that it has been allocated. Correspondingly, in the case where the bounds check is to be performed on a per trust domain basis, each work group within a particular trust domain for which (work group local) storage has been allocated will thus be associated with a respective indication of the address range of the region of the (work group local) storage that it has been allocated for any (and all) work groups coming from that (same) trust domain.
[0090] The memory region access “permission checking” operation that is performed for an incoming memory access for an execution thread of a particular work group may then according to this first main example comprise determining whether the memory access is requesting access to an address that is inside or outside of the address range of the region of the (work group local) storage that has been allocated for use by the particular work group / trust domain including the execution thread for which the memory access is being requested. If the incoming memory access (request) is requesting access to an address that is inside of the address range of the region of the (work group local) storage that has been allocated to the particular work group / trust domain including the execution thread for which the memory access is being requested, the memory access can then proceed. Whereas, if memory access is requesting access to an address that is outside of the address range of the region of the (work group local) storage that has been allocated to the particular work group / trust domain including the execution thread for which the memory access is being requested, the memory region access “permission checking” operation is failed, and a suitable response should be returned in this event.
[0091] Thus, in embodiments, the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use by the group of work items including that execution thread comprises comparing a memory address to which the request is being made with a range of memory addresses identifying the respective region(s) of the storage that has been allocated for use by the group of work items, or in some embodiments the respective region(s) of the storage that has been allocated for use by any groups of work items coming from the same trust domain, for which the request is being made. If the requested address for an execution thread of a work group is outside of the address range of the region of the (work group local) storage that has been allocated for that work group / trust domain, the request is “out of bounds”, and should therefore be (and is) prevented.
[0092] Various other examples would however be possible.
[0093] For instance, in a second main example, an “identifier”-based bounds check is performed, in which when a region of the (work group local) storage is allocated for use by a particular work group, that region is associated with a respective identifier associated with the work group for which it has been allocated.
[0094] For instance, once a particular region of the (work group local) storage has been allocated for use by a particular work group, the respective work group identifier for that work group is then stored appropriately in association with that region of the (work group local) storage.
[0095] Any incoming memory access coming from execution threads of work groups will also have a corresponding identifier identifying the work group to which they relate, and these identifiers can thus be compared with the identifiers stored for the respective regions of the (work group local) storage to see whether the memory access request should be permitted (or not). The identifiers could be stored on a per-work group basis, in which each different work group can be (and is) associated with a respective work group identifier, and the comparison / checking is performed on this basis. Rather than using per-work group identifiers it would also be possible to perform a per-trust domain identity check, and this is done in embodiments. For instance, the Applicants recognise that it may only be necessary to prevent access between work groups when the work groups reside in different trust domains. This check can then work as described above but with the checks being performed using the respective per-trust domain identifiers (instead of the per-work group identifiers).
[0096] Thus, in embodiments, the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use by the group of work items including that execution thread comprises comparing an identifier associated with the group of work items for which the request is being made with a corresponding identifier stored in associated with the respective region of the storage identifying the group of work items for which the respective region of the storage has been allocated. As mentioned above, the identifier may be a per- work group identifier or may be a per-trust domain identifier. In either case, when the compared identifiers do not match, this means that the request is “out of bounds”, and should therefore be (and is) prevented.
[0097] Various other examples would be possible for implementing the memory region access “permission checking” operation.
[0098] As mentioned above, a memory access request is only permitted to proceed further along the memory access path to the (work group local) storage if the request passes such memory region access “permission checking” operation. Conversely, if the memory region access “permission checking” operation, in whatever form that check takes, is failed, the memory region access “permission checking” circuit prevents the memory access request to proceed further along the memory access path to the (work group local) storage (i.e. such that “out of bounds” access to the respective region of the storage is prevented).
[0099] In the event that the memory region access “permission checking” operation is failed, the memory region access “permission checking” circuit should then, and in an embodiment does, return a suitable response. For example, in some embodiments, the memory region access “permission checking” circuit may report that there is an “out of bounds” error, and return a suitable error message indicative of this attempted illegal memory access, e g. to the host processor. This could then cause the processing to be terminated. In other embodiments, therefore, rather than returning an error message, the memory region access “permission checking” circuit may be configured to return another value (e.g. a suitable random or default value) so that the processing can continue (but the “out of bounds” memory access does not return any meaningful values) (in this case, an “out of bounds” error may still be flagged, e.g. for diagnostic processes).
[0100] Various options would be possible in this regard.
[0101] Subject to the particular requirements of the technology described herein the memory region access “permission checking” operation may be performed at any suitable and desired locations within the graphics processor / system.
[0102] In an embodiment, a memory region access “permission checking” operation is performed whenever there is a memory access to the (work group local) storage.
[0103] In an embodiment this is done whether the memory access is for an execution thread of a work group that is currently being executed by the programmable execution unit or for another requester (e.g. in the case that the storage that is used for the “work group local storage” comprises storage that as well as being intended to be used as “work group local storage” is also available to be used for other purposes, such as by other processes that the graphics processor / system may perform (as is the case in some embodiments)).
[0104] Correspondingly, the (or a) memory region access “permission checking” circuit should be provided for each possible memory access path to the (work group local) storage (such that a (in an embodiment hardware-based) memory region access “permission checking” operation is performed whenever there is a memory access to the (work group local) storage, and this check cannot be bypassed or avoided).
[0105] In this respect, it will be appreciated that the graphics processor / system may comprise a single memory region access “permission checking” circuit that is operable and configured to control all possible accesses to the (work group local) storage or there may be a plurality of separate memory region access “permission checking” circuits, each associated with a respective separate memory access path to the (work group local) storage. Various arrangements would be possible in this regard.
[0106] The memory region access “permission checking” circuit (or circuits) that controls memory access to the (work group local) storage and that is operable to perform the memory region “permitted access” check operations according to the technology described herein can be any suitable and desired circuit that can be configured to perform and perform such an operation, and can be arranged in any suitable and desired configuration and location relative to the storage. It should be, and in an embodiment, is located (at least logically) intermediate the programmable processing unit and the (work group local) storage. It is in an embodiment located at an appropriate point along the “access” path between the (instruction execution circuit (execution engine) of the) programmable processing unit and the (work group local) storage.
[0107] Various arrangements are however contemplated in this regard.
[0108] For example, in some embodiments, the or a memory region access “permission checking” circuit is associated with, and in an embodiment part of, the programmable processing unit and is therefore in an embodiment operable and configured to perform the memory region “permitted access” check operations on memory access requests for execution threads of work groups being executed by the programmable processing unit before the memory access requests are passed to the storage.
[0109] In an embodiment, the processing circuit is part of, and / or comprises, a circuit that controls access to the (local) storage (i.e. such that any access requests from the (instruction execution circuit of the) programmable processing unit to the (local) storage will pass through and be controlled by the (local) storage access circuit, of which the processing circuit is (at least) part.
[0110] This then means that if a memory access request fails such a memory region “permitted access” check operation, a suitable result (e.g. an error message, as described above) can be returned without having to pass the memory access on to the overall storage (memory) unit that provides the work group local storage.
[0111] However, it would also be possible for the or a memory region access “permission checking” circuit to be associated with, and in an embodiment part of, the overall storage (memory) unit that provides the work group local storage. For instance, in an embodiment, the storage, as well as comprising appropriate storage elements (e.g. memory banks), also comprises such processing circuit.
[0112] Various examples would be possible in this regard.
[0113] For example, in an embodiment, wherein the storage is configured as multiple, independently accessible banks, each bank of storage having a respective (independent) accessor circuit, a respective (separate) memory region access “permission checking” circuit may be associated with, and in an embodiment part of, each of the multiple, independently accessible banks (for example, by being associated with, and in an embodiment part of, the respective accessor circuits for the respective banks of storage).
[0114] As mentioned above, in some embodiments, the storage that is used for the “work group local storage” comprises storage that as well as being intended to be used as “work group local storage” is also available to be used for other purposes, such as by other processes that the graphics processor / system may perform (as is the case in some embodiments)), and in that case, the or a memory region access “permission checking” circuit should (also) be provided that is operable and configured to perform memory region access “permission checking” operations for requests coming from such “other” requesters.
[0115] For instance, in the case where the (work group local) storage is also shared with another unit or units (or process or processes) that is able to use the storage (in addition to the storage being used for work group local storage for work groups being executed by the programmable processing unit), then in an embodiment, the system (e.g. graphics processor / system) includes an appropriate) arbitration circuit(s) and process(es) to arbitrate between accesses to the storage that relate to its use as work group local storage (e.g. that proceed via the respective accessor circuit(s)) and access requests coming from other “requesters” to the storage. In the case where the storage is comprised of multiple banks, this arbitration is in an embodiment done and provided on a bank-by-bank basis.
[0116] (The arbiters (arbitration process) can be configured to operate as desired, for example to always prioritise memory access requests from “another” requester, or to always prioritise requests relating to work groups, as desired. There could equally be multiple other requesters able to access the storage, with appropriate arbitration (priority) policies in place accordingly.)
[0117] In some embodiments, therefore, a respective (separate) memory region access “permission checking” circuit may be associated with, and in an embodiment part of, the respective arbitration circuit(s) that perform the arbitration for the respective banks of storage. This then means that the same memory region access “permission checking” circuit can perform memory region access “permission checking” operations for memory access requests coming both from execution threads of work groups that are being executed by the programmable processing unit and from any “other” requesters, with the memory region access “permission checking” being performed within the (work group local) storage, thus saving having to provide separate circuits to do this.
[0118] Other arrangements would however be possible. For instance, in embodiments, a separate and dedicated memory region access “permission checking” circuit may be provided for the (or each) “other” requester that is located at a suitable point along the “access” path between the “other” requester and the (work group local) storage. For example, it could be associated with, and in an embodiment located within, the “other” requester. In that case, the memory region access “permission checking” circuit (or circuits) that perform the memory region access “permission checking” for requests for execution threads of work groups are provided separately to the circuit that controls access for the “other requesters”, and may be provided in any suitable manner, e.g. as described above.
[0119] Various arrangements would be possible in this regard.
[0120] The technology described herein may generally find application in any suitable graphics processing system.
[0121] The graphics processing system may further include a host processor that executes applications that can require data or graphics processing by the graphics processor and that instruct the graphics processor accordingly (e.g. via a driver for the graphics processor). The system may further include appropriate storage (e.g. memory), caches, etc..
[0122] The graphics processing system and / or graphics processor may also comprise, and / or be in communication with, one or more memories and / or memory devices that store data, and / or that store software for performing the processes described herein. The graphics processing system and / or graphics processor may also be in communication with a host microprocessor, and / or with a display for displaying images based on the data generated.
[0123] The graphics processor may include (implement) any one or more or all of the processing stages that a graphics processor (processing pipeline) can normally include. Thus, for example, the graphics processor may include a primitive setup stage, a rasteriser, and / or Tenderer (in an embodiment in the form of a fragment shader).
[0124] The graphics processor (processing pipeline) may comprise one or more other programmable shading stages, such as one or more or all of, a vertex shading stage, a hull shader, a tessellation stage (e.g. where tessellation is performed by executing a shader program), a domain (evaluation) shading stage (shader), and a geometry shading stage (shader), as well as a fragment shader.
[0125] The graphics processor (processing pipeline) may also contain any other suitable and desired processing stages that a graphics processing pipeline may contain such as a depth (or depth and stencil) tester(s), a blender, a tile buffer or buffers, a write out unit etc..
[0126] The technology described herein can be used in and with any suitable and desired graphics processing system and processor. In one embodiment, the graphics processor (processing pipeline) is a tiled-based graphics processor (processing pipeline).
[0127] The technology described herein can be used for any form of output that a graphics processor may be used to generate. In one embodiment it is used when a graphics processor is being used to generate images for display, but it can be used for any other form of graphics processing output, such as (e.g. post-processed) graphics textures in a render-to-texture operation, etc., that a graphics processor may produce, as desired. It can also be used when a graphics processor is being used to generate other (e g. non-image or non-graphics) outputs.
[0128] In one embodiment, the various functions of the technology described herein are carried out on a single data or graphics processing platform that generates and outputs the required data, such as processed image data that is, e.g., written to a frame buffer for a display device.
[0129] The technology described herein can be implemented in any suitable system, such as a suitably operable micro-processor based system. In some embodiments, the technology described herein is implemented in a computer and / or microprocessor based system.
[0130] The various functions of the technology described herein can be carried out in any desired and suitable manner. For example, the functions of the technology described herein can be implemented in hardware or software, as desired. Thus, for example, the various functional elements, stages, units, and "means" of the technology described herein may comprise a suitable processor or processors, controller or controllers, functional units, circuitry, circuits, processing logic, microprocessor arrangements, etc., that are operable to perform the various functions, etc., such as appropriately dedicated hardware elements (processing circuits / circuitry) and / or programmable hardware elements (processing circuits / circuitry) that can be programmed to operate in the desired manner.
[0131] It should also be noted here that the various functions, etc., of the technology described herein may be duplicated and / or carried out in parallel on a given processor. Equally, the various processing stages may share processing circuits / circuitry, etc., if desired.
[0132] Furthermore, any one or more or all of the processing stages or units of the technology described herein may be embodied as processing stage or unit circuits / circuitry, e.g., in the form of one or more fixed-function units (hardware) (processing circuits / circuitry), and / or in the form of programmable processing circuits / circuitry that can be programmed to perform the desired operation. Equally, any one or more of the processing stages or units and processing stage or unit circuits / circuitry of the technology described herein may be provided as a separate circuit element to any one or more of the other processing stages or units or processing stage or unit circuits / circuitry, and / or any one or more or all of the processing stages or units and processing stage or unit circuits / circuitry may be at least partially formed of shared processing ci rcuit / circuitry .
[0133] It will also be appreciated by those skilled in the art that all of the described embodiments of the technology described herein can include, as appropriate, any one or more or all of the features described herein.
[0134] The methods in accordance with the technology described herein may be implemented at least partially using software, e.g., computer programs. Thus, further embodiments of the technology described herein comprise computer software specifically adapted to carry out the methods herein described when installed on a data processor, a computer program element comprising computer software code portions for performing the methods herein described when the program element is run on a data processor, and a computer program adapted to perform all the steps of a method or of the methods herein described when the program is run on a data processing system. The data processing system may be a microprocessor, a programmable FPGA (Field Programmable Gate Array), etc.
[0135] The technology described herein also extends to a computer software carrier comprising such software which when used to operate a graphics processor, renderer or other system comprising a data processor causes in conjunction with said data processor said processor, Tenderer or system to carry out the steps of the methods of the technology described herein. Such a computer software carrier could be a physical storage medium such as a ROM chip, CD ROM, RAM, flash memory, or disk, or could be a signal such as an electronic signal over wires, an optical signal or a radio signal such as to a satellite or the like.
[0136] It will further be appreciated that not all steps of the methods of the technology described herein need be carried out by computer software and thus further embodiments of the technology described herein comprise computer software and such software installed on a computer software carrier for carrying out at least one of the steps of the methods set out herein.
[0137] The technology described herein may accordingly suitably be embodied as a computer program product for use with a computer system. Such an implementation may comprise a series of computer readable instructions fixed on a tangible, non- transitory medium, such as a computer readable medium, for example, diskette, CD ROM, ROM, RAM, flash memory, or hard disk. It could also comprise a series of computer readable instructions transmittable to a computer system, via a modem or other interface device, over a tangible medium, including but not limited to optical or analogue communications lines, or intangibly using wireless techniques, including but not limited to microwave, infrared or other transmission techniques. The series of computer readable instructions embodies all or part of the functionality previously described herein.
[0138] Those skilled in the art will appreciate that such computer readable instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Further, such instructions may be stored using any memory technology, present or future, including but not limited to, semiconductor, magnetic, or optical, or transmitted using any communications technology, present or future, including but not limited to optical, infrared, or microwave. It is contemplated that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation, for example, shrink wrapped software, pre-loaded with a computer system, for example, on a system ROM or fixed disk, or distributed from a server or electronic bulletin board over a network, for example, the Internet or World Wide Web.
[0139] Figure 1 shows an exemplary system on-chip (SoC) graphics processing system 8 that comprises a host processor in the form of a central processing unit (CPU) 1 , a graphics processor (GPU) 2, a display processor 3 and a memory controller 5.
[0140] As shown in Figure 1, these units communicate via an interconnect 4 and have access to off-chip memory 6. In this system, the graphics processor 2 will render frames (images) to be displayed, and the display processor 3 will then provide the frames to a display panel 7 for display.
[0141] In use of this system, an application 13 such as a game, executing on the host processor (CPU) 1 will, for example, require the display of frames on the display panel 7. To do this, the application will submit appropriate commands and data to a driver 11 for the graphics processor 2 that is executing on the CPU 1 . The driver 11 will then generate appropriate commands and data to cause the graphics processor 2 to render appropriate frames for display and to store those frames in appropriate frame buffers, e.g. in the main memory 6. The display processor 3 will then read those frames into a buffer for the display from where they are then read out and displayed on the display panel 7 of the display.
[0142] Figure 2 shows schematically the relevant elements and components of a graphics processor (GPU) 60 of the present embodiments.
[0143] As shown in Figure 2, the GPU 60 includes one or more programmable processing units (shader (processing) cores) 61 , 62 together with a memory management unit 63 and a level 2 cache 64 which is operable to communicate with an off-chip memory system 68 (e.g. via an appropriate interconnect and (dynamic) memory controller).
[0144] Figure 2 shows schematically the relevant configuration of one shader core 61 , but as will be appreciated by those skilled in the art, any further shader cores of the graphics processor 60 will be configured in a corresponding manner.
[0145] (The graphics processor (GPU) shader cores 61, 62 are programmable processing units (circuits) that perform processing operations by running small programs for each “item” in an output to be generated such as a render target, e.g. frame. An “item” in this regard may be, e.g. a vertex, one or more sampling positions, a compute shader “work item”, etc.. The shader cores will process each “item” by means of one or more execution threads which will execute the instructions of the shader program(s) in question for the “item” in question. Typically, there will be multiple execution threads each executing at the same time (in parallel).)
[0146] Figure 2 shows the main elements of the graphics processor 60 that are relevant to the operation of the present embodiments. As will be appreciated by those skilled in the art there may be other elements of the graphics processor 60 that are not illustrated in Figure 2. It should also be noted here that Figure 2 is only schematic, and that, for example, in practice the shown functional units may share significant hardware circuits, even though they are shown schematically as separate units in Figure 2. It will also be appreciated that each of the elements and units, etc., of the graphics processor as shown in Figure 2 may, unless otherwise indicated, be implemented as desired and will accordingly comprise, e.g., appropriate circuits (processing logic), etc., for performing the necessary operation and functions.
[0147] As shown in Figure 2, each shader core of the graphics processor 60 includes an appropriate instruction execution unit (execution engine) 65 that is operable to execute shader programs for execution threads to perform processing operations.
[0148] The shader core 61 also includes an instruction cache 66 that stores instructions to be executed by the instruction execution unit 65 to perform processing operations. The instructions to be executed will, be fetched from the memory system 68 via an interconnect 69 and a micro TLB (translation lookaside buffer) 70.
[0149] The shader core 61 also includes an appropriate load / store unit 76 in communication with the instruction execution unit 65, that is operable, e.g., to load into an appropriate cache, data, etc., to be processed by the instruction execution unit 65, and to write data back to the memory system 68 (for data loads and stores for programs executed in the instruction execution unit). Again, such data will be fetched / stored by the load / store unit 76 via the interconnect 69 and the micro TLB 70.
[0150] In order to perform graphics processing operations, the instruction execution unit 65 will execute graphics shader programs (sequences of instructions) for respective execution threads.
[0151] Accordingly, as shown in Figure 2, the shader core 61 further comprises a warp manager 72 operable to generate execution threads for execution by the instruction execution unit 65, and to issue such threads to the instruction execution unit 65 and to control the scheduling of threads on / to the instruction unit 65, for execution.
[0152] The present embodiments are particularly concerned with the operation of the graphics processor (and in particular the shader cores of the graphics processor) when processing so called “work groups” (e.g. when performing compute shading), i.e. collections of execution threads (corresponding to respective work items) that are handled and treated as a “group” as a whole, and that are all accordingly guaranteed to exist at the same time (have the same lifetime) and are able to perform communication and synchronisation with each other. Accordingly, the warp manager 72 is correspondingly operable to generate respective work groups of execution threads, and issue and schedule such work groups of threads on and to the instruction execution unit 65.
[0153] To facilitate such work group operation, and in particular in order to facilitate data sharing between the threads within a work group, as shown in Figure 2, the shader core 61 also includes a shared memory unit (SMU) 74, which is in communication with the instruction execution unit (execution engine) 65 and the warp manager 72.
[0154] This shared memory unit 74 is operable to provide “work group local storage” for execution threads in respective work groups, that the execution threads can read from and write to whilst the work group is being executed (while the work group is in existence), in order to allow data sharing between the threads within the work group.
[0155] In particular, respective work groups can be allocated respective memory regions within the shared memory unit for use as “work group local storage” whilst the work group is executing. When a work group reaches the end of its lifetime, the work group’s “local storage” memory region in the shared memory unit 74 is deallocated, making it available for use by another work group.
[0156] In the present embodiments, as shown in Figure 2, the shared memory unit that provides the work group local storage is local to and on-chip with the shader core 61 . Other arrangements for this would, of course, be possible. (The allocation of regions of the shared memory unit 74 to respective work groups can be performed in any suitable and desired manner, for example in the normal manner for the graphics processor and graphics processing system in question.)
[0157] Figure 3 shows the work group local storage shared memory unit 74 in more detail.
[0158] As shown in Figure 3, it is assumed that the work group local storage shared memory unit 74 comprises a plurality of memory (SRAM) banks 31. Figure 3 shows four memory banks 31 , but other numbers and arrangements of memory banks would, of course, be possible.
[0159] As shown in Figure 3 (and as discussed above in relation to Figure 2), the work group local storage shared memory unit 74 is operable to receive commands from the warp manager 72, and also commands, and read and write operations, and atomic operations from the execution engine 65.
[0160] As shown in Figure 3, the work group local storage shared memory unit 74 also includes a set of memory bank access control units (circuits) 32, one for each memory bank 31. As will be discussed in more detail below, these accessor units (circuits) 32 control accesses to the memory banks 31.
[0161] As shown in Figure 3, the execution engine 65 includes, in accordance with the technology described herein, a memory region “access permission checking” circuit 65A. As discussed above, this circuit is operable and configured to perform a suitable “bounds check” for any memory accesses to the work group local storage shared memory unit 74 for execution threads of work groups that are being executed by the execution engine 65
[0162] The “bounds check” that the memory region “access permission checking” circuit 65A is configured to perform may be performed in any suitable and desired manner.
[0163] For instance, according to one main example (corresponding to the “second main example” referred to above), a “trust domain” identifier-based “bounds check” is performed. This will be described in more detail in relation to Figure 4, Figure 5 and Figure 6.
[0164] Figure 4 shows the operation of the warp manager 72 when a new compute run command is received to process a new work group.
[0165] As shown in Figure 4, when the new compute run command is received (step 401), the warp manager 72 then allocates a respective region of the work group local storage shared memory unit 74 for use by the work group to which the compute run command relates (step 402). This allocation can be done in any suitable and desired manner, e.g. in the normal manner for such allocation.
[0166] The allocated region of the work group local storage shared memory unit 74 should then be, and in an embodiment is, cleared at this point (step 403), to ensure that any previous work group’s data is removed. This clear operation may be done in various suitable ways, as desired. For example, in some embodiments, a dedicated clear operation may be automatically triggered by the step of allocating a respective region of the work group local storage shared memory unit 74 for use by the new work group (i.e. by step 402), e.g. in response to a suitable “clear” command issued by the warp manager 72. However, various other examples would be possible.
[0167] Then, a trust domain identifier identifying the trust domain for the new work group is assigned to all regions of the work group local storage shared memory unit 74 that have been allocated for use by the work group (step 404). These trust domain identifiers are then suitably stored in association with the respective regions of the work group local storage shared memory unit 74 (e.g. as shown in Figure 6).
[0168] Figure 5 then shows the corresponding “bounds check” operation, i.e. that will in this embodiment be performed by the memory region “access permission checking” circuit 65A for incoming memory accesses.
[0169] As shown in Figure 5, when an execution thread of a work group being processed by the execution engine 65 requests access to the work group local storage shared memory unit 74 (step 501), it is checked whether the trust domain identifier for the memory access matches the trust domain identifier that is stored for the region of the work group local storage shared memory unit 74 to which access is being requested (step 502).
[0170] If the trust domain identifiers do not match (step 502 - No), the memory region “access permission checking” circuit 65A then reports an “out of bounds” error (step 503) and the memory access request is prevented from reaching the work group local storage shared memory unit 74.
[0171] On the other hand, if the trust domain identifiers match (step 502 - Yes), the memory region “access permission checking” circuit 65A then permits the memory access request to proceed further to the work group local storage shared memory unit 74 (step 504).
[0172] Figure 6 then shows the arrangement of the work group local storage shared memory unit 74 in this example. As shown in Figure 6, each SRAM bank 32 includes a plurality of entries (locations) that can be (and in this example have been) allocated for work groups in different trust domains. Stored in association with the respective entries (locations) are respective trust domain identifiers, that can be used as described above to perform the “bounds check”.
[0173] This example has the benefit of being able to allocate non-contiguous chunks of memory to a work group. It does need some per-allocation-block storage to do so though.
[0174] The “bounds check” may be performed in other ways, as desired.
[0175] For example, in another main example (corresponding to the “first main example” referred to above), a range-based bounds check is performed. This will be described in more detail in relation to Figure 7, Figure 8 and Figure 9.
[0176] Figure 7 shows the corresponding operation of the warp manager 72 according to this example.
[0177] As shown in Figure 7, when a new compute run command is received (step 701), the warp manager 72 may then spawn one or more, and typically a plurality of, work groups that are to be executed for the compute run command (and all of these work groups are thus known to come from the same trust domain).
[0178] The warp manager 72 then allocates respective regions of the work group local storage shared memory unit 74 for use by each of the work groups spawned by the compute run command relates (step 702), and the allocated regions of the work group local storage shared memory unit 74 should then be, and in an embodiment is, cleared at this point (step 703), as described above.
[0179] At this point, the allocation base address and size of the full region of local storage that has been allocated for all of the work groups spawned by the run compute command (and that hence come from the same trust domain) are then stored in per-trust domain state associated with all execution threads (warps) for the work group associated with the run command (step 704).
[0180] Figure 8 then shows the corresponding “bounds check” operation, i.e. that will in this embodiment be performed by the memory region “access permission checking” circuit 65A for incoming memory accesses.
[0181] As shown in Figure 8, when an execution thread of a work group being processed by the execution engine 65 requests access to the work group local storage shared memory unit 74 (step 801), it is checked whether the memory access is attempting to access memory outside of the allowed address range as determined by the base address and size stored in the per-trust domain state of the execution thread (warp) that is initiating the memory access (step 802).
[0182] If the request is outside the allowed address range (step 802 - Yes), the memory region “access permission checking” circuit 65A then reports an “out of bounds” error (step 803) and the memory access request is prevented from reaching the work group local storage shared memory unit 74.
[0183] On the other hand, so long as the request is inside the allowed address range (step 802 - No), the memory region “access permission checking” circuit 65A then permits the memory access request to proceed further to the work group local storage shared memory unit 74 (step 804).
[0184] Figure 9 then shows the arrangement of the work group local storage shared memory unit 74 in this second main example. As shown in Figure 9, each SRAM bank 32 includes a plurality of entries (locations) that can be (and in this example have been) allocated for work groups in different trust domains. A separate allocation table 901 is provided that stores the respective base address and size values for the different trust domains, that can be used as described above to perform the “bounds check”.
[0185] This example has the advantage (compared to the example described above at least) of avoiding the per-allocation-block storage, as instead storage is only needed per-trust domain. This then requires fewer storage bits but at the cost of not being able to allocate non-contiguous chunks of memory.
[0186] Various other examples would be possible.
[0187] For instance, whilst the examples described above in relation to Figures 4 to 6 and Figures 7 to 9 perform the memory region “access permission checking” on a per trust domain basis, it will be appreciated that this could also be done on a per work group basis (i.e. by storing suitable “work group” identifiers or storing an address range per work group) (and in that case a given work group would only be permitted to access the particular region of the storage that has been allocated for its use).
[0188] In Figure 3, as described above, a memory region “access permission checking” circuit 65A is provided within the execution engine 65. However, such a memory region “access permission checking” circuit may be provided at any suitable location within the graphics processor, so long as it is able to perform a suitable “bounds check” as required for any accesses to the work group local storage shared memory unit 74 (and so long as the bounds that are being checked against do not come from an un-trusted source such as the instruction stream).
[0189] For example, Figure 10 shows another example in which a memory region “access permission checking” circuit 74A is provided within the work group local storage shared memory unit 74.
[0190] Figure 11 shows another example in which each of the respective bank access control units (circuits) 32 includes a respective memory region “access permission checking” circuit 32A, so that a respective memory region “access permission checking” circuit 32A is provided for each memory bank 31.
[0191] Figure 3, Figure 10 and Figure 11 described above show various arrangement in which the work group local storage shared memory unit 74 is exclusively used for and available for work group local storage for work groups executing on the execution engine 65.
[0192] Figure 12 shows an alternative arrangement and embodiment, in which the work group local storage shared memory unit 90 is shared with another unit (requester) 91 that is able to make accesses to the memory banks of the shared memory unit 90.
[0193] In this case, accordingly, the shared memory unit 90 also includes a set of arbiters (arbitration circuits) 92 that are operable to arbitrate between memory access requests coming from the other requester 91 and from the work group local storage accessor units 32.
[0194] In this arrangement, the arbiters 92 can be configured to operate as desired, for example to always prioritise memory access requests from the “other” requester” 91 , or to always prioritise requests relating to work groups, as desired. There could equally be multiple other requesters able to access the memory, with appropriate arbitration (priority) policies in place accordingly.
[0195] In this arrangement, in the case where a higher priority request from another requester 91 is received, the memory access relating to a work group will be appropriately stalled until the other request has been serviced, and may then be resumed (and vice-versa).
[0196] In this case, a “bounds check” should also be performed on memory accesses from the other requester 91, to determine whether they are being made to a region of the shared memory unit of a memory bank that has already been allocated to a particular work group (and to prevent any accesses from another requester that is to a region of a memory bank that has already been allocated to a work group), if desired. This would then avoid other memory accesses from other requesters being able to interfere with the work group local storage region.
[0197] Such memory region bounds checks could be implemented in any suitable and desired manner, for example in the manner described above.
[0198] For example, as shown in Figure 12, each of the arbiters (arbitration circuits) 92 may include a respective memory region “access permission checking” circuit 92A, with the respective memory region “access permission checking” circuits 92A thus being operable to perform “bounds checks” both for requests coming from the other requester 91 and from the work group local storage accessor units 32. Other arrangements would however be possible. For example, as shown in Figure 13, Figure 14 and Figure 15 the “other” requester 91 may itself include its own respective memory region “access permission checking” circuit 91A that performs the “bounds check” for any requests coming from the other requester 91. In that case, the memory region “access permission checking” circuit (or circuits) that perform the “bounds check” for any requests coming from the execution engine 65 may be located in any suitable location, e.g. within the execution engine 65, within the work group local storage shared memory unit 74, within the bank access control units (circuits) 32, etc., as described above in relation to the earlier figures.
[0199] It can be seen from the above, that the technology described herein, in its embodiments at least, can provide improved operations and implementation when using work group local storage for work groups being executed by a programmable processing unit of a graphics processor.
[0200] The foregoing detailed description has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the technology described herein to the precise form disclosed. Many modifications and variations are possible in the light of the above teaching. The described embodiments were chosen in order to best explain the principles of the technology described herein and its practical applications, to thereby enable others skilled in the art to best utilise the technology described herein, in various embodiments and with various modifications as are suited to the particular use contemplated. It is intended that the scope be defined by the claims appended hereto.
Claims
P08058W001 ; 170020 / 01Claims;1. A graphics processing system, the graphics processing system comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processing system; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the graphics processing system further comprising one or more memory region access permission checking circuit that control memory accesses to the storage, the one or more memory region access permission checking circuits configured to: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage when it is determined that the execution thread is other than permitted to access that region of the storage.
2. The system of claim 1 , wherein different groups of work items may reside in different trust domains, and wherein determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage comprises determining whether the respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or by another group of work items within the same trust domain as the group of work items including that execution thread; and wherein access to the respective region of the storage is prevented when it is determined that the request is being made for an execution thread other than an execution thread of the work group for which the respective region of the storage has been allocated for use by or an execution thread of another group of work items within the same trust domain.
3. The system of claim 2, wherein the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or by another group of work items within the same trust domain as the group of work items including that execution thread comprises comparing an identifier associated with the group of work items for which the request is being made with a corresponding identifier, stored in association with the respective region of the storage that the execution thread is requesting access to, identifying the group of work items or trust domain for which the respective region of the storage has been allocated.
4. The system of claim 2, wherein the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or by another group of work items within the same trust domain as the group of work items including that execution thread comprises comparing a memory address for which the request is being made with a range of memory addresses identifying the respective region(s) of the storage that has been allocated for the group of work items or trust domain including the group of work items for which the request is being made.
5. The system of any one of the preceding claims, wherein the one or more memory region access permission checking circuits comprises a memory region access permission checking circuit that is associated with the programmable execution unit.
6. The system of any one of the preceding claims, wherein the one or more memory region access permission checking circuits are associated with the storage.
7. The system of claim 6, wherein the storage is configured as multiple, independently accessible banks, and wherein a respective memory region access permission checking circuit is associated with each independently accessible bank of storage.
8. The system of any one of the preceding claims, wherein the storage is available to be used for other purposes as well as storage in which a respective storage region can be allocated for temporary use by a respective group of executionthreads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, and wherein the one or more memory region access permission checking circuits are also configured to control access to the storage for requesters other than execution threads of work groups being executed by the programmable processing unit.
9. The system of claim 8, wherein the system further comprises: one or more arbitration circuits configured to receive both storage access requests relating to work group storage regions, and access requests from other requester(s), and to arbitrate between those requests, and wherein a respective memory region access permission checking circuit is associated with each of the one or more arbitration circuits.
10. The system of any one of the preceding claims, wherein the storage is local to and on-chip with the programmable processing unit of the graphics processor.
11. A graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; the graphics processor further comprising one or more memory region access permission checking circuits that control memory accesses to storage, wherein respective regions of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the one or more memory region access permission checking circuits configured to: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: determine whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and prevent the execution thread accessing the respective region of the storage when it is determined that the execution thread is other than permitted to access that region of the storage.
12. A method of operating a graphics processing system, the graphics processing system comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processing system; storage in which a respective region of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed; and one or more memory region access permission checking circuits that control memory accesses to the storage, the method comprising: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits: determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and preventing the execution thread accessing the respective region of the storage when it is determined that the execution thread is other than permitted to access that region of the storage.
13. The method of claim 12, wherein different groups of work items may reside in different trust domains, and wherein determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage comprises determining whether the respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or another group of work items within the same trust domain as the group of work items including that execution thread; and wherein the method comprises preventing access to the respective region of the storage when it is determined that the request is being made for an execution thread other than an execution thread of the work group for which the respective region of the storage has been allocated for use by or an execution thread of another group of work items within the same trust domain14. The method of claim 13, wherein the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or by another group of work items within the same trust domain as the group of work items including that execution thread comprises comparing an identifier associated with the group of work items for which the request is being made with a corresponding identifier stored in association with the respective region of the storage identifying the group of work items or trust domain for which the respective region of the storage has been allocated.
15. The method of claim 13, wherein the determining whether a respective region of the storage that the execution thread is requesting access to has been allocated for use either by the group of work items including that execution thread or by another group of work items within the same trust domain as the group of work items including that execution thread comprises comparing a memory address for which the request is being made with a range of memory addresses identifying the respective region(s) of the storage that has been allocated for the work group or trust domain including the group of work items for which the request is being made.
16. The method of any one of claims 12 to 15, wherein the one or more memory region access permission checking circuits comprises a memory region access permission checking circuit that is associated with the programmable execution unit.
17. The method of any one of claims 12 to 16, wherein the one or more memory region access permission checking circuits are associated with the storage.
18. The method of claim 17, wherein the storage is configured as multiple, independently accessible banks, and wherein a respective memory region access permission checking circuit is associated with each independently accessible bank of storage.
19. The method of any one of claims 12 to 18, wherein the storage is available to be used for other purposes as well as storage in which a respective storage region can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, and wherein the one or more memory region access permission checking circuits arealso configured to control access to the storage for requesters other than execution threads of work groups being executed by the programmable processing unit.
20. The method of claim 19, wherein the system further comprises: one or more arbitration circuits configured to receive both storage access requests relating to work group storage regions, and access requests from other requester(s), and to arbitrate between those requests, and wherein a respective memory region access permission checking circuit is associated with each of the one or more arbitration circuits.21 . The method of any one or claims 12 to 20, wherein the storage is local to and on-chip with the programmable processing unit of the graphics processor.
22. A method of operating a graphics processor, the graphics processor comprising: a programmable processing unit operable to execute processing programs for execution threads corresponding to work items to be processed by the graphics processor; and one or more memory region access permission checking circuits that control memory accesses to storage, wherein respective regions of the storage can be allocated for temporary use by a respective group of execution threads corresponding to a group of work items being executed by the programmable processing unit while the group of execution threads are being executed, the method comprising: when a request is made for an execution thread of a group of work items that is being executed by the programmable processing circuit to access a respective region of the storage: the one or more memory region access permission checking circuits: determining whether the execution thread that is requesting access to the respective region of the storage is permitted to access that region of the storage; and preventing the execution thread accessing the respective region of the storage when it is determined that the execution thread is other than permitted to access that region of the storage.
23. A computer program comprising computer software code for performing the method of any one of claims 12 to 22 when the program is run on one or more processors.
Citation Information
Patent Citations
Memory protection at a thread level for a memory protection key architecture
US20170286326A1
Implementing per-thread memory access permissions
US20180329835A1
Boosting local memory performance in processor graphics
US20200371804A1