Resource allocation
By prioritizing the allocation of continuous memory blocks for workgroup tasks in SIMD processors and using state arrays to manage memory resources, deadlocks and fragmentation problems caused by shared memory allocation are solved, and memory usage efficiency and task processing speed are improved.
Patent Information
- Application Number
- CN202510415638.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-15
- Filing Date
- 2018-09-17
- Publication Date
- 2025-07-22
AI Technical Summary
In multi-processing unit systems, deadlocks and resource fragmentation problems caused by dynamic allocation mechanisms of shared memory, especially in SIMD processors, lead to task processing delays and resource waste.
The resource allocator is used to divide the shared memory into multiple memory parts, and allocate memory resources to the workgroup tasks by prioritizing the allocation of continuous blocks. It combines fine state arrays and coarse state arrays for efficient management to ensure that the tasks can continuously receive memory resources and avoid deadlocks.
It effectively avoids deadlocks, improves the utilization rate of memory resources, reduces task processing delays, and improves the robustness and efficiency of the system.
Smart Images

Figure CN120353582A_ABST
Abstract
Description
[0001] Division Application Instructions
[0002] This application is a divisional application of the invention patent application with the application date of September 17, 2018, the application number of 201811083351.8, and the title of "Resource Allocation". Technical Field
[0003] The present disclosure relates to a memory subsystem for multiple processing units. Background Art
[0004] In a system including multiple processing units, a shared memory that can be accessed by at least some of the processing units is typically provided. Due to the range of processing tasks that can be run on the processing units and that have different memory requirements respectively, the shared memory resources are generally not fixed for each processing unit at design time. An allocation mechanism is typically provided that allows processing tasks running on different processing units to request the allocation of one or more regions of the shared memory respectively. This enables the shared memory to be dynamically allocated for use by the tasks executed at the processing units.
[0005] Efficient use of the shared memory can be achieved through careful design of the allocation mechanism. Summary of the Invention
[0006] The present summary is provided to introduce a selection of concepts that are further described below in the detailed description. The present summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0007] There is provided a memory subsystem for a single instruction multiple data (SIMD) processor including multiple processing units, the multiple processing units being configured to process one or more work groups respectively including multiple SIMD tasks, the memory subsystem including:
[0008] A shared memory divided into multiple memory portions for allocation to tasks to be processed by the processor; and
[0009] A resource allocator configured to, in response to receiving a memory resource request for a first memory resource for a first received task of a work group, allocate a block of memory portions to the work group, the size of the block of memory portions being sufficient for each task in the work group to receive a memory resource equivalent to the first memory resource in the block of memory portions.
[0010] The resource allocator may be configured to allocate a contiguous block of memory resources as the block of memory portions.
[0011] The resource allocator can be configured to allocate the requested first memory resource to the task from the block memory section when servicing the first received task of the workgroup, and reserve the remaining memory portion of the block memory section against tasks assigned to other workgroups.
[0012] The resource allocator can be configured to, in response to subsequently receiving a memory resource request for a second task of the workgroup, allocate the memory resource of the block memory section to the second task.
[0013] The resource allocator can be arranged to receive memory resource requests from multiple different requesters, and to preferentially service the memory requests received from the following requester in response to allocating the block memory section to the workgroup, where the first received task of the workgroup was received from that requester.
[0014] The resource allocator can be further configured to, in response to receiving an indication that the processing of a task of the workgroup has been completed, release the memory resources allocated to the task without waiting for the processing of the workgroup to complete.
[0015] The shared memory can be further divided into a plurality of non - overlapping windows, each non - overlapping window including a plurality of memory sections, and the resource allocator is configured to maintain a window pointer indicating the current window, where memory sections will be attempted to be allocated in the current window in response to the next received memory request.
[0016] The resource allocator can be implemented in binary logic circuitry, and the window length can be such that the availability of all memory sections of each window can be checked in a single clock cycle of the binary logic circuitry.
[0017] The resource allocator can be further configured to maintain a fine - grained status array, which is arranged to indicate whether each memory section of the shared memory is allocated to a task.
[0018] The resource allocator can be configured to, in response to receiving a memory resource request for the first received task of the workgroup, search for a contiguous block of memory sections in the current window indicated by the fine - grained status array as available for allocation, and the resource allocator is configured to, if such a contiguous block is identified in the current window, allocate that contiguous block to the workgroup.
[0019] The resource allocator can be configured to allocate a contiguous block of memory sections such that the block starts at the lowest possible position in the window.
[0020] The resource allocator may further be configured to maintain a coarse state array that is arranged to indicate for each window of the shared memory whether all memory portions of that window are unallocated. The resource allocator is configured to check the coarse state array in parallel with searching for contiguous blocks of memory portions in the current window to determine whether the size of the requested block can be provided by one or more subsequent windows; the resource allocator is configured to allocate to a workgroup a block that includes a first memory portion of the current window that is in a contiguous block with one or more subsequent windows and extends into the memory portions of those subsequent windows if a sufficiently large contiguous block cannot be identified in the current window and the requested block can be provided by one or more subsequent windows.
[0021] The resource allocator may further be configured to form an overflow metric in parallel with searching the current window, the overflow metric representing the memory resources of the memory portions of the required block that cannot be provided in the current window starting from the first memory portion of the current window in the unallocated memory portions of the contiguous blocks of the immediately subsequent windows. The resource allocator is configured to, if a sufficiently large contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then attempt to allocate a memory portion to a workgroup by searching in the subsequent windows for unallocated memory portions of contiguous blocks whose total size is sufficient to provide the overflow metric starting from the first memory portion of the subsequent windows.
[0022] The fine state array may be a bit array in which each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated.
[0023] The coarse state array may be a bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is fully allocated.
[0024] The resource allocator may be configured to form each bit of the coarse state array by performing an OR reduction on all the bits of the fine state array that correspond to the memory portions that are present in the window corresponding to that bit in the coarse state array.
[0025] The window length may be a power of 2.
[0026] The resource allocator may maintain a data structure that identifies which workgroups among one or more workgroups are currently being allocated blocks of memory portions.
[0027] According to a second aspect, there is provided a method of allocating shared memory resources to tasks executed in a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units, each processing unit being configured to process one or more workgroups respectively including a plurality of SIMD tasks, the method including:
[0028] Receiving a shared memory resource request for a first memory resource for a first received task of a workgroup; and
[0029] Allocating a memory portion of a block of shared memory to the workgroup, the size of the block of memory portion being sufficient for each task in the workgroup to receive a memory resource equivalent to the first memory resource in the block of memory portion.
[0030] Allocating a memory portion of a block to the workgroup may include allocating the block of memory portion as a contiguous block of memory.
[0031] The method may further include: allocating the requested first memory resource from the block of memory portion to the task when servicing the first received task of the workgroup, and reserving the remaining memory portion of the block of memory portion against tasks allocated to other workgroups.
[0032] The method may further include: in response to a subsequent received memory resource request for a second task of the workgroup, allocating the memory resource of the block of memory portion to the second task.
[0033] The method may further include: receiving memory resource requests from a plurality of different requesters, and preferentially servicing memory requests received from the following requester in response to allocating the block of memory portion to the workgroup, the first received task of the workgroup being received from the requester.
[0034] The method may further include: in response to receiving an indication that processing of a task of the workgroup has been completed, releasing the memory resources allocated to the task without waiting for the processing of the workgroup to complete.
[0035] The method may further include: maintaining a fine-grained status array arranged to indicate whether each memory portion of the shared memory is allocated to a task.
[0036] Allocating a memory portion of a block to the workgroup may include: searching for a contiguous block of memory portion indicated as available for allocation by the fine-grained status array in the current window, and if such a contiguous block is identified in the current window, allocating the contiguous block to the workgroup.
[0037] Allocating a contiguous block may include: allocating a memory portion of the contiguous block such that the contiguous block starts at the lowest possible position in the window.
[0038] The method may further include:
[0039] Maintaining a coarse state array that indicates for each window of the shared memory whether all memory portions of the window are allocated;
[0040] Checking the coarse state array in parallel with searching for memory portions of consecutive blocks in the current window to determine whether the size of the requested block can be provided by one or more subsequent windows; and
[0041] If a large enough consecutive block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, allocating to a workgroup a block that includes memory portions starting from a first memory portion of the current window in the consecutive block with the subsequent window and extending to memory portions in those subsequent blocks.
[0042] The method may further include:
[0043] Forming a representation of an overflow metric in parallel with searching the current window, the overflow metric representing memory resources for memory portions of a required block that cannot be provided in the current window starting from the first memory portion of the current window in unallocated memory portions of consecutive blocks of the immediately subsequent window; and
[0044] If a large enough consecutive block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then subsequently attempting to allocate a block of memory portions to the workgroup by searching unallocated memory portions of consecutive blocks in the subsequent window starting from the first memory portion of the subsequent window that are large enough in total size to accommodate the overflow metric.
[0045] According to a third aspect, there is provided a memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units for processing SIMD tasks, the memory subsystem including:
[0046] A shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor;
[0047] A translation unit configured to associate a task with one or more physical addresses of the shared memory; and
[0048] A resource allocator configured to, in response to receiving a memory resource request for a first memory resource regarding a task, allocate a consecutive virtual memory block to the task and cause the translation unit to associate the task with a plurality of physical addresses of memory portions respectively corresponding to partitions of the virtual memory block, the memory portions together implementing the complete virtual memory block.
[0049] At least some of the multiple memory portions of the block may be discontinuous in the shared memory.
[0050] A virtual memory block may include a base address of a physical address as one of the memory portions associated with a task.
[0051] A virtual memory block may include a base address of a logical address of the virtual memory block.
[0052] The translation unit is operable to subsequently service an access request received from a task regarding the virtual memory block, the access request including an identifier of the task and an offset of a storage area in the virtual memory block, and the translation unit is configured to service the access request in a memory portion corresponding to a partition of the virtual memory block indicated by the offset.
[0053] The translation unit may include: a content-addressable memory configured to return one or more corresponding physical addresses in response to receiving an identifier of an item and an offset in the virtual memory block.
[0054] A SIMD processor may be configured to process a workgroup including multiple tasks, the task being the first received task of the workgroup, and a resource allocator is configured to reserve sufficient memory portions for the workgroup such that each task of the workgroup can receive a memory resource equivalent to a first memory resource, and allocate the requested memory resource from the memory portions reserved for the workgroup to the first received task.
[0055] The resource allocator may be configured to allocate virtual memory blocks to each task of the workgroup such that the virtual memory blocks allocated to the tasks of the workgroup together represent the virtual memory blocks of a continuous superblock.
[0056] The resource allocator may be configured to, in response to subsequently receiving a memory resource request regarding a second task of the workgroup, allocate continuous virtual memory blocks from the superblock to the second task.
[0057] According to a fourth aspect, there is provided a method of allocating shared memory resources to tasks executed in a single instruction multiple data (SIMD) processor, the SIMD processor including multiple processing units respectively configured to process SIMD tasks, the method including:
[0058] Receiving a memory resource request for a first memory resource regarding a task;
[0059] Allocating continuous virtual memory blocks to the task; and
[0060] Associating the task with multiple physical addresses of memory portions of the shared memory, each memory portion corresponding to a partition of the virtual memory block, and the memory portions together implement a complete virtual memory block.
[0061] The method may further comprise:
[0062] subsequently receiving an access request received from a task regarding a virtual memory block, the access request including an identifier of the task and an offset of a storage area in the virtual memory block; and
[0063] serving the access request by accessing a memory portion corresponding to a partition of the virtual memory block indicated by the offset.
[0064] Tasks may be grouped together in a workgroup for execution at a SIMD processor, a task may be a first received task of the workgroup, and allocating contiguous virtual memory blocks to the tasks may include:
[0065] reserving a sufficient memory portion for the workgroup for each task of the workgroup to receive a memory resource equivalent to a first memory resource; and
[0066] allocating the requested memory resource from the memory portion reserved for the workgroup to the first received task.
[0067] Reserving a memory portion for the workgroup may include reserving virtual memory blocks for each task of the workgroup such that the virtual memory blocks allocated to the tasks of the workgroup together represent the virtual memory blocks of a contiguous superblock.
[0068] The method may further comprise: in response to a subsequent receipt of a memory resource request for a second task of the workgroup, allocating contiguous virtual memory blocks from the superblock to the second task.
[0069] The memory subsystem may be implemented in hardware on an integrated circuit.
[0070] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided.
[0071] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided, the method comprising:
[0072] processing a computer-readable description of a graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and
[0073] manufacturing a graphics processing system according to the circuit layout description using an integrated circuit generation system.
[0074] An integrated circuit definition data set is provided that, when processed in an integrated circuit manufacturing system, configures the system to manufacture the memory subsystem described herein.
[0075] A non-transitory computer-readable storage medium is provided, having stored thereon a computer-readable description of an integrated circuit, which when processed in an integrated circuit manufacturing system causes the integrated circuit manufacturing system to manufacture the memory subsystem described herein.
[0076] A non-transitory computer-readable storage medium is provided, having stored thereon a computer-readable description of the memory subsystem described herein, which when processed in an integrated circuit manufacturing system causes the integrated circuit manufacturing system to:
[0077] Process a computer-readable description of a graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and
[0078] Manufacture a graphics processing system using an integrated circuit generation system based on the circuit layout description.
[0079] An integrated circuit manufacturing system is provided, configured to manufacture the memory subsystem described herein.
[0080] An integrated circuit manufacturing system is provided, comprising:
[0081] A non-transitory computer-readable storage medium having stored thereon a computer-readable integrated circuit description of the memory subsystem described herein;
[0082] A layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit implementing the memory subsystem; and
[0083] An integrated circuit generation system configured to manufacture the memory subsystem based on the circuit layout description.
[0084] A memory subsystem is provided, configured to perform the method described herein. A computer program code is provided for performing the method described herein. A non-transitory computer-readable storage medium is provided, having stored thereon computer-readable instructions, which when executed at a computer system cause the computer system to perform the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The present invention is described by way of example with reference to the accompanying drawings. In the drawings:
[0086] Figure 1 Shows a conventional allocation of shared memory for tasks executed at a SIMD processor.
[0087] Figure 2 Is a schematic diagram of a first memory subsystem including a resource allocator configured according to the principles described herein.
[0088] Figure 3Illustrates a memory subsystem in the context of a computer system having multiple processing cores.
[0089] Figure 4 Illustrates a shared memory including blocks of tasks assigned on a per-slice basis to workgroups.
[0090] Figure 5 Illustrates the allocation of shared memory by a resource allocator to tasks executed at a SIMD processor as workgroups.
[0091] Figure 6 Illustrates the allocation of shared memory by a resource allocator configured according to a particular embodiment described herein.
[0092] Figure 7 Is a schematic diagram of a second memory subsystem including a resource allocator configured according to the principles described herein.
[0093] Figure 8 Illustrates the correspondence between virtual memory blocks and memory portions of underlying shared memory.
[0094] Figure 9 Is a flowchart illustrating the operation of a first memory subsystem configured according to the principles described herein.
[0095] Figure 10 Is a flowchart illustrating the operation of a second memory subsystem configured according to the principles described herein.
[0096] Figure 11 Is a schematic diagram of an integrated circuit manufacturing system. Detailed Description
[0097] The following description is given by way of example so that those skilled in the art can make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art. The embodiments are described only by way of example.
[0098] The term "task" is used herein to refer to a set of data items and the work to be performed on those data items. For example, in addition to a set of data to be processed according to a program (e.g., the same sequence of ALU instructions or references to those ALU instructions), a task may include a program or be associated with a program or may also include a reference to a program, where the set of data may include one or more data elements (or data items, e.g., multiple pixels or vertices).
[0099] The term "program instance" is used herein to refer to an individual instance of code walking through. Thus, a program instance refers to a single data item and a reference (e.g., a pointer) to the program to be executed on the data item. Thus, a task can be considered to include multiple program instances (e.g., up to 32 program instances), although in practice each task only requires a single instance of the common program (or reference). A group of tasks sharing a common purpose share local memory and can execute the same program (although they can execute different parts of the program) or compatible programs on different data blocks can be connected by a group ID. A group of tasks with the same group ID can be referred to as a "workgroup" (thus, the group ID can be referred to as the "workgroup ID"). Thus, there is a hierarchy of terms, where a task includes multiple program instances and a group (or workgroup) includes multiple tasks.
[0100] Some processing systems that provide shared memory can include one or more SIMD (Single Instruction Multiple Data) processors, each SIMD processor being configured to execute the tasks of a workgroup in parallel. Each task of a workgroup may require an allocation of shared memory. However, generally, there are certain interdependencies among the processing of tasks executed at the SIMD processor. For example, for some workgroups, a barrier synchronization point is defined that all tasks must reach in order to continue processing any of these tasks beyond the barrier synchronization point and to complete the processing of the workgroup as a whole. Sometimes more than one barrier synchronization point can be defined for a workgroup, where all tasks need to reach a given barrier in order to continue processing any of these tasks beyond the barrier.
[0101] For traditional allocation mechanisms of shared memory, the use of barrier synchronization points can lead to deadlocks in the processing of workgroups. This is shown in the Figure 1 simple example, where Figure 1It is shown that a SIMD processor 102, in which a shared memory 100 is divided into a plurality of memory sections 101 and has five processing elements 109 - 113, is configured to process a work group 'A' including five tasks 103 - 107 in parallel. In this example, only the processing of four tasks is enabled, and all these tasks have reached a barrier synchronization point 108. This is because each task requires two memory sections of a continuous block in the shared memory (marked with the letter 'A'), but once the blocks in the shared memory are allocated for the first four tasks, it is impossible to allocate two memory sections of a continuous block for the fifth task. The memory sections belonging to the second work group are marked with the letter 'B'. A number of memory sections 114 in the shared memory are available, but these memory sections are fragmented and cannot provide a continuous block of memory. As a result, a deadlock has occurred: the processing of four tasks 103 - 106 has stopped, waiting for the fifth task 107 to reach the barrier synchronization point 108, but the processing of the fifth task 107 cannot be started because its request for allocating the shared memory cannot be satisfied.
[0102] Such a deadlock can cause significant delays, and generally can only be resolved when the processing of the deadlocked work group reaches a predetermined timeout. This not only wastes the time spent waiting for the timeout, but also partial processing has been performed on the tasks that have reached the barrier synchronization point (this processing will need to be repeated). The situation of deadlocks is particularly problematic in systems including multiple SIMD processors and all these SIMD processors sharing a common memory, because when the available shared memory is limited or when tasks require a relatively large continuous block of memory, more than one work group can be locked each time.
[0103] Avoid Fragments
[0104] Figure 2FIG. is a schematic diagram showing a memory subsystem 200 configured to allocate shared memory in a manner that addresses the above problems. The memory subsystem 200 includes a resource allocator 201 and a shared memory 202, which is divided into a plurality of memory portions 203. The resource allocator may be configured to receive memory requests from a plurality of requesters 205. For example, each task running at a processing unit may be a requester, each processing unit may be a requester, or each type of processing unit may be a requester. Different requesters may have different hard-wired input terminals to the resource allocator, and the allocator needs to arbitrate among these hard-wired input terminals. The resource allocator may service memory requests from requesters based on any suitable manner. For example, memory requests may be serviced based on simple polling or according to a set of arbitration rules for selecting the next requester to be serviced. In other embodiments, memory requests from a plurality of requesters may be aggregated at a request queue, where memory allocation requests are received into the request queue, and the resource allocator may receive memory requests to be serviced from the request queue.
[0105] The processing unit may access the allocated portion of the shared memory through an input / output arbiter 204. The I / O arbiter 204 arbitrates access to the shared memory among a plurality of units (such as processing units or other units capable of accessing the shared memory) that submit access requests (such as read / write) through an interface 211, where the interface 211 may include one or more hard-wired connections, and the I / O arbiter arbitrates among the one or more hard-wired connections.
[0106] Figure 3 FIG. shows the memory subsystem 200 in the context of a computer system 300 that includes a plurality of processing units 301 and a job scheduler 302, where the job scheduler 302 is configured to schedule tasks it receives for processing at the processing units. One or more of the processing units may be SIMD processors. The computer system may be, for example, a graphics processing unit (GPU), and the processing units may include one or more units (all of which may be SIMD processors) for performing integer operations, floating-point operations, compound arithmetic, texture address calculations, sampling operations, etc. The processing units may together form a set of parallel pipelines 305, where each processing unit represents one pipeline in the set of pipelines. The computer system may be, for example, a vector processor.
[0107] Each processing unit may include one or more arithmetic logic units (ALUs), and each SIMD processor may include one or more different types of ALUs. For example, each type of ALU is optimized for a specific type of computation. In the example where the GPU provides processing unit 301, processing block 104 may include multiple shader cores, and each shader core includes one or more ALUs.
[0108] The work scheduler 302 may include a scheduler unit 303 and an instruction decoder 304. The scheduler unit is configured to perform scheduling of tasks to be executed on the processing unit 301, and the instruction decoder is configured to decode the tasks into a form suitable for execution on the processing unit 301. It will be appreciated that the specific decoding of the tasks performed by the instruction decoder will be determined by the type and configuration of the processing unit 301.
[0109] The resource allocator 201 is operable to receive requests for allocating memory resources to tasks executed at the processing unit 301 of the computer system. Each task or processing unit may be a requester 205. One or more processing units may be SIMD processors configured to execute tasks of a workgroup in parallel. Each task of the workgroup may request the resource allocator to allocate memory (depending on the specific architecture of the system), and such requests may be made, for example, by the work scheduler on behalf of the task when assigning tasks to the parallel processor or by the task itself.
[0110] In some architectures, tasks related to various workgroups or for execution on different types of processing units may be queued for processing in one or more common work queues, such that tasks for parallel processing as a workgroup are assigned to appropriate processing units by the work scheduler 302 during a certain time period. Tasks related to other workgroups and / or processing units may be interleaved with the tasks of the workgroup. Requests for memory resources of the resource allocator may generally be made for tasks, as the tasks are scheduled at the processor. The task itself may not need to make a request for memory resources. For example, in some architectures, the work scheduler 302 may make a request for memory resources on behalf of the tasks it schedules at the processing unit. Many other architectures are possible.
[0111] The resource allocator is configured to allocate memory resources to an entire workgroup when it first receives an allocation request for a task of the workgroup. The allocation request generally indicates the size of the memory resources required. For example, a task may require a certain number of bytes or a memory portion. Since the SIMD processor executes the same instructions in parallel on different source data, each task of the workgroup for execution at the SIMD processor has the same memory requirement. The resource allocator is configured to allocate a contiguous block of memory for the entire workgroup in response to the first request for memory for a task of the workgroup that it receives. The size of the block is at least N times the memory resources requested in the first request, where N is the number of tasks in the workgroup (e.g., 32). The number of tasks in the workgroup can be fixed for a given system and is thus known in advance to the resource allocator, and the number of tasks in the workgroup can be provided to the resource allocator together with the first request for memory resources, and / or the number of tasks in the workgroup is available to the resource allocator (e.g., as a parameter held at a data store accessible to the resource allocator).
[0112] If there is sufficient contiguous space in the shared memory for the entire workflow block, the resource allocator is configured to respond to the first received request for memory resources for the task of the workflow by utilizing the memory resources requested by the task of the workflow. This enables the processing of the first task of the workgroup to begin. The resource allocator also reserves the remaining portion of the shared memory block (which will be allocated to other tasks of the workgroup) to prevent the block from being allocated to tasks of other workgroups. If there is not sufficient contiguous space in the shared memory for the entire workflow block, the first received memory request is rejected. This process ensures that once the first task of the workgroup receives an allocation of memory, the system can guarantee that all tasks of the workgroup will receive their memory allocations. This avoids the deadlock scenario described above.
[0113] When the resource allocator receives a request for memory resources for a subsequently received task of the same workgroup, the resource allocator allocates the required resources to these tasks from the reserved block. The resource allocator can then sequentially allocate adjacent "slices" of the reserved shared memory block to the tasks in the order in which the corresponding memory requests are received, where the first received task receives the first slice of the block. This is shown in Figure 4 shown Figure 4It is shown that the shared memory 202 includes a plurality of memory portions 203 and a block 402 reserved for the workgroup after receiving the first memory request for the task of the workgroup. In this example, each slice includes two memory portions. The first slice 403 of the block is assigned to the first task, and the subsequent slices of the block are reserved in the shared memory so that when the next memory request for the task of the workgroup is received, the next slice 404 is assigned, and so on until the slices 403 - 407 of the block are assigned to tasks and the workgroup can be processed completely.
[0114] In the case where the tasks of the workgroup have a desired or predetermined order, the resource allocator can alternatively allocate the memory slices of the reserved memory block to the tasks of the workgroup according to this order. In some architectures, the memory requests for tasks may be received out of order. For example, even if the memory request for the third task is received before the memory request for the second task, the third slice of the block can be assigned to the third task among the five tasks. In this way, a predetermined slice of the block reserved for the workgroup can be assigned to each task of the workgroup. Since the first received task may not be the first task of the workgroup, it is not necessary to assign the first slice of the block reserved for the workgroup to the first task.
[0115] In some architectures, a task may make further requests for memory during processing. These memory requests can be processed in the same way as the initial request for memory, where when the first request regarding the workgroup is received, based on the fact that each task processed by the SIMD processor will require the same memory resources, a block of memory resources is allocated to all tasks of the workgroup. The size of the allocated block will be at least N times the memory resources requested in the first request, where N is the number of tasks in the workgroup.
[0116] Once a certain task has been scheduled at the processing unit, the task can access its allocated memory resources (e.g., by performing reads and / or writes) at the shared memory 202 as indicated by the data path 210.
[0117] The resource allocator can be configured to immediately release the slice of the shared memory allocated to the task when the processing of each task is completed at the processing unit (i.e., it is not necessary to wait until the entire workgroup is completed). This ensures that the memory portions that are no longer needed are released as soon as possible for use by other tasks. For example, the release of the last "slice" can create a block of sufficient size to be allocated to the subsequent consecutive workgroup to be processed next.
[0118] Figure 5Shows the sharing of memory for task assignment according to the principles described herein. The shared memory 202 is shown divided into a plurality of memory sections 203, and the SIMD processor 502 has five processing elements 503 - 507. Contrary to the conventional manner of shared memory allocation, this arrangement corresponds to Figure 1 the arrangement shown. The shared memory includes a memory section 508 of an existing block assigned to workgroup B (the corresponding memory section in the figure is labeled 'B'). Consider the point at which a first memory request for a task 509 of a second workgroup A is received. When the memory request for task 509 is received, a memory section of block 514 (labeled 'A') is assigned to the shared memory, as shown. Thus, as described above, in response to receiving a first memory request for a task of a workgroup, the assignment of the memory section of that block for the entire workgroup is performed. A slice 515 of two memory sections of block 514 is assigned to the first task. Slices 516 - 519 of that block are reserved for other tasks of workgroup A.
[0119] When memory requests for subsequent tasks 510 - 513 are received, each memory request can be serviced by assigning to the corresponding task and scheduling one of the reserved slices 516 - 519 for task assignment to be executed on the SIMD processor 502. Since all the memory required by the workgroup has been allocated in advance, the last task 513 receives slice 519 of the block. Since all tasks have received the memory allocations they need, the processing of all tasks can proceed to the barrier synchronization point 520. Once all tasks have been processed to the barrier 520, the processing of all tasks can be permitted to continue to completion, so that the workgroup can be processed as a whole. Due to the method of allocating contiguous space to the entire block, the shared memory is more robust to fragmentation, and contiguous space in the shared memory is more available for allocation to new blocks.
[0120] If there is not sufficient contiguous space available for the entire block 514, the memory request for the first task of workgroup A is rejected. The rejected memory request can be handled in various ways, but generally, it can be serviced when the next resource allocator expects to service the corresponding requester according to its defined mechanism (e.g., based on polling or according to a set of rules for arbitration between requesters). Once sufficient memory becomes available, subsequent memory allocation requests can succeed (e.g., by tasks being completed on other processing units and the corresponding memory sections being released).
[0121] The resource allocator can maintain a data structure 206 that identifies which memory blocks have been allocated to which workgroups. For example, when reserving a memory block for a workgroup, the resource allocator can add an entry to the data structure that associates the memory block with the workgroup, so that when a task for the workgroup is subsequently received, the task allocator can identify which memory slices of which block will be allocated. Similarly, the resource allocator can use the data structure 206 to identify which tasks do not belong to workgroups that already have memory blocks allocated, so as to identify the tasks for which new block allocations will be performed. The data structure 206 can have any suitable configuration, for example, it can be a simple lookup table or register.
[0122] To allocate slices of reserved blocks to subsequently received tasks of a workgroup, the resource allocator needs to know which workgroup the memory requests it receives belong to. The memory requests can include an identifier of the workgroup to which the corresponding task belongs, to allow the resource allocator to identify memory requests that belong to workgroups for which blocks have been reserved. Alternatively, the resource allocator can access a data store that identifies which tasks belong to which workgroups (e.g., a lookup table maintained by the work scheduler 302). Generally, the memory request will identify the task for which the memory request is being made, and such a task identifier can be used to look up the corresponding workgroup in the data store.
[0123] Memory requests received for workgroups that already have memory blocks allocated can preferably be processed with priority. This helps to minimize latency in the system by ensuring that once a memory has been allocated to a workgroup and resources for completing the processing of that workgroup are reserved in the system, all tasks of the workgroup can start processing as soon as possible and reach any defined barrier synchronization points of the workgroup. In other words, ensuring that memory can be allocated to tasks as soon as possible after reserving the memory as a block for the workgroup helps to avoid tasks that are subsequently allocated their memory slices from delaying the processing of other tasks in the workgroup that already have other slices allocated and are ready to start processing.
[0124] Any suitable mechanism for processing memory requests with priority can be implemented. For example, the resource allocator can be configured to process requests from the same requester 205 from which an initial request was received that caused a block to be allocated, with priority. The requester can be made to have priority over other requesters by increasing the frequency at which the requester's memory requests are processed. The requester can be made to have priority by repeatedly processing memory requests received from the same requester after a block has been allocated until a memory request for a task belonging to a different workgroup is received. The resource allocator can use the above data structure 206 to identify which memory request corresponds to a workgroup for which a memory block has been reserved, so as to identify which memory request should be processed with priority.
[0125] Note that, in some embodiments of the systems described herein, some of the tasks targeted by the memory requests received by the resource allocator will not belong to a workgroup. For example, Figure 3 one or more processing units 301 used to process tasks in Figure 3 may not be SIMD processors. The resource allocator may be configured to identify, in the memory requests it receives, tasks that do not belong to a workgroup from one or more identifiers, or to identify the absence of one or more identifiers (e.g., no workgroup identifier). Alternatively or in addition, the resource allocator may determine whether a received memory request is made for a task that does not belong to a workgroup by looking up the identifier received in the memory request in a data store (e.g., a lookup table maintained by the work scheduler 302) that identifies which tasks belong to which workgroups.
[0126] A method of allocating memory portions at the resource allocator 201 will now be described.
[0127] The resource allocator may be configured to further divide the shared memory 202 into a plurality of non-overlapping windows, each non-overlapping window including a plurality of memory portions. The length of each window may be a power of two to enable an efficient implementation of the resource allocator in binary logic circuitry. The windows may all be of the same length. The resource allocator may maintain a fine-grained status array 207 indicating which memory portions are allocated and which are not. For example, the fine-grained status array may include a one-dimensional array of bits, where each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated (e.g., '1' indicates the portion is allocated; '0' indicates the portion is not allocated). The size of each window may be chosen such that the resource allocator can search for an unallocated memory portion in the window that is large enough to contain a contiguous block of a given size in a single clock cycle of the digital hardware implementing the resource allocator. Generally, the size of the shared memory 202 is too large for the resource allocator to search the entire memory space in one clock cycle.
[0128] The resource allocator may maintain a window pointer 208 (e.g., a register internal to the resource allocator) that identifies in which window of the shared memory the resource allocator will start a new memory allocation when receiving a first memory request for a workgroup for which memory has not yet been allocated. When receiving a memory request for the first task of a workgroup, the resource allocator checks in the current window where a contiguous memory block to be reserved for the workgroup can fit. This is referred to as a "fine check". The resource allocator uses a fine state array 207 to identify unallocated memory portions. The resource allocator may start at the lowest position of the current window and scan in the direction of increasing memory addresses to identify the lowest possible start position of the block and minimize fragmentation of the shared memory. Any suitable algorithm for identifying memory portions that can provide a contiguous set of the memory block may be used.
[0129] When allocating a block to a workgroup, the fine state array is updated to mark all corresponding memory portions of the block as allocated. In this way, the block can be reserved for the tasks of the workgroup.
[0130] The resource allocator may also maintain a coarse state array 209 that indicates whether each window of the shared memory is completely unallocated, in other words, whether all memory portions of the window are unallocated. The coarse state array may be a one-dimensional bit array where each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is completely allocated. The bit value of a window in the coarse state array may be calculated as the OR reduction of all bits in the state array corresponding to the memory portions located in that window.
[0131] The resource allocator may be configured to perform a coarse check in combination with the fine check. The coarse check includes identifying from the coarse state array 209 whether the next one or more windows following the current window indicated by the window pointer 208 can provide the new memory block required by the workgroup, i.e., if the coarse state array 209 indicates that the next window is completely unallocated, then all these memory portions can be used by the block. One or more windows following the current window may be checked because in some systems the required block may be larger than the size of a window.
[0132] The resource allocator may be configured to allocate a shared memory block to a workgroup using the following fine check and coarse check:
[0133] 1. If the fine-grained check of the memory block to be allocated to the workgroup is successful and the memory block can be provided in the current window, the memory block is reserved for use by the workgroup (e.g., by marking the corresponding memory portion as allocated in the fine state array and adding the workgroup to the data structure 206) and a slice (e.g., the first slice) of the memory block is allocated to the task for which the memory request regarding it was made. If the coarse-grained check of the memory indicates that the current window of the shared memory is completely unallocated, the coarse state array can be updated to indicate that the current window is now no longer completely unallocated.
[0134] 2. If the fine-grained check of the memory block to be allocated to the workgroup is unsuccessful but the coarse-grained check of the memory is successful, the memory block can be provided by one or more unallocated windows following the current window and the memory block is reserved for the workgroup. The memory block can be allocated starting from the first unallocated memory portion of the unallocated memory portion that forms a contiguous partition with the next window. The allocated memory block extends into one or more of the next windows identified as available by the coarse-grained check.
[0135] If both the fine-grained check and the coarse-grained check fail, the memory request can be postponed and the memory request can be attempted again later (possibly on the next clock cycle or after a predetermined number of clock cycles). The resource allocator can service other memory requests before returning the failed memory request (e.g., the resource allocator can move to service the memory request from the next requester according to some rules defined for arbitration among the requesters).
[0136] Figure 6 An example allocation according to the above method is shown, where the shared memory is divided into 128 memory portions and 4 windows each including 32 memory portions. The '1' bit 601 in the fine state array 207 indicates that the first 24 memory portions are allocated to the workgroup, but all higher memory portions are unallocated. The coarse state array indicates that only the first window includes the allocated memory portions (see the '1' bit 602 corresponding to the first window). In this example, the workgroup needs a 48-memory-portion block of memory 603. The resource allocator can determine that such a block is needed when receiving a memory request for 3 memory portions for the first task of the workgroup: it is known that the memory request has been received from a requester that is a SIMD processor of a workgroup for processing 16 tasks, and the resource allocator determines that the workgroup as a whole needs 48 memory portions.
[0137] Assume that the current window indicated by the window pointer of the resource allocator is the first window '0' (possibly because the first unallocated memory portion after the allocated memory portion is memory portion 24). The detailed check of the current window will fail because there are no remaining 48 memory portions in the first window '0'. However, the rough check will succeed because the second and third windows are indicated by the rough state array 209 as not being allocated at all. Therefore, a block of 48 will be allocated from memory portion 24 of the first window until memory portion 71 of the third window. The window pointer can be updated to identify the third window as the current window because the next available unallocated memory portion is the number 72.
[0138] The resource allocator can be configured to compute an overflow metric in parallel with the detailed check and the rough check. The overflow metric is given by the number of memory blocks required by the workgroup minus the number of unallocated memory portions at the top of the current window, so as to give a measure of the number of memory portions that would overflow the current window in the case of allocating the block starting from the lowest memory portion of the unallocated memory of consecutive blocks at the top of the current window.
[0139] The resource allocator can be configured to use the overflow metric to attempt to allocate memory blocks to the workflow in an additional step after steps 1 and 2 above:
[0140] 3. If the detailed check and the rough check fail but there are still some free memory portions at the top of the current window, use the overflow metric as the requested size and make a new attempt to allocate the memory block starting from the lowest address of the next window of the shared memory at the next opportunity (e.g., the next clock cycle). For the initial allocation attempt, the detailed check and the rough check as described above can be performed. This additional new attempt is beneficial because although the rough check on the next window fails, the lower partition of the next window and the free memory portions of the current window are sufficient to provide enough contiguous unallocated space to receive the memory block.
[0141] If step 3 fails, the memory request can be postponed and then the memory request can be attempted again as described above.
[0142] Once the allocation is successful, the resource allocator immediately updates the window pointer to the window containing the memory portion after the successful allocation of the memory portion block (this can be the current window or the next window).
[0143] Using a detailed check and a coarse check that are executed in parallel (optionally, in combination with an overflow metric that is executed in parallel with the detailed check and the coarse check), enables a resource allocator to efficiently attempt to allocate shared memory to a workgroup in a minimum number of clock cycles. On appropriate hardware, the above method can enable block allocation to be performed in two clock cycles (in the first clock cycle, the detailed check and the coarse check are executed and optionally the overflow metric is calculated; in the second clock cycle, an attempt is made to allocate according to steps 1-2 above and optional step 3).
[0144] Figure 9 A flowchart showing a resource allocator allocating shared memory is shown. The resource allocator receives a shared memory resource request 901 for a first task of a workgroup (i.e., the first task among the tasks of the workgroup for which a shared memory request has been received). In response, a shared memory block is allocated to the workgroup, the size of the shared memory block being sufficient to provide the resources requested by the first task to all tasks of the workgroup. Generally, the shared memory block will be a contiguous shared memory block. A portion 903 of the shared memory block is allocated for the first task, and this step may or may not be considered part of the memory allocation to the workgroup 902, and thus may be performed before, at the same time as, or after allocating memory resources to the workgroup. In some embodiments, the remaining portion of the shared memory may be reserved for other tasks of the workgroup.
[0145] When a subsequent shared memory request 904 for another task of the same workgroup is received, a portion 905 of the shared memory block allocated to the workgroup is allocated to the task. In some embodiments, in response to receiving a first memory request 901 from a task of a workgroup, memory allocation for all tasks may be performed. In such embodiments, further allocation may include responding by allocating memory resources to tasks established in response to the first received memory request for the workgroup.
[0146] The allocation 907 of memory resources from the block is repeated until all tasks of the workgroup have received the required memory resources. Then, the execution 906 of the workgroup may be performed, where each task uses the shared memory resources allocated to that task in the block allocated to the workgroup.
[0147] Fragment Insensitivity
[0148] A second configuration of a memory subsystem that also addresses the above limitations of the prior art for allocating shared memory will now be described.
[0149] Figure 7A memory subsystem 700 is shown that includes a resource allocator 701, a shared memory 702, an input / output arbiter 703, and a translation unit 704. The shared memory is divided into a plurality of memory sections 705. The resource allocator may be configured to receive memory requests from a plurality of requesters 205. For example, each task running on a processing unit may be a requester, each processing unit may be a requester, or each type of processing unit may be a requester. Different requesters may have different hardwired input terminals to the resource allocator, and the resource allocator needs to arbitrate among these hardwired input terminals. The resource allocator may service memory requests from requesters on any basis. For example, memory requests may be serviced based on simple polling or according to a set of arbitration rules for selecting the next requester to be serviced. In other embodiments, memory requests from a plurality of requesters may be aggregated at a request queue, memory allocation requests may be received into the request queue, and the resource allocator may receive memory requests to be serviced from the request queue. The allocated portions of the shared memory may be accessed by processing units using the input / output arbiter 703.
[0150] The memory subsystem 700 may be arranged in a computer system in the Figure 3 manner shown and described above.
[0151] In this configuration, the resource allocator is configured to allocate the shared memory 702 to tasks using the translation unit 704, and the input / output arbiter 703 is configured to access the shared memory 702 using the translation unit 704. The translation unit 704 is configured to associate each task regarding the received memory requests with the plurality of memory sections 705 of the shared memory. Each task generally has an identifier provided in the memory requests made regarding that task. The plurality of memory sections associated with a task need not be contiguous portions in the shared memory. However, from the perspective of the task, the task will receive an allocation of a contiguous memory block. This may be achieved by the resource allocator configured to allocate a contiguous virtual memory block to the task when the same task is actually associated with a potentially non-contiguous set of memory sections at the translation unit.
[0152] When a task accesses the memory region allocated to it (e.g., via a read or write operation), the translation unit 704 is configured to translate between the logical address provided by the task and the physical address where its data is stored or to which its data will be written. Note that the virtual address provided to the task by the resource allocator does not need to belong to a constant virtual address space, and in some examples, the virtual address will correspond to a physical address at the shared memory. It is advantageous for the logical address used by the task to have an offset relative to some base address (which itself can be a zero offset or some other predetermined value) with respect to a predetermined position (generally the start position of the block) in the virtual block. Since the virtual blocks are contiguous, only the relative position of the data in the block that needs to be accessed needs to be provided to the translation unit. The memory portion of the shared memory that constitutes a particular set of virtual blocks of a task can be determined by the translation unit based on the association between the task identifier in the data store area 706 and the memory portion.
[0153] In response to receiving a request for memory resources, the resource allocator is configured to allocate to the task a set of memory portions sufficient to satisfy the request. The memory portions do not need to be contiguous or stored sequentially in the shared memory. The resource allocator is configured to cause the translation unit to associate the task with each of the memory portions in the set of memory portions. The resource allocator can utilize a data structure 707 indicating which memory portions are available to track the unallocated memory portions at the shared memory. For example, the data structure can be a status array (e.g., the aforementioned fine-grained status array) whose bits identify whether each memory portion at the shared memory is allocated.
[0154] The translation unit can include a data store area 706 and can be configured to associate the task with a set of memory portions by correspondingly storing in the data store area the identifier of the task (e.g., task_id) and the physical address of each of the memory portions in the set of memory portions, as well as an indication of which partition of the contiguous virtual memory block each memory portion allocated to the task corresponds to (e.g., the offset of the partition in the virtual memory block or the logical address of the partition). In this way, the translation unit can, in response to the task providing its identifier and an indication of which partition of its virtual memory block it needs (e.g., the offset in its virtual block), return the appropriate physical address to the task. The offset can be the number of memory portions of the memory region relative to the base address (e.g., in the case where the memory portion is a basic allocation interval), the number of bytes or bits, or any other position indication. The base address can be but does not have to be the memory address where the virtual memory block starts; the base address can be a zero offset. The data store area can be an associative memory (e.g., a content-addressable memory) whose inputs can be the task identifier and an indication of the desired memory partition in its block (e.g., the memory offset relative to the base address of its block).
[0155] More generally, the translation unit can be provided at any suitable point along the data paths between the resource allocator 701 and the shared memory 702, and between the I / O arbiter 703 and the shared memory 702. More than one translation unit can be provided. Multiple translation units can share a common data storage area such as a content addressable memory.
[0156] The resource allocator can be configured to allocate logical base addresses to different tasks of a workgroup to identify consecutive virtual memory blocks within a consecutive virtual memory superblock for the workgroup as a whole that are assigned to each task of the workgroup.
[0157] Figure 8 The correspondence between the memory portion 805 of the shared memory 702 assigned to a task and the virtual block 801 assigned to the task is shown. The memory portions 805 of the shared memory labeled A - F provide the underlying storage area for partitions 806 of the virtual consecutive memory block 801. The memory blocks A - F are non - consecutive or are stored sequentially in the shared memory and are located among the memory portions 804 assigned to other tasks. The virtual block 801 has a base address 802, and the partitions can be identified by the offset 803 of the partition within the block relative to the base address.
[0158] In one example, in response to receiving a request for memory resources, the resource allocator is configured to allocate to the task a consecutive range of memory addresses starting from the physical address of the memory portion corresponding to the first partition of the block. In other words, the base address 802 of the virtual block (and Figure 8 the base address of partition 'A' or 809 in Figure 8 is the physical address of memory portion 'A' or 808 in
[0159] Figure 8 ). Assuming there is sufficient unallocated memory portion in the shared memory to service the memory request, the resource allocator allocates to the task a virtual memory block of the requested size starting from the physical memory address of the memory partition 809. Thus, from the perspective of the task, it is allocated a consecutive block of six memory portions starting from memory portion 808 in the shared memory; in reality, the task is allocated six non - consecutive memory portions labeled A, B, C, D, E, and F in the shared memory 702. Thus, the actual allocation of the memory portion for the task is labeled by the virtual block 801 it receives. In this example, the virtual block has a physical address and there is no such virtual address space.
[0159] Once a virtual memory block has been allocated for a task, it can access the memory using the I / O arbiter 703, where the I / O arbiter generally arbitrates access to the shared memory among multiple units (e.g., processing units or other units capable of accessing the shared memory) that submit access requests (e.g., read / write) via an interface 708, which may include one or more hardwired connections and among which the I / O arbiter arbitrates.
[0160] Continuing with the above example, consider a task that requests access to a region of block 801 located in partition 'D'. The task can submit a request to the I / O arbiter 703 that includes its task identifier (e.g., task_id) and an indication of the region of the block it needs to access. The indication can be, for example, an offset 803 relative to a base address 802 (e.g., number of bits, bytes, memory partial units), or a physical address formed based on the physical base address 802 assigned to the task and the offset 803 in the block where the requested memory region is found. The translation unit will look up the task unit identifier and the desired offset / address in the data store 706 to identify the corresponding physical address that identifies a memory portion 811 in the shared memory 702. The translation unit will service the access request from the memory portion 811 in the shared memory. In the case where the memory region in the block spans more than one memory partition, the data store will return more than one physical address for the corresponding data portions.
[0161] In an implementation where the task provides the physical address of the storage region, the physical address submitted in this case refers to a storage region in a memory portion 810 assigned to another task. Thus, different tasks can ostensibly access overlapping blocks in the shared memory. However, since the translation unit 704 translates each access request, the access to the physical address submitted by the task is redirected by the translation unit to the correct memory portion (811 in this example). Thus, in this case, the logical address that references the virtual memory block is actually a physical address, but not necessarily the physical address where the corresponding data of the block is stored in the shared memory.
[0162] In a second example, in response to receiving a request for memory resources, the resource allocator is configured to allocate to a task a logical memory address that starts from a logical base address and represents a contiguous range of virtual memory blocks. The logical address can be allocated in any suitable manner, since a given logical address will be mapped to a physical address of a corresponding underlying memory portion. For example, a logical address in a virtual address space that is consistent across tasks can be allocated to a task such that the virtual blocks allocated to different tasks do not share overlapping logical address ranges; a logical base address in a set of logical addresses and an offset relative to the base address that is used to refer to a location in the virtual block can be allocated to a task; or a logical base address that is a zero offset can be allocated.
[0163] The translation unit responds to an access request (e.g., read / write) received from a task via the I / O arbiter 703. Each access request can include an identifier of the task and an offset or logical address that identifies the region of the virtual block that it needs to access. When an access request is received from a task via the I / O arbiter 703, the translation unit is configured to identify the desired physical address in the data structure based on the identifier of the working unit and the offset / logical address. The offset / logical address is used to identify which memory portions associated with the task will be accessed.
[0164] For example, again consider the case where a task requests access to Figure 8 the storage area in partition 'D' of the virtual block 801 shown. The translation unit will receive from the task an access request that includes the identifier of the task and an offset 803 that points to a location in partition 'D'. Using the identifier of the task and the offset, the translation unit looks up in the data store 706 the physical address of the corresponding memory portion 811 in the shared memory (i.e., the physical address in the data store that is associated with this task and the offset). The translation unit can then proceed to service the access request in the corresponding memory portion 811 (e.g., by performing the requested read / write to / from the memory portion 811).
[0165] It will be appreciated that the "fragmentation-insensitive" method described herein allows the underlying shared memory 702 to become fragmented while the tasks themselves receive contiguous memory blocks. In the case where there are a sufficient number of memory portions at the shared memory, this enables memory blocks to be allocated to tasks and work groups that include multiple tasks even when there is not enough contiguous space in the shared memory to allocate a contiguous block of the required size.
[0166] Figure 10A flowchart showing a resource allocator allocating memory resources to tasks is shown. When a request 1001 for shared memory resources for a task is received, a contiguous block of virtual memory 1002 is allocated for the task. The virtual memory block is associated 1003 with a physical memory address of the shared memory such that each part of the virtual block corresponds to a lower-level part of the physical shared memory. Once a task receives its allocated virtual memory, the task can be executed 1004 on a SIMD processor, where the task uses the lower-level shared memory resources at the physical address corresponding to its allocated virtual memory block. Figure 2 , 3 The memory subsystem of FIGS. 6 and 7 is shown as including a plurality of functional blocks. This is merely illustrative and is not intended to define a strict partitioning between the different logic elements of these entities. Each functional block can be provided in any suitable manner. It will be understood that the intermediate values formed by each functional block described herein need not be physically generated by the functional block at any time, but can merely represent logical values that conventionally describe the processing performed by the functional block between its inputs and outputs.
[0167] The memory subsystem described herein can be implemented in hardware on an integrated circuit. The memory subsystem described herein can be configured to perform any of the methods described herein. In general, any of the functions, methods, techniques, or components described above can be implemented in software, firmware, hardware (e.g., fixed logic circuitry), or any combination thereof. The terms "module," "function," "component," "element," "unit," "block," and "logic" can be used herein generally to represent software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, unit, block, or logic represents program code that, when executed on a processor, performs a specified task. The algorithms and methods described herein can be executed by one or more processors executing code that causes the one or more processors to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disks, flash memory, hard disk memory, and other memory devices that can use magnetic, optical, or other technologies to store instructions or other data that can be accessed by a machine.
[0168] The terms "computer program code" and "computer-readable instructions" are used herein to refer to any kind of executable code for a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, byte code, code defining an integrated circuit (e.g., a hardware description language or a netlist), and code expressed in a programming language code (e.g., C, Java, or OpenCL). The executable code can be, for example, any kind of software, hardware, script, module, or library that, when properly executed, processed, interpreted, assembled, or executed in a virtual machine or other software environment, causes a processor of a computer system supporting the executable code to perform the tasks specified by the code.
[0169] A processor, computer, or computer system can be any kind of device, machine, or dedicated circuit having processing capabilities such that it can execute instructions, or a collection or portion thereof. The processor can be any kind of general-purpose or special-purpose processor, e.g., a CPU, GPU, system-on-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system can include one or more processors.
[0170] It is also desirable to cover software that defines the hardware configurations described herein, e.g., HDL (hardware description language) software for designing integrated circuits or for configuring programmable chips to implement desired functions. That is, a computer-readable storage medium having encoded thereon computer-readable program code in the form of an integrated circuit definition data set can be provided, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture a memory subsystem configured to execute any of the methods described herein, or to manufacture a memory subsystem including any of the devices described herein. The integrated circuit definition data set can be, for example, an integrated circuit description.
[0171] A method of manufacturing a memory subsystem described herein in an integrated circuit manufacturing system is provided. An integrated circuit definition data set can be provided that, when processed by the integrated circuit manufacturing system, causes the method of manufacturing the memory subsystem to be executed.
[0172] An integrated circuit definition dataset can be in the form of computer code, e.g., as a netlist, code for configuring a programmable chip, as a hardware description language (including register transfer level (RTL) code) that defines an integrated circuit at any level, as a high-level circuit representation (e.g., Verilog or VHDL), and as a low-level circuit representation (e.g., OASIS(RTM) and GDSII). A computer system configured to generate a manufacturing definition of an integrated circuit in the context of a software environment can process a higher-level representation (e.g., RTL) that logically defines the integrated circuit, where the manufacturing definition includes definitions of circuit elements and rules for joining these elements to generate a manufacturing definition representing the defined integrated circuit. In the general case where a computer system executes software to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to cause the computer system configured to generate a manufacturing definition of an integrated circuit to execute the execution code that defines the integrated circuit and thereby generate the manufacturing definition of the integrated circuit.
[0173] Now referring to Figure 11 , an example of an integrated circuit manufacturing system processing an integrated circuit definition dataset to configure the system to manufacture a memory subsystem is described.
[0174] Figure 11 An example of an integrated circuit (IC) manufacturing system 1102 that can be configured to manufacture the memory subsystem described in any of the examples herein is shown. Specifically, the IC manufacturing system 1102 includes a layout processing system 1104 and an integrated circuit generation system 1106. The IC manufacturing system 1102 is configured to receive an IC definition dataset (e.g., defining the memory subsystem described in any of the examples herein), process the IC definition dataset, and generate an IC according to the IC definition dataset (e.g., an IC definition dataset embodying the memory subsystem described in any of the examples herein). The processing of the IC definition dataset configures the IC manufacturing system 1102 to manufacture an integrated circuit embodying the memory subsystem described in any of the examples herein.
[0175] The layout processing system 1104 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout from an IC definition data set are known in the art and may include, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated (e.g., in terms of logic components such as NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). A circuit layout may be determined from the gate-level representation of the circuit by determining location information for the logic components. This may be achieved automatically or under the control of a user to optimize the circuit layout. When the layout processing system 1104 has determined a circuit layout, it may output a circuit layout definition to the IC generation system 1106. The circuit layout definition may be, for example, a circuit layout description.
[0176] The IC generation system 1106 generates an IC from the circuit layout definition, as is known in the art. For example, the IC generation system 1106 may perform a semiconductor device assembly process to generate the IC, which may include a multi-step sequential lithography process and chemical processing steps during which an electronic circuit is gradually formed on a wafer made of semiconductor material. The circuit layout definition may be in the form of a mask, which may be used in the lithography process for generating an IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 1106 may be in the form of computer-readable code, and the IC generation system 1106 may use the computer-readable code to form an appropriate mask for generating the IC.
[0177] The different processes that may be performed by the IC manufacturing system 1102 may all be at one location, e.g., implemented by one party. Alternatively, the IC manufacturing system 1102 may be a distributed system, so that some processes may be performed at different locations and may be performed by different parties. For example, some of the following stages may be performed at different locations and / or by different parties: (i) synthesizing RTL code representing the IC definition data set to form a gate-level representation of the circuit to be generated, (ii) generating a circuit layout based on the gate-level representation, (iii) forming a mask according to the circuit layout, and (iv) assembling an integrated circuit using the mask.
[0178] In other examples, processing of an IC definition data set at an integrated circuit manufacturing system may configure the system to manufacture a memory subsystem without processing the IC definition data set to determine a circuit layout. For example, the IC definition data set may define the configuration of a reconfigurable processor (e.g., an FPGA), and processing of the data set may configure the IC manufacturing system to generate a reconfigurable processor with the defined configuration (e.g., by loading configuration data into the FPGA).
[0179] In some embodiments, when the integrated circuit manufacturing definition dataset is processed in an integrated circuit manufacturing system, the integrated circuit manufacturing system can generate the devices described herein. For example, through the integrated circuit manufacturing definition dataset, the configuration of the integrated circuit manufacturing system in the manner described for Figure 11 can cause the devices described herein to be manufactured.
[0180] In some examples, the integrated circuit definition dataset can include software that runs on the hardware defined at the dataset or interfaces with the hardware defined at the dataset. In Figure 11 the example shown, the IC generation system can be further configured by the integrated circuit definition dataset to load the hardware into the integrated circuit or provide the program code for the integrated circuit to use according to the program code defined at the integrated circuit definition dataset when manufacturing the integrated circuit.
[0181] The graphics processing system described herein can be implemented in hardware on an integrated circuit. The graphics processing system described herein can be configured to perform any of the methods described herein.
[0182] The implementation of the concepts set forth in this application in devices, apparatuses, modules, and / or systems (and the methods implemented herein) can result in performance improvements compared to known implementations. Performance improvements can include one or more of increased computing performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of these devices, apparatuses, modules, and systems (e.g., in an integrated circuit), a balance can be struck between performance improvements and physical implementation to improve the manufacturing method. For example, a balance can be struck between performance improvements and layout area to match the performance of known implementations but use less silicon. This can be achieved, for example, by reusing functional blocks in a sequential manner or sharing functional blocks among the elements of the device, apparatus, module, and / or system. Conversely, the concepts of this application that result in improvements in the physical implementation of the device, apparatus, module, and system (e.g., reduced silicon area) can be traded for improved performance. This can be achieved, for example, by manufacturing multiple instances of the module within a predetermined area budget.
[0183] The applicant separately discloses each individual feature described herein, as well as any combination of two or more such features, such that these features or combinations can be implemented based on the common general knowledge of those skilled in the art from the specification, regardless of whether these features or feature combinations solve any of the problems disclosed herein. In view of the above description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units for processing SIMD tasks, the memory subsystem comprising: A shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor; A translation unit configured to associate a task with one or more physical addresses of the shared memory; And A resource allocator configured to, in response to receiving a memory resource request for a first memory resource for a task, allocate a contiguous virtual memory block to the task and cause the translation unit to associate the task with a plurality of physical addresses of memory portions of a partition corresponding to the contiguous virtual memory block, the memory portions together implementing the complete contiguous virtual memory block, wherein the contiguous virtual memory block includes a base address that is one of the physical addresses of the memory portions associated with the task.
2. The memory subsystem as claimed in claim 1, wherein, At least some of the memory portions of the contiguous virtual memory block are non - contiguous in the shared memory.
3. The memory subsystem according to claim 1 or 2, wherein, The translation unit is operable to subsequently service an access request received from the task regarding the contiguous virtual memory block, the access request including an identifier of the task and an offset of a storage area in the contiguous virtual memory block, the translation unit being configured to service the access request in a memory portion of a partition corresponding to the contiguous virtual memory block indicated by the offset.
4. The memory subsystem according to claim 1 or 2, wherein The translation unit includes: a content - addressable memory configured to, in response to receiving an identifier of an item and an offset in the contiguous virtual memory block, return one or more corresponding physical addresses.
5. The memory subsystem according to claim 1 or 2, wherein The SIMD processor is configured to process one or more workgroups, each workgroup including a plurality of SIMD tasks, the task being a first - received task of the workgroup, and the resource allocator is configured to reserve sufficient memory portions for the workgroup such that each task of the workgroup can receive a memory resource equivalent to the first memory resource, and allocate the requested memory resource to the first - received task from the memory portions reserved for the workgroup.
6. The memory subsystem according to claim 5, wherein, The resource allocator is configured to allocate a contiguous virtual memory block to each task of the workgroup such that the contiguous virtual memory blocks allocated to the tasks of the workgroup together represent the virtual memory blocks of a contiguous superblock.
7. The memory subsystem of claim 6, wherein, The resource allocator is configured to, in response to subsequently receiving a memory resource request for a second task of the workgroup, allocate a contiguous virtual memory block to the second task from the virtual memory blocks of the contiguous superblock.
8. The memory subsystem according to claim 5, wherein, The resource allocator maintains a data structure that identifies which of the one or more workgroups are currently allocated contiguous virtual memory blocks.
9. The memory subsystem according to claim 8, wherein, The data structure includes a fine - grained status array arranged to indicate whether each memory portion of the shared memory is allocated to a task.
10. A method for allocating shared memory resources to tasks to be executed in a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units respectively configured to process SIMD tasks, the method including: Receiving a memory resource request for a first memory resource for a task; Allocating a contiguous virtual memory block to the task; And Associating the task with a plurality of physical addresses of a memory portion of a shared memory, each memory portion corresponding to a partition of the contiguous virtual memory block, the memory portions together implementing the complete contiguous virtual memory block, wherein the contiguous virtual memory block includes a base address that is one of the physical addresses of the memory portions associated with the task.
11. The method according to claim 10, wherein, At least some of the memory portions of the contiguous virtual memory block are non - contiguous in the shared memory.
12. The method according to claim 10 or 11, further including: Subsequently receiving an access request received from the task for the contiguous virtual memory block, the access request including an identifier of the task and an offset of a storage area in the contiguous virtual memory block; And Servicing the access request by accessing a memory portion corresponding to the partition of the contiguous virtual memory block indicated by the offset.
13. The method according to claim 10 or 11, wherein Tasks are grouped together in a workgroup for execution at the SIMD processor, the task is the first received task of the workgroup, and allocating a contiguous virtual memory block to the task includes: Reserving sufficient memory portions for the workgroup for each task of the workgroup to receive a memory resource equivalent to the first memory resource; and Allocating the requested memory resource to the first received task from the memory portions reserved for the workgroup.
14. The method according to claim 13, wherein, Reserving for the workgroup includes: reserving a contiguous virtual memory block for each task of the workgroup such that the contiguous virtual memory blocks allocated to the tasks of the workgroup together represent a virtual memory block of a contiguous superblock.
15. The method according to claim 14, further comprising: In response to subsequently receiving a memory resource request for a second task of the workgroup, allocating a contiguous virtual memory block to the second task from the virtual memory block of the contiguous superblock.
16. The method according to claim 10 or 11, further comprising: Maintaining a data structure that identifies which of the one or more workgroups are currently allocated contiguous virtual memory blocks.
17. The method according to claim 16, wherein, The data structure includes a fine - grained status array arranged to indicate whether each memory portion of the shared memory is allocated to the task.
18. A non - transitory computer - readable storage medium having stored thereon a computer - readable description of an integrated circuit, the computer - readable description when processed in an integrated circuit manufacturing system causes the integrated circuit manufacturing system to manufacture the memory subsystem of claim 1 or claim 2.
19. An integrated circuit manufacturing system, including: A non - transitory computer - readable storage medium having stored thereon a computer - readable integrated circuit description of the memory subsystem of claim 1 or claim 2; A layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit implementing the memory system; and an integrated circuit generation system configured to fabricate the memory subsystem according to the circuit layout description.
20. A non-transitory computer-readable storage medium having stored thereon computer-readable instructions that, when executed at a computer system, cause the computer system to perform the method of claim 10 or claim 11.