Resource Allocation
By designing a dynamic memory allocation mechanism in a multi-processing unit system, the processing deadlock problem caused by competition for shared memory resources is solved, and more efficient memory resource utilization and task processing order are achieved.
Patent Information
- Application Number
- CN201811083351.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-15
- Filing Date
- 2018-09-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-12-30
AI Technical Summary
In a multi-processing unit system, the dynamic allocation mechanism of shared memory is difficult to effectively solve the competition between memory resources between processing tasks, resulting in processing deadlocks and resource waste.
A memory subsystem is designed, including shared memory and resource allocator. The resource allocator can respond to the memory request of the task, dynamically allocate the memory resources of continuous blocks, and prioritize the tasks of the allocated memory blocks to avoid resource waste.
Through this memory subsystem, deadlocks can be effectively avoided, memory resource utilization can be improved, tasks can be processed in a predetermined order, and system delay can be reduced.
Smart Images

Figure CN109656710B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a memory subsystem for multiple processing units. Background Art
[0002] In a system including multiple processing units, a shared memory is typically provided that can be accessed by at least some of the processing units. Due to the range of processing tasks that can be run on a processing unit, each with different memory requirements, the shared memory resources are generally not fixed for each processing unit at design time. An allocation mechanism is typically provided that allows processing tasks running on different processing units to request one or more regions of shared memory respectively. This enables shared memory to be dynamically allocated for use by tasks executed at a processing unit.
[0003] Efficient use of shared memory can be achieved through careful design of the allocation mechanism. Summary of the invention
[0004] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0005] A memory subsystem is provided for a single instruction multiple data (SIMD) processor including a plurality of processing units, the plurality of processing units being configured to process one or more work groups each including a plurality of SIMD tasks, the memory subsystem comprising:
[0006] a shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor; and
[0007] A resource allocator is configured to allocate a memory portion to the work group in response to receiving a memory resource request for a first memory resource from a first receiving task of the work group, the memory portion having a size sufficient for each task in the work group to receive a memory resource in the memory portion equivalent to the first memory resource.
[0008] The resource allocator may be configured to allocate a contiguous block of memory resources as the block memory portion.
[0009] The resource allocator may be configured to allocate the requested first memory resource from the block memory portion to a first received task of the work group when servicing the task, and reserve the remaining memory portion of the block memory portion from allocation to tasks of other work groups.
[0010] The resource allocator may be configured to allocate the memory resource of the block memory portion to the second task in response to subsequently receiving a memory resource request from the second task of the associated work group.
[0011] The resource allocator may be arranged to receive memory resource requests from a plurality of different requestors and, in response to allocating the block memory portion to the workgroup, prioritise servicing memory requests received from the requestor from which the first received task for the workgroup was received.
[0012] The resource allocator may be further configured to, in response to receiving an indication that processing of a task of the workgroup has completed, release the memory resources allocated to the task without waiting for the processing of the workgroup to complete.
[0013] The shared memory may be further divided into a plurality of non-overlapping windows, each non-overlapping window comprising a plurality of memory portions, the resource allocator being configured to maintain a window pointer indicating a current window in which to attempt to allocate a memory portion in response to a next received memory request.
[0014] The resource allocator may be implemented in a binary logic circuit, and the window length may be such that the availability of all memory portions of each window may be checked in a single clock cycle of the binary logic circuit.
[0015] The resource allocator may further be configured to maintain a fine state array arranged to indicate whether each memory portion of the shared memory is allocated to a task.
[0016] The resource allocator may be configured to, in response to receiving a memory resource request for a first receiving task associated with a work group, search in a current window for a portion of memory that is indicated by the fine status array as a contiguous block available for allocation, and the resource allocator may be configured to allocate such a contiguous block to the work group if such a contiguous block is identified in the current window.
[0017] The resource allocator may be configured to allocate the memory portion of a contiguous block such that the block starts at the lowest possible position in the window.
[0018] The resource allocator may be further configured to maintain a coarse state array, the coarse state array being arranged to indicate, for each window of the shared memory, whether all memory portions of the window are unallocated, the resource allocator being configured to check the coarse state array in parallel with searching for memory portions of a contiguous block in a current window to determine whether the size of the requested block can be provided by one or more subsequent windows; the resource allocator being configured to allocate to the work group a block starting from a first memory portion of the current window that is in a contiguous block with one or more subsequent windows and extending to memory portions in these subsequent windows if a sufficiently large contiguous block cannot be identified in the current window and the requested block can be provided by one or more subsequent windows.
[0019] The resource allocator may be further configured to form an overflow metric in parallel with searching the current window, the overflow metric representing memory resources of a memory portion of a required block that cannot be provided in the current window starting from the first memory portion of the current window in an unallocated memory portion of a contiguous block adjacent to a subsequent window, the resource allocator being configured to, if a sufficiently large contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then next attempt to allocate a memory portion to the work group by searching in the subsequent window starting from the first memory portion of the subsequent window for an unallocated memory portion of a contiguous block whose total size is sufficient to provide the overflow metric.
[0020] The fine status array may be a bit array in which each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated.
[0021] The coarse state array may be a bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is completely allocated.
[0022] The resource allocator may be configured to form each bit of the coarse state array by performing an OR reduction on all bits of the fine state array corresponding to a memory portion present in a window corresponding to the bit in the coarse state array.
[0023] The window length can be a power of 2.
[0024] The resource allocator may maintain a data structure identifying which workgroups of the one or more workgroups are currently being allocated portions of blocks of memory.
[0025] According to a second aspect, there is provided a method of allocating shared memory resources to tasks executed in a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units, each processing unit being configured to process one or more work groups respectively comprising a plurality of SIMD tasks, the method comprising:
[0026] receiving a shared memory resource request for a first memory resource from a first receiving task of the associated working group; and
[0027] A block memory portion of the shared memory is allocated to the work group, the block memory portion being large enough for each task in the work group to receive memory resources in the block memory portion equivalent to the first memory resources.
[0028] Allocating a block memory portion to a work group may include allocating the block memory portion as a contiguous block of memory portions.
[0029] The method may further include allocating the requested first memory resource from the block memory portion to a first received task of the work group when servicing the task, and reserving a remaining memory portion of the block memory portion from allocation to tasks of other work groups.
[0030] The method may further include: in response to subsequently receiving a memory resource request from a second task of the associated work group, allocating memory resources of the block memory portion to the second task.
[0031] The method may also include receiving memory resource requests from a plurality of different requestors, and in response to allocating the block memory portion to the work group, prioritizing servicing memory requests received from the requestor from which the first received task of the work group was received.
[0032] The method may further include, in response to receiving an indication that processing of a task of the work group has completed, releasing memory resources allocated to the task without waiting for processing of the work group to complete.
[0033] The method may further comprise maintaining a fine status array arranged to indicate whether each memory portion of the shared memory is allocated to a task.
[0034] Allocating a portion of memory to the work group may include searching the current window for a portion of memory indicated by the fine status array as a contiguous block available for allocation, and allocating the contiguous block to the work group if such a contiguous block is identified in the current window.
[0035] Allocating the contiguous block to the work group may include allocating a memory portion of the contiguous block such that the contiguous block starts at a lowest possible position in the window.
[0036] The method may further comprise:
[0037] maintaining a coarse state array indicating, for each window of shared memory, whether all memory portions of the window are allocated;
[0038] In parallel with searching the memory portion for successive blocks in the current window, checking the coarse state array to determine whether the requested block size can be accommodated by one or more subsequent windows; and
[0039] If a sufficiently large contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, a block is allocated to the work group starting from the first memory portion of the current window that is in a contiguous block with the subsequent windows and extending to the memory portions in these subsequent blocks.
[0040] The method may further include:
[0041] In parallel with searching the current window, forming an overflow metric representing memory resources of memory portions of a required block that cannot be accommodated in the current window starting from a first memory portion of the current window in unallocated memory portions of consecutive blocks of an immediately subsequent window; and
[0042] If a sufficiently large contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then a subsequent attempt is made to allocate a memory portion to the work group by searching in the subsequent window, starting with the first memory portion of the subsequent window, for an unallocated memory portion having a total size sufficient to accommodate the contiguous block of the overflow metric.
[0043] According to a third aspect, there is provided a memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units for processing SIMD tasks, the memory subsystem comprising:
[0044] a shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor;
[0045] a translation unit configured to associate the task with one or more physical addresses of the shared memory; and
[0046] A resource allocator is configured to, in response to receiving a memory resource request for a first memory resource related to a task, allocate a continuous virtual memory block to the task and cause a conversion unit to associate the task with a plurality of physical addresses of memory portions that respectively correspond to partitions of the virtual memory block, the memory portions together implementing a complete virtual memory block.
[0047] At least some of the plurality of memory portions of the block may be non-contiguous in the shared memory.
[0048] The virtual memory block may include a base address that is a physical address of one of the memory portions associated with the task.
[0049] The virtual memory block may include a base address that is a logical address of the virtual memory block.
[0050] The translation unit is operable to subsequently service an access request received from a task regarding the virtual memory block, the access request comprising an identifier of the task and an offset to a storage area in the virtual memory block, the translation unit being configured to service the access request at a memory portion of the partition corresponding to the virtual memory block indicated by the offset.
[0051] The translation unit may include a content addressable memory configured to, in response to receiving an identifier of an item and an offset in a virtual memory block, return one or more corresponding physical addresses.
[0052] The SIMD processor can be configured to process a workgroup including multiple tasks, the task being a first receiving task of the workgroup, and the resource allocator is configured to reserve a sufficient memory portion for the workgroup to enable each task of the workgroup to receive memory resources equivalent to the first memory resources, and to allocate the requested memory resources to the first receiving task from the memory portion reserved for the workgroup.
[0053] The resource allocator may be configured to allocate a virtual memory block to each task of the work group such that the virtual memory blocks allocated to the tasks of the work group collectively represent a virtual memory block of a contiguous superblock.
[0054] The resource allocator may be configured to allocate a contiguous virtual memory block from the superblock to the second task in response to subsequently receiving a request for a memory resource from the second task of the work group.
[0055] According to a fourth aspect, there is provided a method for allocating shared memory resources to tasks executed in a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units respectively configured to process SIMD tasks, the method comprising:
[0056] receiving a memory resource request for a first memory resource associated with a task;
[0057] allocating a contiguous block of virtual memory to the task; and
[0058] Tasks are associated with a plurality of physical addresses of memory portions of a shared memory, each memory portion corresponding to a partition of a virtual memory block, the memory portions collectively implementing a complete virtual memory block.
[0059] The method may further comprise:
[0060] subsequently receiving an access request received from a task regarding the virtual memory block, the access request including an identifier of the task and an offset of a storage region in the virtual memory block; and
[0061] The access request is serviced by accessing the memory portion of the partition corresponding to the virtual memory block indicated by the offset.
[0062] The tasks may be grouped together in a workgroup for execution at a SIMD processor, the task may be a first received task of the workgroup, and allocating a contiguous virtual memory block to the task may include:
[0063] reserving a sufficient portion of memory for the work group for each task of the work group to receive memory resources equivalent to the first memory resources; and
[0064] The requested memory resources are allocated to the first receiving task from the portion of memory reserved for the working group.
[0065] Reserving for the workgroup may include reserving a virtual memory block for each task of the workgroup such that the virtual memory blocks assigned to the tasks of the workgroup collectively represent a virtual memory block of a contiguous superblock.
[0066] The method may further include: in response to subsequently receiving a memory resource request from a second task of the associated work group, allocating a contiguous virtual memory block from the superblock to the second task.
[0067] The memory subsystem may be implemented in hardware on an integrated circuit.
[0068] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided.
[0069] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided, the method comprising:
[0070] processing the computer-readable description of the graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and
[0071] A graphics processing system is fabricated from the circuit layout description using an integrated circuit generation system.
[0072] An integrated circuit definition data set is provided that, when processed in an integrated circuit manufacturing system, configures the system to manufacture the memory subsystem described herein.
[0073] A non-transitory computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, which when processed in an integrated circuit manufacturing system causes the integrated circuit manufacturing system to manufacture the memory subsystem described herein.
[0074] A non-transitory computer-readable storage medium is provided having stored thereon a computer-readable description of a memory subsystem described herein, which when processed in an integrated circuit manufacturing system causes the integrated circuit manufacturing system to:
[0075] processing the computer-readable description of the graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and
[0076] A graphics processing system is fabricated from the circuit layout description using an integrated circuit generation system.
[0077] An integrated circuit manufacturing system is provided that is configured to manufacture the memory subsystem described herein.
[0078] An integrated circuit manufacturing system is provided, comprising:
[0079] a non-transitory computer-readable storage medium having stored thereon a computer-readable integrated circuit description describing the memory subsystem described herein;
[0080] a layout processing system configured to process the integrated circuit description to generate a circuit layout description of the integrated circuit implementing the memory subsystem; and
[0081] An integrated circuit generation system is configured to manufacture a memory subsystem based on a circuit layout description.
[0082] A memory subsystem is provided, configured to perform the method described herein. A computer program code is provided for performing the method described herein. A non-transitory computer readable storage medium is provided, having computer readable instructions stored thereon, which when executed at a computer system causes the computer system to perform the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0084] Figure 1 A conventional allocation of shared memory for tasks executed at a SIMD processor is shown.
[0085] Figure 2 is a schematic diagram of a first memory subsystem including a resource allocator configured according to the principles described herein.
[0086] Figure 3The memory subsystem is shown in the context of a computer system having multiple processing cores.
[0087] Figure 4 A shared memory is shown including blocks that are allocated in slices to tasks of a work group.
[0088] Figure 5 The allocation of shared memory by a resource allocator to tasks executed as work groups at a SIMD processor is shown.
[0089] Figure 6 Allocation of shared memory by a resource allocator configured in accordance with certain embodiments described herein is shown.
[0090] Figure 7 is a schematic diagram of a second memory subsystem including a resource allocator configured according to the principles described herein.
[0091] Figure 8 The correspondence between a virtual memory block and the memory portion of the underlying shared memory is shown.
[0092] Fig. 9 is a flow chart illustrating the operation of a first memory subsystem configured according to the principles described herein.
[0093] Fig.10 is a flow chart illustrating the operation of a second memory subsystem configured according to the principles described herein.
[0094] Fig.11 It is a schematic diagram of the integrated circuit manufacturing system. DETAILED DESCRIPTION
[0095] The following description is given by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art. The embodiments are described only by way of example.
[0096] The term "task" is used herein to refer to a set of data items and work to be performed on these data items. For example, in addition to a set of data to be processed according to a program (e.g., the same sequence of ALU instructions or references to these ALU instructions), a task may also include or be associated with a program or may also include a reference to a program, wherein the set of data may include one or more data elements (or data items, such as a plurality of pixels or vertices).
[0097] The term "program instance" is used here to refer to an individual instance that walks through the code. Therefore, a program instance refers to a single data item and a reference (e.g., a pointer) to the program to be executed on the data item. Therefore, it can be considered that a task includes multiple program instances (e.g., up to 32 program instances), although in practice each task only requires a single instance of a common program (or reference). Task groups that share a common purpose share local memory and can execute the same program (although they can execute different parts of the program) or compatible programs on different data blocks can be connected by a group ID. The task group with the same group ID can be referred to as a "workgroup" (therefore, the group ID can be referred to as a "workgroup ID"). Therefore, there is a hierarchy of terms, in which a task includes multiple program instances, and a group (or workgroup) includes multiple tasks.
[0098] Some processing systems that provide shared memory may include one or more SIMD (single instruction multiple data) processors, each SIMD processor being configured to execute the tasks of a work group in parallel. Each task of a work group may need to be allocated a shared memory. However, generally, there is a certain interdependency between the processing of the tasks executed at the SIMD processor. For example, for some work groups, a barrier synchronization point that all tasks must reach is defined, so as to continue to process any of these tasks across the barrier synchronization point and complete its processing of the work group as a whole. Sometimes more than one barrier synchronization point can be defined for a work group, wherein all tasks need to reach a given barrier, so as to continue to process any of these tasks across the barrier.
[0099] For traditional allocation mechanisms of shared memory, the use of barrier synchronization points can cause the processing of work groups to deadlock. Figure 1 A simple example is shown in which Figure 1It is shown that a shared memory 100 is divided into a plurality of memory portions 101 and a SIMD processor 102 having five processing elements 109-113 is configured to process a work group 'A' comprising five tasks 103-107 in parallel. In this example, processing of only four tasks has been started, and all of these tasks have reached the barrier synchronization point 108. This is because each task requires two memory portions (marked with the letter 'A') of consecutive blocks in the shared memory, but once the first four tasks have been allocated their blocks in the shared memory, it is not possible to allocate two memory portions of consecutive blocks for the fifth task. The memory portions belonging to the second work group are marked with the letter 'B'. Several memory portions 114 in the shared memory are available, but these memory portions are fragmented and cannot provide consecutive blocks of memory. As a result, a deadlock has occurred: the processing of the four tasks 103-106 has stopped, waiting for the fifth task 107 to reach the barrier synchronization point 108, but the processing of the fifth task 107 cannot be started because its request to allocate shared memory cannot be satisfied.
[0100] Such deadlocks can cause significant delays and are generally resolved only when the processing of the deadlocked workgroup reaches a predetermined timeout. Not only is the time spent waiting for the timeout wasted, but partial processing is performed on the task that reached the barrier synchronization point (which will need to be repeated). Deadlock situations are particularly problematic in systems that include multiple SIMD processors and all of which share common memory, because more than one workgroup can be deadlocked at a time when the available shared memory is limited or when tasks require relatively large contiguous blocks of memory.
[0101] Avoid fragmentation
[0102] Figure 22 is a schematic diagram showing a memory subsystem 200 configured to allocate shared memory in a manner that solves the above-mentioned problems. The memory subsystem 200 includes a resource allocator 201 and a shared memory 202, and the shared memory 202 is divided into a plurality of memory portions 203. The resource allocator can be configured to receive memory requests from a plurality of requesters 205, for example, each task running at a processing unit can be a requester, each processing unit can be a requester, or each type of processing unit can be a requester. Different requesters can have different hard-wired inputs to the resource allocator, and the allocator needs to arbitrate between these hard-wired inputs. The resource allocator can serve the memory requests from the requesters based on any appropriate manner, for example, the memory requests can be served based on simple polling or according to a set of arbitration rules for selecting the next requester to be served. In other embodiments, the memory requests from the plurality of requesters can be gathered at a request queue, wherein the memory allocation request is received in the request queue, and the resource allocator can receive the memory request that needs to be served from the request queue.
[0103] The processing units may access the allocated portions of the shared memory through the input / output arbiter 204. The I / O arbiter 204 arbitrates access to the shared memory between multiple units (e.g., processing units or other units capable of accessing the shared memory) that submit access requests (e.g., read / write) through the interface 211, wherein the interface 211 may include one or more hard connections between which the I / O arbiter arbitrates.
[0104] Figure 3 The memory subsystem 200 is shown in the context of a computer system 300 including a plurality of processing units 301 and a work scheduler 302, wherein the work scheduler 302 is configured to schedule tasks received by it for processing at the processing units. One or more processing units may be SIMD processors. The computer system may be, for example, a graphics processing unit (GPU), and the processing unit may include one or more units (each of which may be SIMD processors) for performing integer operations, floating point operations, complex arithmetic, texture address calculations, sampling operations, etc. The processing units may collectively form a set of parallel pipelines 305, wherein each processing unit represents one pipeline in the set of pipelines. The computer system may be, for example, a vector processor.
[0105] Each processing unit may include one or more arithmetic logic units (ALUs), and each SIMD processor may include one or more different types of ALUs, for example, each type of ALU is optimized for a specific type of calculation. In the example where a GPU provides a processing unit 301, a processing block 104 may include multiple shader cores, each of which includes one or more ALUs.
[0106] The work scheduler 302 may include a scheduler unit 303 configured to perform scheduling of tasks for execution on the processing unit 301 and an instruction decoder 304 configured to decode the tasks into a form suitable for execution on the processing unit 301. It will be appreciated that the specific decoding of the tasks performed by the instruction decoder will be determined by the type and configuration of the processing unit 301.
[0107] The resource allocator 201 is operable to receive requests to allocate memory resources to tasks executed at a processing unit 301 of a computer system. Each task or processing unit may be a requestor 205. One or more processing units may be SIMD processors configured to execute tasks of a work group in parallel. Each task of a work group may request the resource allocator to allocate memory (depending on the specific architecture of the system), such requests may be made on behalf of the task by, for example, a work scheduler when assigning tasks to parallel processors or by the task itself.
[0108] In some architectures, tasks associated with various workgroups or for execution on different types of processing units can be queued for processing in one or more common work queues, so that tasks used for parallel processing as a workgroup are assigned to appropriate processing units by the work scheduler 302 during a certain period of time. Tasks associated with other workgroups and / or processing units can be interspersed with the tasks of the workgroup. Requests for memory resources of the resource allocator can generally be made for tasks because tasks are scheduled at the processor. The tasks themselves may not need to make requests for memory resources. For example, in some architectures, the work scheduler 302 can make requests for memory resources on behalf of the tasks it schedules at the processing unit. Many other architectures are possible.
[0109] The resource allocator is configured to allocate memory resources to the entire work group when an allocation request for a task of the work group is received for the first time. The allocation request generally indicates the size of the required memory resources, for example, the task may require a certain number of bytes or memory portions. Since the SIMD processor executes the same instruction in parallel on different source data, each task of the work group executed at the SIMD processor has the same memory requirement. The resource allocator is configured to allocate a continuous block of memory for the entire work group in response to the first request for memory received by the task of the work group. The size of the block is at least N times the memory resource requested in the first request, where N is the number of tasks in the work group (e.g., 32). The number of tasks in the work group can be fixed for a given system, so it is known in advance for the resource allocator, and the number of tasks in the work group can be provided to the resource allocator together with the first request for memory resources, and / or the number of tasks in the work group is available to the resource allocator (e.g., as a parameter maintained at a data storage area accessible to the resource allocator).
[0110] If there is enough continuous space in the shared memory for the entire block of workflow, the resource allocator is configured to respond to the first received request for memory resources for the task using the memory resources requested by the task of the workflow. This starts the processing of the first task of the work group. The resource allocator also reserves the remainder of the shared memory block (which will be allocated to other tasks of the work group) to prevent the block from being allocated to tasks of other work groups. If there is not enough continuous space in the shared memory for the entire block of workflow, the first received memory request is rejected. This process ensures that once the first task of the work group receives an allocation of memory, the system can ensure that all tasks of the work group will receive their memory allocations. This avoids the above-mentioned deadlock scenario.
[0111] When the resource allocator receives requests for memory resources for subsequently received tasks of the same workgroup, the resource allocator allocates the required resources from the reserved block to the tasks. The resource allocator can then sequentially allocate adjacent "slices" of the reserved shared memory block to the tasks in the order in which the corresponding memory requests were received, with the first receiving task receiving the first slice of the block. This Figure 4 It is shown in Figure 4The shared memory 202 is shown to include a plurality of memory portions 203 and a block 402 reserved for a work group after receiving a first memory request for a task of the work group. In this example, each slice includes two memory portions. The first slice 403 of the block is allocated to the first task, and subsequent slices of the block are retained in the shared memory so that when the next memory request for a task of the work group is received, the next slice 404 is allocated, and so on, until slices 403-407 of the block are allocated to tasks and the work group can be processed to completion.
[0112] In the case where the tasks of the work group have an expected or predetermined order, the resource allocator may instead allocate memory slices of the reserved memory block to the tasks of the work group according to the order. In some architectures, memory requests for tasks may be received out of order. For example, even if a memory request for the third task is received before the memory request for the second task, the third slice of the block may be allocated to task number three of the five tasks. In this way, a predetermined slice of the block reserved for the work group may be allocated to each task of the work group. Since the first receiving task may not be the first task of the work group, it is not necessary to allocate the first slice of the block reserved for the work group to the first task.
[0113] In some architectures, tasks may make further requests for memory during processing. These memory requests may be handled in the same manner as the initial request for memory, where, when a first request is received for a workgroup, a block of memory resources is allocated for all tasks of the workgroup, based on the fact that each task for processing on the SIMD processor will require the same memory resources. The size of the block allocated will be at least N times the memory resources requested in the first request, where N is the number of tasks in the workgroup.
[0114] Once a task has been scheduled at a processing unit, the task may access its allocated memory resources at shared memory 202 (eg, by performing reads and / or writes) as indicated by data path 210 .
[0115] The resource allocator can be configured to release the slice of shared memory allocated to each task immediately upon completion of processing at the processing unit (i.e., without waiting until the entire work group is completed). This ensures that portions of memory that are no longer needed are released as quickly as possible for use by other tasks. For example, the release of the last "slice" can create a subsequent contiguous block of sufficient size to be allocated to the next processed work group.
[0116] Figure 5202 is shown divided into a plurality of memory portions 203, and a SIMD processor 502 has five processing elements 503-507. In contrast to conventional approaches to shared memory allocation, this arrangement corresponds to Figure 1 Arrangement shown. The shared memory includes a memory portion 508 of an existing block allocated to working group B (the corresponding memory portion is marked as 'B' in the figure). Consider the point at which a first memory request is received for task 509 associated with a second working group A. Upon receipt of the memory request for task 509, a memory portion of block 514 (marked as 'A') is allocated to the shared memory, as shown. Thus, as described above, in response to receiving the first memory request for a task associated with the working group, an allocation of the memory portion of the block for the entire working group is performed. Slices 515 of the two memory portions of block 514 are allocated to the first task. Slices 516-519 of the block are reserved for other tasks of working group A.
[0117] When a memory request is received for subsequent tasks 510-513, each memory request can be serviced by allocating one of the reserved slices 516-519 to the corresponding task and the task scheduled for execution on the SIMD processor 502. Since all the memory required by the work group is allocated in advance, the last task 513 receives the slice 519 of the block. Since all tasks receive the memory allocations they need, the processing of all tasks can proceed to the barrier synchronization point 520. Once all tasks have been processed to the barrier 520, the processing of all tasks can be allowed to continue to complete, so that the work group can be processed as a whole. Due to the method of allocating continuous space to the entire block, the shared memory is more robust to fragmentation, and continuous space is more available in the shared memory for allocation to new blocks.
[0118] If there is not enough contiguous space available for the entire block 514, the memory request for the first task of the associated working group A is rejected. The rejected memory request can be handled in a variety of ways, but in general, the rejected memory request can be serviced until the next time the resource allocator expects to service the corresponding requester according to its defined mechanism (e.g., based on polling or according to a set of rules for arbitration between requesters). Once sufficient memory becomes available, subsequent memory allocation requests can succeed (e.g., by completing tasks on other processing units and the corresponding memory portions being released).
[0119] The resource allocator may maintain a data structure 206 that identifies which workgroup has been allocated a memory block. For example, when reserving a memory block for a workgroup, the resource allocator may add an entry to a data structure that associates the memory block with the workgroup so that when a task for the workgroup is subsequently received, the task allocator may identify which block of memory slices will be allocated. Similarly, the resource allocator may use the data structure 206 to identify which tasks do not belong to a workgroup that has already been allocated a memory block, thereby identifying the tasks for which a new block allocation will be performed. The data structure 206 may have any suitable configuration, for example, it may be a simple lookup table or a register.
[0120] In order to allocate slices of the reserved block to subsequently received tasks of the work group, the resource allocator needs to know which work group the memory request it receives belongs to. The memory request may include an identifier of the work group to which the corresponding task belongs, allowing the resource allocator to identify memory requests belonging to the work group for which the block has been reserved. Alternatively, the resource allocator may access a data store (e.g., a lookup table maintained by the work scheduler 302) that identifies which tasks belong to which work groups. Generally, the memory request will identify the task about which the memory request is being made, and such a task identifier may be used to look up the corresponding work group in the data store.
[0121] Memory requests received with respect to workgroups that have been allocated blocks of memory may preferably be prioritized. This helps to minimize delays in the system by ensuring that once memory has been allocated to a workgroup and resources are reserved in the system for processing that workgroup to completion, all tasks of the workgroup can begin processing as quickly as possible and reach any defined barrier synchronization points for the workgroup. In other words, ensuring that memory can be allocated to tasks as quickly as possible after reserving memory for them as blocks for the workgroup helps to avoid tasks that are subsequently allocated their slices of memory stalling the processing of other tasks in the workgroup that have been allocated other slices and are ready to begin processing.
[0122] Any suitable mechanism for prioritizing memory requests can be implemented. For example, the resource allocator can be configured to prioritize requests from the same requestor 205 from which the initial request is received, wherein the initial request causes the block to be allocated. The requestor can be prioritized over other requestors by increasing the frequency with which the requestor's memory requests are processed. The requestor can be prioritized by repeatedly processing memory requests received from the same requestor after allocating the block until a memory request for a task belonging to a different work group is received. The resource allocator can use the above data structure 206 to identify which memory request corresponds to the work group that has reserved the memory block to identify which memory request should be prioritized.
[0123] Note that in some embodiments of the system described herein, some tasks for which the resource allocator receives memory requests will not belong to a work group. Figure 3 The one or more processing units 301 used to process tasks in the resource allocator 302 may not be SIMD processors. The resource allocator may be configured to identify in a memory request it receives a task that does not belong to a work group from one or more identifiers, or to identify the absence of one or more identifiers (e.g., no work group identifier). Alternatively or in addition, the resource allocator may determine whether a received memory request is made for a task that does not belong to a work group by looking up the identifier received in the memory request in a data storage area (e.g., a lookup table maintained by the work scheduler 302) that identifies which task belongs to which work group.
[0124] The method of allocating memory portions at the resource allocator 201 will now be described.
[0125] The resource allocator can be configured to further divide the shared memory 202 into a plurality of non-overlapping windows, each of which includes a plurality of memory portions. The length of each window can be a power of 2, so as to enable the efficient implementation of the resource allocator in a binary logic circuit. The windows can all be of the same length. The resource allocator can maintain a fine state array 207 indicating which memory portions are allocated and which memory portions are not allocated, for example, the fine state array can include a one-dimensional bit array, each bit of which corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated (for example, '1' indicates that the portion is allocated; '0' indicates that the portion is not allocated). The size of each window can be selected so that the resource allocator can search in the window for the unallocated memory portion of the continuous block that is sufficient to contain a given size in a single clock cycle of the digital hardware that implements the resource allocator. Generally, the size of the shared memory 202 is too large for the resource allocator, so that the resource allocator cannot search the entire memory space in one clock cycle.
[0126] The resource allocator can maintain a window pointer 208 (e.g., a register inside the resource allocator) that identifies in which window of the shared memory the resource allocator will open a new memory allocation when it receives the first memory request for a work group that has not yet allocated memory. Upon receiving a memory request for the first task of the work group, the resource allocator checks the location in the current window where the contiguous memory block reserved for the work group can be adapted. This is called a "fine check". The resource allocator uses a fine state array 207 to identify unallocated memory portions. The resource allocator can start at the lowest position in the current window and scan in the direction of increasing memory addresses to identify the lowest possible starting position of the block and minimize fragmentation of the shared memory. Any appropriate algorithm for identifying a portion of memory that can provide a contiguous group of memory blocks can be used.
[0127] When a block is allocated to a work group, the fine state array is updated to mark all corresponding memory portions of the block as allocated, so that the block can be reserved for use by the tasks of the work group.
[0128] The resource allocator may also maintain a coarse state array 209 that indicates whether each window of the shared memory is completely unallocated, in other words, whether all memory portions of the window are unallocated. The coarse state array may be a one-dimensional bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is completely allocated. The bit value of a window in the coarse state array may be calculated as the OR reduction of all bits in the state array corresponding to the memory portion located in the window.
[0129] The resource allocator can be configured to perform a coarse check in conjunction with a fine check. The coarse check includes identifying from the coarse state array 209 whether the next one or more windows behind the current window indicated by the window pointer 208 can provide the new memory block required by the work group, that is, if the coarse state array 209 indicates that the next window is not allocated at all, all these memory portions can be used by the block. The one or more windows behind the current window can be checked because the required block may exceed the size of the window in some systems.
[0130] The resource allocator may be configured to allocate shared memory blocks to workgroups using fine and coarse checks as follows:
[0131] 1. If the fine check of the memory block for allocation to the workgroup succeeds and the memory block can be provided in the current window, the memory block is reserved for use by the workgroup (e.g., by marking the corresponding memory portion as allocated in the fine state array and adding the workgroup to data structure 206) and a slice (e.g., the first slice) of the memory block is allocated to the task for which the memory request was made. If the coarse check of the memory indicates that the current window of shared memory is completely unallocated, the coarse state array may be updated to indicate that the current window is now no longer completely unallocated.
[0132] 2. If the fine check of the memory block allocated to the working group is unsuccessful but the rough check of the memory is successful, the memory block can be provided by one or more unallocated windows behind the current window, and the memory block is reserved for the working group. The memory block can be allocated starting from the first unallocated memory portion of the unallocated memory portion that forms a continuous partition with the next window. The allocated memory block extends into one or more next windows that are identified as available for rough check.
[0133] If both the fine check and the coarse check fail, the memory request may be postponed and then attempted again (perhaps on the next clock cycle or after a predetermined number of clock cycles have passed). The resource allocator may service other memory requests before returning the failed memory request (e.g., the resource allocator may move to service the memory request from the next requester according to some rules defined for arbitrating between requesters).
[0134] Figure 6 An example allocation according to the above method is shown, wherein the shared memory is divided into 128 memory portions and 4 windows of 32 memory portions each. The '1' bit 601 in the fine state array 207 indicates that the first 24 memory portions are allocated to the work group, but all higher memory portions are not allocated. The coarse state array indicates that only the first window includes the allocated memory portions (see the '1' bit 602 corresponding to the first window). In this example, the work group requires 48 memory portion blocks of memory 603. The resource allocator can determine that such blocks are needed when receiving a memory request for 3 memory portions for the first task of the work group: it is known that a memory request has been received from a requester of a SIMD processor that is a work group for processing 16 tasks, and the resource allocator determines that the work group as a whole requires 48 memory portions.
[0135] Assume that the current window indicated by the window pointer of the resource allocator is the first window '0' (probably because the first unallocated memory portion after the allocated memory portion is memory portion 24). The fine check of the current window will fail because there are no remaining 48 memory portions in the first window '0'. However, the rough check will succeed because the second and third windows are indicated as not being allocated at all by the rough state array 209. Therefore, the block 48 starting from the memory portion 24 of the first window until the memory portion 71 of the third window will be allocated. The window pointer can be updated to identify the third window as the current window because the next available unallocated memory portion is number 72.
[0136] The resource allocator may be configured to calculate an overflow metric in parallel with the fine check and the coarse check. The overflow metric is given by the number of memory blocks required by the working group minus the number of unallocated memory portions at the top of the current window, so as to give a metric of the number of memory portions that would overflow the current window if the block were allocated starting from the lowest memory portion of the unallocated memory of the contiguous blocks at the top of the current window.
[0137] The resource allocator may be configured to use the overflow metric in an additional step after steps 1 and 2 above to attempt to allocate a memory block to the workflow:
[0138] 3. If the fine check and the rough check fail but there is still some free memory portion at the top of the current window, a new attempt to allocate the memory block is made at the next opportunity (e.g., the next clock cycle) starting from the lowest address of the next window of the shared memory using the overflow metric as the requested size. For the first allocation attempt, the fine check and the rough check as described above can be performed. This additional new attempt is advantageous because although the rough check on the next window fails, the lower partitions of the next window and the free memory portion of the current window are sufficient to provide enough continuous unallocated space to receive the memory block.
[0139] If step 3 fails, the memory request may be postponed and then attempted again later as described above.
[0140] Upon successful allocation, the resource allocator updates the window pointer to the window containing the memory portion (this may be the current window or the next window) immediately after the successful allocation of the memory portion block is complete.
[0141] The use of fine and coarse checks performed in parallel (optionally, in combination with the use of overflow metrics performed in parallel with the fine and coarse checks) enables the resource allocator to efficiently attempt to allocate shared memory to the workgroup in a minimum number of clock cycles. On appropriate hardware, the above method can enable block allocation to be performed in two clock cycles (in the first clock cycle, fine and coarse checks are performed and overflow metrics are optionally calculated; in the second clock cycle, allocation is attempted according to steps 1-2 above and optionally step 3).
[0142] Fig. 9 A flow chart showing a resource allocator allocating shared memory is shown. The resource allocator receives a shared memory resource request 901 for a first task of a relevant work group (i.e., a first task among the tasks of the work group for which a shared memory request has been received). In response, a shared memory block is allocated to the work group, the size of which is sufficient to provide the resources requested by the first task to all tasks of the work group. Generally, the shared memory block will be a contiguous shared memory block. A portion of the shared memory block is allocated 903 to the first task, which step may or may not be considered as a portion of the memory allocation to the work group 902, and therefore may be performed before, simultaneously with, or after allocating memory resources to the work group. In some embodiments, the remaining portion of the shared memory may be reserved for other tasks of the work group.
[0143] When a subsequent shared memory request 904 is received for another task of the same work group, the task is allocated a portion of the shared memory block allocated to the work group 905. In some embodiments, the allocation of shared memory for all tasks may be performed in response to receiving a first memory request 901 from a task of the work group. In such embodiments, further allocation may include responding by allocating memory resources to the task established in response to the first received memory request of the associated work group.
[0144] The allocation of memory resources from the block is repeated 907 until all tasks of the workgroup have received the required memory resources.Then, execution of the workgroup 906 can be performed, wherein each task uses the shared memory resources allocated to the task in the block allocated to the workgroup.
[0145] Fragment insensitive
[0146] A second configuration of a memory subsystem will now be described that also addresses the above-mentioned limitations of prior art systems for allocating shared memory.
[0147] Figure 7A memory subsystem 700 including a resource allocator 701, a shared memory 702, an input / output arbiter 703, and a conversion unit 704 is shown. The shared memory is divided into a plurality of memory portions 705. The resource allocator may be configured to receive memory requests from a plurality of requesters 205, for example, each task running on a processing unit may be a requester, each processing unit may be a requester, or each type of processing unit may be a requester. Different requesters may have different hardwired inputs to the resource allocator, and the resource allocator may need to arbitrate between these hardwired inputs. The resource allocator may service memory requests from the requesters on any basis, for example, it may service memory requests based on simple polling or according to a set of arbitration rules for selecting the next requester to be serviced. In other embodiments, memory requests from a plurality of requesters may be gathered at a request queue, memory allocation requests may be received in the request queue, and the resource allocator may receive memory requests to be serviced from the request queue. The allocated portion of the shared memory may be accessed by the processing unit using the input / output arbiter 703.
[0148] The memory subsystem 700 may be configured as follows: Figure 3 The approach shown and described above is arranged in a computer system.
[0149] In this configuration, the resource allocator is configured to use the conversion unit 704 to allocate shared memory 702 to the task, and the input / output arbiter 703 is configured to utilize the conversion unit 704 to access the shared memory 702. The conversion unit 704 is configured to associate each task of the memory request received with multiple memory sections 705 of the shared memory. Each task generally has an identifier provided in the memory request made about the task. The multiple memory sections associated with the task do not need to be continuous sections in the shared memory. However, from the perspective of the task, the task will receive the allocation of continuous memory blocks. When the same task is actually associated with the memory section of a potential non-continuous set at the conversion unit, this can be realized by the resource allocator configured to allocate continuous virtual memory blocks to the task.
[0150] When a task accesses a memory area assigned to it (e.g., by a read or write operation), the conversion unit 704 is configured to convert between a logical address provided by the task and a physical address where its data is stored or to which its data will be written. Note that the virtual address provided to the task by the resource allocator does not need to belong to a constant virtual address space, and in some examples, the virtual address will correspond to a physical address at the shared memory. It is advantageous for the logical address used by the task to have an offset relative to some base address (which itself can be a 0 offset or some other predetermined value) of a predetermined position in the virtual block (generally the starting position of the block). Since the virtual blocks are continuous, only the relative position of the data in the block that needs to be accessed needs to be provided to the conversion unit. The memory portion of a specific set of virtual blocks that constitute a task in the shared memory can be determined by the conversion unit based on the association between the task identifier and the memory portion in the data storage area 706.
[0151] In response to receiving a request for a memory resource, the resource allocator is configured to allocate to the task a group of memory portions sufficient to satisfy the request. The memory portions do not need to be continuous or sequentially stored in the shared memory. The resource allocator is configured to cause the conversion unit to associate the task with each memory portion in the group of memory portions. The resource allocator can track the unallocated memory portions at the shared memory using a data structure 707 indicating which memory portions are available. For example, the data structure can be a state array (e.g., the above-mentioned fine state array) whose bits identify whether each memory portion at the shared memory is allocated.
[0152] The conversion unit may include a data storage area 706 and may be configured to associate a task with a set of memory portions by storing in the data storage area an identifier of the task (e.g., task_id) and a physical address of each memory portion in the set of memory portions, and an indication of which partition of a continuous virtual memory block each memory portion allocated to the task corresponds to (e.g., an offset of the partition in the virtual memory block or a logical address of the partition). In this way, the conversion unit may return an appropriate physical address to the task in response to the task providing its identifier and an indication of which partition of its virtual memory block it needs (e.g., an offset in its virtual block). The offset may be the number of memory portions of the memory region relative to the base address (e.g., in the case where the memory portion is a basic allocation interval), the number of bytes or bits, or any other position indication. The base address may be, but not necessarily, the memory address at which the virtual memory block begins; the base address may be a zero offset. The data storage area may be an associative memory (e.g., a content addressable memory), the input of which may be a task identifier and an indication of the required memory partition in its block (e.g., a memory offset relative to the base address of its block).
[0153] More generally, a conversion unit may be provided at any suitable point along the data path between the resource allocator 701 and the shared memory 702, and between the I / O arbiter 703 and the shared memory 702. More than one conversion unit may be provided. Multiple conversion units may share a common data storage area such as a content addressable memory.
[0154] The resource allocator may be configured to assign logical base addresses to different tasks of the workgroup to identify contiguous virtual memory blocks allocated to each task of the workgroup that are present in a contiguous virtual memory superblock for the workgroup as a whole.
[0155] Figure 8 The corresponding relationship between the memory portion 805 of the shared memory 702 allocated to the task and the virtual block 801 allocated to the task is shown. The memory portion 805 of the shared memory marked as AF provides a lower storage area for the partition 806 of the virtual continuous memory block 801. The memory block AF is discontinuous or stored sequentially in the shared memory and is located in the middle of the memory portion 804 allocated to other tasks. The virtual block 801 has a base address 802, and the partitions in the block can be identified by their offset 803 relative to the base address.
[0156] In one example, in response to receiving a request for a memory resource, the resource allocator is configured to allocate to the task a contiguous range of memory addresses starting from a physical address of the memory portion corresponding to the first partition of the block. In other words, the base address 802 of the virtual block (and Figure 8 The base address of partition 'A' or 809 in Figure 8 808 in the shared memory. Assuming that there are enough unallocated memory portions in the shared memory to service the memory request, the resource allocator allocates a virtual memory block of the requested size starting from the physical memory address of the memory partition 809 to the task. Therefore, from the perspective of the task, it is allocated a contiguous block of six memory portions in the shared memory starting from the memory portion 808; in fact, the task is allocated six non-contiguous memory portions in the shared memory 702 labeled A, B, C, D, E, and F. Therefore, the actual allocation of memory portions to the task is marked by the virtual block 801 it receives. In this example, the virtual blocks have physical addresses and there is no such virtual address space.
[0157] Once a task has been allocated a block of virtual memory, it may access the memory using an I / O arbiter 703, which typically arbitrates access to shared memory between multiple units (e.g., processing units or other units capable of accessing shared memory) that submit access requests (e.g., read / write) via an interface 708 (which may include one or more hard-wired connections between which the I / O arbiter arbitrates).
[0158] Continuing with the above example, consider a task requesting access to a region of block 801 located in partition 'D'. The task may submit a request to the I / O arbiter 703 including its task identifier (e.g., task_id) and an indication of the region of the block it needs to access. The indication may be, for example, an offset 803 (e.g., number of bits, bytes, memory portion units) relative to a base address 802, or a physical address formed based on the physical base address 802 assigned to the work and the offset 803 in the block where the requested memory region is found. The conversion unit will look up the work unit identifier and the expected offset / address in the data storage area 706 to identify the corresponding physical address identifying the memory portion 811 in the shared memory 702. The switch unit will service the access request from the memory portion 811 in the shared memory. In the case where the memory region in the block spans more than one memory partition, the data storage area will return more than one physical address for the corresponding data portion.
[0159] In an embodiment where a task provides a physical address of a storage area, the submitted physical address in this case refers to a storage area in the memory portion 810 allocated to another task. Therefore, different tasks can apparently access overlapping blocks in the shared memory. However, since the conversion unit 704 converts each access request, the access to the physical address submitted by the task is redirected to the correct memory portion (811 in this example) by the conversion unit. Therefore, in this case, the logical address that references the virtual memory block is actually a physical address, but not necessarily the physical address at which the corresponding data of the block is stored in the shared memory.
[0160] In a second example, in response to receiving a request for a memory resource, a resource allocator is configured to allocate to a task a logical memory address starting from a logical base address and representing a continuous range of virtual memory blocks. Logical addresses can be allocated in any suitable manner, since a given logical address will be mapped to a physical address of a corresponding lower memory portion. For example, a task may be allocated a logical address in a virtual address space that remains consistent between tasks so that virtual blocks allocated to different tasks do not share overlapping logical address ranges; a task may be allocated a logical base address from a set of logical addresses and an offset relative to the base address that is used to refer to a location in a virtual block; or a logical base address may be allocated as a zero offset.
[0161] The conversion unit responds to access requests (e.g., read / write) received from tasks via the I / O arbiter 703. Each access request may include an identifier of the task and an offset or logical address identifying the region of the virtual block that it needs to access. Upon receiving an access request from a task via the I / O arbiter 703, the conversion unit is configured to identify the desired physical address in the data structure based on the identifier of the work unit and the offset / logical address. The offset / logical address is used to identify which memory portions associated with the task will be accessed.
[0162] For example, consider again the task requesting access Figure 8 801 of the virtual block 801 shown. The conversion unit will receive an access request from a task including an identifier of the task, and an offset 803 pointing to a location in partition 'D'. Using the identifier and offset of the task, the conversion unit looks up the physical address of the corresponding memory portion 811 in the shared memory in the data storage area 706 (i.e., the physical address in the data storage area associated with the task and the offset). The conversion unit can then proceed to service the access request at the corresponding memory portion 811 (e.g., by performing the requested read / write from / to the memory portion 811).
[0163] It will be appreciated that the "fragmentation insensitive" approach described herein allows the underlying shared memory 702 to become fragmented while the tasks themselves receive contiguous memory blocks. This enables memory blocks to be allocated to tasks and workgroups including multiple tasks even when there is not enough contiguous space in the shared memory to allocate contiguous blocks of memory portions of the required size, provided there is a sufficient number of memory portions at the shared memory.
[0164] Fig.10A flow chart showing the allocation of memory resources to tasks by a resource allocator is shown. Upon receiving a request for shared memory resources from a task 1001, a contiguous block of virtual memory is allocated to the task 1002. The virtual memory block is associated 1003 with a physical memory address of the shared memory so that each portion of the virtual block corresponds to an underlying portion of the physical shared memory. Once the task receives its allocated virtual memory, the task can be executed 1004 on a SIMD processor, wherein the task uses the underlying shared memory resources at the physical address corresponding to its allocated virtual memory block. Figure 2 , 3 , 7 is shown as including a plurality of functional blocks. This is merely illustrative and does not intend to define a strict division between the different logical elements of these entities. Each functional block may be provided in any suitable manner. It will be understood that the intermediate values formed by each functional block described herein need not be physically generated by the functional block at any time, but may simply represent a logical value that conventionally describes the processing performed by the functional block between its input and output.
[0165] The memory subsystem described herein can be implemented in hardware on an integrated circuit. The memory subsystem described herein can be configured to perform any method described herein. Generally, any function, method, technology, or component described above can be implemented in software, firmware, hardware (e.g., fixed logic circuit), or any combination thereof. The terms "module", "function", "component", "element", "unit", "block", and "logic" can be used here to generally represent software, firmware, hardware, or any combination thereof. In the case of software implementation, a module, function, component, element, unit, block, or logic represents a program code that performs a specified task when executed on a processor. The algorithms and methods described herein can be performed by one or more processors that execute a code, and the code causes the one or more processors to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disk, flash memory, hard disk storage, and other memory devices that can use magnetic, optical, or other technologies to store instructions or other data that can be accessed by a machine.
[0166] The terms "computer program code" and "computer readable instructions" are used herein to refer to any kind of executable code for a processor, including code expressed in a machine language, an interpreted language, or a scripting language. Executable code includes binary code, machine code, byte code, code that defines an integrated circuit (e.g., a hardware description language or a netlist), and code expressed in a programming language code (e.g., C, Java, or OpenCL). Executable code can be, for example, any kind of software, hardware, script, module, or library that, when properly executed, processed, interpreted, assembled, executed in a virtual machine or other software environment, causes a processor of a computer system supporting the executable code to perform the tasks specified by the code.
[0167] A processor, computer, or computer system may be any kind of device, machine, or special circuit, or a collection or part thereof, that has processing capabilities so that it can execute instructions. A processor may be any kind of general or special processor, such as a CPU, GPU, system on a chip, state machine, media processor, application specific integrated circuit (ASIC), programmable logic array, field programmable gate array (FPGA), etc. A computer or computer system may include one or more processors.
[0168] It is also intended to cover software that defines the hardware configuration described herein, for example, HDL (Hardware Description Language) software for designing integrated circuits or for configuring programmable chips to achieve desired functions. That is, a computer-readable storage medium may be provided having encoded thereon a computer-readable program code in the form of an integrated circuit definition data set, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture a memory subsystem configured to perform any of the methods described herein, or to manufacture a memory subsystem including any of the apparatus described herein. The integrated circuit definition data set may be, for example, an integrated circuit description.
[0169] A method of manufacturing the memory subsystem described herein at an integrated circuit manufacturing system is provided. An integrated circuit definition data set may be provided which, when processed by the integrated circuit manufacturing system, causes the method of manufacturing the memory subsystem to be performed.
[0170] The integrated circuit definition data set may be in the form of computer code, for example, as a netlist, code for configuring a programmable chip, as a hardware description language that defines an integrated circuit at any level (including register transfer level (RTL) code), as a high-level circuit representation (e.g., Verilog or VHDL), and as a low-level circuit representation (e.g., OASIS (RTM) and GDSII). A higher-level representation (e.g., RTL) that logically defines an integrated circuit may be processed at a computer system configured to generate a manufacturing definition of an integrated circuit in the context of a software environment, wherein the manufacturing definition includes definitions of circuit elements and rules for joining these elements to generate a manufacturing definition that represents the defined integrated circuit. In the general case where a computer system executes software to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to cause the computer system configured to generate a manufacturing definition of an integrated circuit to execute the executable code that defines the integrated circuit to generate the manufacturing definition of the integrated circuit.
[0171] Reference now Fig.11 , describes an example of processing an integrated circuit definition data set in an integrated circuit manufacturing system to configure the system to manufacture a memory subsystem.
[0172] Fig.11 An example of an integrated circuit (IC) manufacturing system 1102 that can be configured to manufacture a memory subsystem described in any of the examples herein is shown. Specifically, the IC manufacturing system 1102 includes a layout processing system 1104 and an integrated circuit generation system 1106. The IC manufacturing system 1102 is configured to receive an IC definition data set (e.g., defining a memory subsystem described in any of the examples herein), process the IC definition data set, and generate an IC based on the IC definition data set (e.g., an IC definition data set embodying a memory subsystem described in any of the examples herein). The processing of the IC definition data set configures the IC manufacturing system 1102 to manufacture an integrated circuit embodying a memory subsystem described in any of the examples herein.
[0173] The layout processing system 1104 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout based on an IC definition data set are known in the art, and may include, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated (e.g., in terms of logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components)). The circuit layout may be determined from the gate-level representation of the circuit by determining location information for the logic components. This may be accomplished automatically or under user control to optimize the circuit layout. When the layout processing system 1104 has determined the circuit layout, it may output a circuit layout definition to the IC generation system 1106. The circuit layout definition may be, for example, a circuit layout description.
[0174] The IC generation system 1106 generates the IC according to the circuit layout definition, as is known in the art. For example, the IC generation system 1106 can perform a semiconductor device assembly process to generate the IC, which can include a multi-step sequence of photolithography processes and chemical processing steps during which electronic circuits are gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask that can be used in the photolithography process to generate the IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 1106 can be in the form of a computer readable code, and the IC generation system 1106 can use the computer readable code to form an appropriate mask for generating the IC.
[0175] The different processes that may be performed by the IC manufacturing system 1102 may all be performed in one location, e.g., by one party. Alternatively, the IC manufacturing system 1102 may be a distributed system so that some of the processes may be performed in different locations and may be performed by different parties. For example, some of the following stages may be performed in different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition data set to form a gate-level representation of a circuit to be generated, (ii) generating a circuit layout based on the gate-level representation, (iii) forming a mask based on the circuit layout, and (iv) assembling the integrated circuit using the mask.
[0176] In other examples, processing of an integrated circuit definition data set at an integrated circuit manufacturing system may configure the system to manufacture a memory subsystem without processing the IC definition data set to determine a circuit layout. For example, an integrated circuit definition data set may define a configuration of a reconfigurable processor (e.g., an FPGA), and processing of the data set may configure the IC manufacturing system to generate a reconfigurable processor having the defined configuration (e.g., by loading the configuration data into the FPGA).
[0177] In some embodiments, when the integrated circuit manufacturing definition data set is processed in the integrated circuit manufacturing system, it can cause the integrated circuit manufacturing system to generate the device described herein. Fig.11 The configuration of an integrated circuit manufacturing system in the described manner may enable the apparatus described herein to be manufactured.
[0178] In some examples, an integrated circuit definition data set may include software that runs on or interfaces with hardware defined at the data set. Fig.11 In the example shown, the IC generation system can be further configured by the integrated circuit definition data set to load hardware into the integrated circuit or provide program code to the integrated circuit for use by the integrated circuit according to the program code defined in the integrated circuit definition data set when manufacturing the integrated circuit.
[0179] The graphics processing system described herein may be implemented in hardware on an integrated circuit.The graphics processing system described herein may be configured to perform any of the methods described herein.
[0180] The concepts set forth in this application can produce performance improvements compared to known implementations in devices, apparatuses, modules, and / or systems (and methods implemented herein).Performance improvements can include one or more of increased computing performance, reduced delays, increased throughput, and / or reduced power consumption.During (e.g., in integrated circuits) manufacturing these devices, devices, modules, and systems, a balance can be made between performance improvements and physical implementations, thereby improving manufacturing methods.For example, a balance can be made between performance improvements and layout areas, thereby matching the performance of known implementations but using less silicon.This can be achieved, for example, by reusing functional blocks in a continuous manner or sharing functional blocks between elements of devices, devices, modules, and / or systems.On the contrary, the concepts of this application that lead to improvements in the physical implementation of devices, devices, modules, and systems (e.g., silicon area reduction) can be exchanged for improved performance.This can be achieved, for example, by manufacturing multiple instances of modules in a predetermined area budget.
[0181] The applicant discloses each individual feature described herein separately, and any combination of two or more such features, so that these features or combinations can be implemented based on the specification according to the common knowledge of those skilled in the art, regardless of whether these features or feature combinations solve any problems disclosed herein. In view of the above description, it is obvious to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units configured to process one or more work groups each comprising a plurality of SIMD tasks, the memory subsystem comprising: a shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor; as well as The resource allocator is configured to do the following: receiving a single memory resource request for a first memory resource from a first receiving task of a work group, the work group including the first receiving task and at least one other task; In response to receiving the single memory resource request, allocating a block of memory to the first receiving task and the at least one other task of the working group, the block of memory being large enough for the first receiving task and the at least one other task of the working group to each receive a memory resource in the block of memory equivalent to the first memory resource, wherein the block of memory is: (i) a contiguous block of physical memory, or (ii) A contiguous block of virtual memory.
2. The memory subsystem of claim 1, wherein: The resource allocator is configured to perform at least one of the following processes: while servicing the first receiving task of the work group, allocating the requested first memory resource from the block of memory portions to the task, and reserving a remaining memory portion of the block of memory portions to prevent allocation to tasks of other work groups; as well as In response to subsequently receiving a request for memory resources from a second task of the work group, the memory resource of the block of memory is allocated to the second task.
3. The memory subsystem of claim 1 or 2, wherein: The resource allocator is configured to receive memory resource requests from a plurality of different requestors, and in response to allocating the block of memory portion to the work group, prioritize the memory request received from the requestor from which the first received task of the work group is received.
4. The memory subsystem of claim 1 or 2, wherein: The resource allocator is further configured to, in response to receiving an indication that processing of a task of the work group has completed, release memory resources allocated to the task without waiting for processing of the work group to complete.
5. The memory subsystem of claim 1 or 2, wherein: The shared memory is further divided into a plurality of non-overlapping windows, wherein each non-overlapping window comprises a plurality of memory portions, and the resource allocator is configured to maintain a window pointer indicating a current window in which allocation of a memory portion will be attempted in response to a next received memory request.
6. The memory subsystem of claim 5, wherein: The resource allocator is implemented in a binary logic circuit, and the window length is such that the availability of all memory portions of each window can be checked in a single clock cycle of the binary logic circuit.
7. The memory subsystem of claim 1 or 2, wherein: The resource allocator is further configured to maintain a fine state array arranged to indicate whether each memory portion of the shared memory is allocated to a task.
8. The memory subsystem of claim 7, wherein: The resource allocator is configured to, in response to receiving a memory resource request for the first receiving task regarding the work group, search in a current window for a memory portion that is indicated by the fine status array as a continuous block available for allocation, and the resource allocator is configured to allocate the continuous block to the work group if such a continuous block is identified in the current window.
9. The memory subsystem of claim 8, wherein: The resource allocator is configured to allocate the memory portion of the contiguous block such that the block starts at a lowest possible position of the window.
10. The memory subsystem of claim 7, wherein: The shared memory is further divided into a plurality of non-overlapping windows, wherein each non-overlapping window includes a plurality of memory portions, wherein the resource allocator is further configured to maintain a coarse state array, the coarse state array being arranged to indicate, for each non-overlapping window of the shared memory, whether all memory portions of the non-overlapping window are unallocated, and the resource allocator is configured to check the coarse state array in parallel with searching for memory portions of a continuous block in a current non-overlapping window to determine whether the size of the requested block can be provided by one or more subsequent non-overlapping windows; the resource allocator is configured to allocate to the work group a block including a memory portion starting from a first memory portion of the current window in a continuous block with the one or more subsequent windows and extending to the memory portion of the one or more subsequent windows if a sufficiently large continuous block cannot be identified in the current window and the requested block can be provided by one or more subsequent windows.
11. The memory subsystem of claim 10, wherein: The resource allocator is further configured to form an overflow metric in parallel with searching the current window, the overflow metric representing memory resources of a memory portion of a required block that cannot be provided in the current window starting from the first memory portion of the current window in an unallocated memory portion of a continuous block adjacent to the subsequent window, and the resource allocator is configured to try to allocate a memory portion to the work group by searching in the subsequent window starting from the first memory portion of the subsequent window for an unallocated memory portion of a continuous block whose total size is sufficient to provide the overflow metric if a sufficiently large continuous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows.
12. The memory subsystem of claim 7, wherein: The fine status array is a bit array in which each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated.
13. The memory subsystem of claim 10, wherein: The coarse state array is a bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is completely allocated.
14. The memory subsystem of claim 12, wherein: The shared memory is further divided into a plurality of non-overlapping windows, wherein each non-overlapping window comprises a plurality of memory portions, wherein the resource allocator is further configured to maintain a coarse state array arranged to indicate, for each non-overlapping window of the shared memory, whether all memory portions of the non-overlapping window are unallocated, wherein the coarse state array is a bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is fully allocated, and wherein the resource allocator is configured to form each bit of the coarse state array by performing an OR reduction on all bits of the fine state array corresponding to memory portions located in the window corresponding to the bit of the coarse state array.
15. The memory subsystem of claim 6, wherein: The window length is a power of 2.
16. The memory subsystem of claim 1 or 2, wherein: The resource allocator may maintain a data structure identifying which of the one or more workgroups are currently allocated a block of memory portions.
17. A method of allocating shared memory resources to tasks for execution in a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units respectively configured to process one or more work groups, each of the one or more work groups comprising a plurality of SIMD tasks, the method comprising: receiving a single shared memory resource request for a first memory resource from a first receiving task of a work group, the work group including the first receiving task and at least one other task; as well as In response to receiving the single shared memory resource request, allocating a memory portion of a shared memory to the first receiving task and the at least one other task of the working group, the size of the memory portion being sufficient for the first receiving task and the at least one other task in the working group to each receive a memory resource in the memory portion equivalent to the first memory resource, wherein the memory portion is: (i) a contiguous block of physical memory, or (ii) A contiguous block of virtual memory.
18. A non-transitory computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, the computer-readable description, when processed in an integrated circuit manufacturing system, causing the integrated circuit manufacturing system to manufacture the memory subsystem of claim 1.
19. An integrated circuit manufacturing system comprising: a non-transitory computer readable storage medium having stored thereon a computer readable integrated circuit description describing the memory subsystem of claim 1; a layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit implementing the memory system; as well as An integrated circuit production system is configured to manufacture the memory subsystem according to the circuit layout description.
20. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon, the computer-readable instructions being executed at a computer system so that the computer system performs the method of claim 17.
Citation Information
Patent Citations
Interlocked Increment Memory Allocation and Access
US20110055511A1