Resource allocation

By using resource allocator in SIMD processor to conduct fine and rough inspections on shared memory and prioritize the allocation of continuous block resources, the deadlock and fragmentation problems in shared memory are solved, and efficient utilization of memory resources and smooth task processing are achieved.

CN120353584APending Publication Date: 2025-07-22IMAGINATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510416749.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-09-15
Filing Date
2018-09-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In multi-processing unit systems, deadlocks and resource fragmentation problems caused by dynamic allocation mechanisms of shared memory, especially in SIMD processors, lead to task processing delays and resource waste.

Method used

The resource allocator is used to divide the shared memory into multiple memory parts, and through fine and coarse inspection mechanisms, continuous block resources are allocated first, combining fine state arrays and coarse state arrays to ensure that each workgroup task can continuously receive memory resources and avoid deadlocks and fragmentation.

Benefits of technology

It improves the utilization rate of memory resources, reduces task processing delays, and ensures the smooth execution of tasks and efficient allocation of resources in SIMD processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353584A_ABST
    Figure CN120353584A_ABST
Patent Text Reader

Abstract

The invention relates to resource allocation. A memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor comprising a plurality of processing units configured to process one or more work groups each comprising a plurality of SIMD tasks, the memory subsystem comprising: a shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor; a resource allocator configured to allocate a block of memory portion to the work group in response to receiving a memory resource request for a first memory resource with respect to a first receive task of the work group, the one block of memory portion is sized enough for each task of the work group to receive a memory resource in the one block of memory portion that corresponds to the first memory resource.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Application Instructions

[0002] This application is a divisional application of the invention patent application with the application date of September 17, 2018, the application number of 201811083351.8, and the title of "Resource Allocation". Technical Field

[0003] The present disclosure relates to a memory subsystem for multiple processing units. Background Art

[0004] In a system including multiple processing units, a shared memory that can be accessed by at least some of the processing units is typically provided. Due to the range of processing tasks that can run on the processing units and have different memory requirements respectively, the shared memory resources are generally not fixed for each processing unit during design. An allocation mechanism that allows processing tasks running on different processing units to request the allocation of one or more regions of the shared memory respectively is typically provided. This enables the shared memory to be dynamically allocated for use by tasks executed at the processing units.

[0005] Efficient use of the shared memory can be achieved through careful design of the allocation mechanism. Summary of the Invention

[0006] The present invention content is provided to introduce a selection of concepts further described below in the detailed description. The present invention content is not used to identify the key features or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter.

[0007] A memory subsystem is provided for a single instruction multiple data (SIMD) processor including multiple processing units, the multiple processing units being configured to process one or more workgroups respectively including multiple SIMD tasks, the memory subsystem including:

[0008] A shared memory divided into multiple memory portions for allocation to tasks to be processed by the processor; and

[0009] A resource allocator configured to, in response to receiving a memory resource request for a first memory resource of a first received task of a workgroup, allocate a block of memory portion to the workgroup, the size of the block of memory portion being sufficient for each task in the workgroup to receive a memory resource equivalent to the first memory resource in the block of memory portion.

[0010] The resource allocator may be configured to allocate a continuous block of memory resources as the block of memory portion.

[0011] The resource allocator can be configured to allocate the requested first memory resource to the task from the block memory section when servicing the first received task of the workgroup, and reserve the remaining memory section of the block memory section against tasks assigned to other workgroups.

[0012] The resource allocator can be configured to, in response to subsequently receiving a memory resource request for a second task of the workgroup, allocate the memory resource of the block memory section to the second task.

[0013] The resource allocator can be arranged to receive memory resource requests from multiple different requesters, and to preferentially service memory requests received from the following requester in response to allocating the block memory section to the workgroup, where the first received task of the workgroup was received from that requester.

[0014] The resource allocator can further be configured to, in response to receiving an indication that the processing of a task of the workgroup has been completed, release the memory resources allocated to the task without waiting for the processing of the workgroup to complete.

[0015] The shared memory can further be divided into a plurality of non-overlapping windows, each non-overlapping window including a plurality of memory sections, and the resource allocator is configured to maintain a window pointer indicating the current window, where memory sections will be attempted to be allocated in the current window in response to the next received memory request.

[0016] The resource allocator can be implemented in binary logic circuitry, and the window length can be such that the availability of all memory sections of each window can be checked in a single clock cycle of the binary logic circuitry.

[0017] The resource allocator can further be configured to maintain a fine-grained status array, which is arranged to indicate whether each memory section of the shared memory is allocated to a task.

[0018] The resource allocator can be configured to, in response to receiving a memory resource request for the first received task of the workgroup, search in the current window for a contiguous block of memory sections indicated by the fine-grained status array as available for allocation, and the resource allocator is configured to, if such a contiguous block is identified in the current window, allocate that contiguous block to the workgroup.

[0019] The resource allocator can be configured to allocate a contiguous block of memory sections such that the block starts at the lowest possible position in the window.

[0020] The resource allocator may further be configured to maintain a coarse state array that is arranged to indicate for each window of the shared memory whether all memory portions of the window are unallocated. The resource allocator is configured to check the coarse state array in parallel with searching for a contiguous block of memory portions in the current window to determine whether the size of the requested block can be provided by one or more subsequent windows; the resource allocator is configured to allocate to a workgroup a block that includes a first memory portion of the current window that is in a contiguous block with one or more subsequent windows and extends to memory portions in those subsequent windows if a sufficiently large contiguous block cannot be identified in the current window and the requested block can be provided by one or more subsequent windows.

[0021] The resource allocator may further be configured to form an overflow metric in parallel with searching the current window, the overflow metric representing the memory resources of the memory portions of the required block that cannot be provided in the current window starting from the first memory portion of the current window in the unallocated memory portions of the contiguous block of the immediately subsequent window. The resource allocator is configured to, if a sufficiently large contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then attempt to allocate a block of memory portions to a workgroup by searching in the subsequent windows for unallocated memory portions of a contiguous block whose total size is sufficient to provide the overflow metric starting from the first memory portion of the subsequent window.

[0022] The fine state array may be a bit array in which each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated.

[0023] The coarse state array may be a bit array in which each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is fully allocated.

[0024] The resource allocator may be configured to form each bit of the coarse state array by performing an OR reduction on all the bits of the fine state array that correspond to the memory portions that are present in the window corresponding to that bit in the coarse state array.

[0025] The window length may be a power of 2.

[0026] The resource allocator may maintain a data structure that identifies which of one or more workgroups are currently being allocated blocks of memory portions.

[0027] According to a second aspect, there is provided a method of allocating shared memory resources to a task executed in a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units, each processing unit being configured to process one or more workgroups respectively including a plurality of SIMD tasks, the method including:

[0028] Receiving a shared memory resource request for a first memory resource for a first received task of a workgroup; and

[0029] Allocating a memory portion of a block of shared memory to the workgroup, the size of the block of memory portion being sufficient for each task in the workgroup to receive a memory resource equivalent to the first memory resource in the block of memory portion.

[0030] Allocating a memory portion of a block to the workgroup may include allocating the block of memory portion as a contiguous block of memory.

[0031] The method may further include: allocating the requested first memory resource to the task from the block of memory portion when servicing the first received task of the workgroup, and reserving the remaining memory portion of the block of memory portion in case of being allocated to tasks of other workgroups.

[0032] The method may further include: in response to a subsequent receipt of a memory resource request for a second task of the workgroup, allocating the memory resource of the block of memory portion to the second task.

[0033] The method may further include: receiving memory resource requests from a plurality of different requesters, and preferentially servicing memory requests received from the following requester in response to allocating the block of memory portion to the workgroup, the first received task of the workgroup being received from the requester.

[0034] The method may further include: in response to receiving an indication that the processing of a task of the workgroup has been completed, releasing the memory resources allocated to the task without waiting for the processing of the workgroup to be completed.

[0035] The method may further include: maintaining a fine state array, the fine state array being arranged to indicate whether each memory portion of the shared memory is allocated to a task.

[0036] Allocating a memory portion of a block to the workgroup may include: searching for a contiguous block of memory portion indicated as available for allocation by the fine state array in a current window, and if such a contiguous block is identified in the current window, allocating the contiguous block to the workgroup.

[0037] Allocating a contiguous block may include: allocating the memory portion of the contiguous block such that the contiguous block starts at the lowest possible position in the window.

[0038] The method may further include:

[0039] Maintaining a coarse state array that indicates, for each window of the shared memory, whether all memory portions of the window are allocated;

[0040] Checking the coarse state array in parallel with searching for memory portions of contiguous blocks in the current window to determine whether the size of the requested block can be provided by one or more subsequent windows; and

[0041] If a large enough contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, allocating to a workgroup a block that includes a memory portion starting from a first memory portion of the current window in a contiguous block with the subsequent window and extending to memory portions in those subsequent blocks.

[0042] The method may further include:

[0043] Forming, in parallel with searching the current window, a representation of an overflow metric that represents memory resources of memory portions of a required block that cannot be provided in the current window starting from a first memory portion of the current window in unallocated memory portions of contiguous blocks of immediately subsequent windows; and

[0044] If a large enough contiguous block cannot be identified in the current window and the requested block cannot be provided by one or more subsequent windows, then subsequently attempting to allocate a memory portion of a block to a workgroup by searching unallocated memory portions of contiguous blocks in subsequent windows starting from a first memory portion of the subsequent window that are large enough in total size to accommodate the overflow metric.

[0045] According to a third aspect, there is provided a memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units for processing SIMD tasks, the memory subsystem including:

[0046] A shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor;

[0047] A translation unit configured to associate a task with one or more physical addresses of the shared memory; and

[0048] A resource allocator configured to, in response to receiving a memory resource request for a first memory resource regarding a task, allocate a contiguous virtual memory block to the task and cause the translation unit to associate the task with a plurality of physical addresses of memory portions respectively corresponding to partitions of the virtual memory block, the memory portions together implementing the complete virtual memory block.

[0049] At least some of the multiple memory portions of the block may be discontinuous in the shared memory.

[0050] A virtual memory block may include a base address of a physical address as one of the memory portions associated with a task.

[0051] A virtual memory block may include a base address of a logical address of the virtual memory block.

[0052] A translation unit is operable to subsequently service an access request received from a task regarding a virtual memory block, the access request including an identifier of the task and an offset of a storage area in the virtual memory block, the translation unit being configured to service the access request in a memory portion corresponding to a partition of the virtual memory block indicated by the offset.

[0053] The translation unit may include: a content-addressable memory configured to return one or more corresponding physical addresses in response to receiving an identifier of an item and an offset in the virtual memory block.

[0054] A SIMD processor may be configured to process a workgroup including multiple tasks, the task being the first received task of the workgroup, and a resource allocator is configured to reserve sufficient memory portions for the workgroup such that each task of the workgroup can receive a memory resource equivalent to a first memory resource, and allocate the requested memory resource from the memory portions reserved for the workgroup to the first received task.

[0055] The resource allocator may be configured to allocate virtual memory blocks to each task of the workgroup such that the virtual memory blocks allocated to the tasks of the workgroup together represent virtual memory blocks of a continuous superblock.

[0056] The resource allocator may be configured to, in response to subsequently receiving a memory resource request regarding a second task of the workgroup, allocate continuous virtual memory blocks from the superblock to the second task.

[0057] According to a fourth aspect, there is provided a method for allocating shared memory resources to a task executed in a single instruction multiple data (SIMD) processor, the SIMD processor including multiple processing units respectively configured to process SIMD tasks, the method including:

[0058] Receiving a memory resource request for a first memory resource regarding a task;

[0059] Allocating continuous virtual memory blocks to the task; and

[0060] Associating the task with multiple physical addresses of memory portions of the shared memory, each memory portion corresponding to a partition of the virtual memory block, the memory portions together implementing a complete virtual memory block.

[0061] The method may further comprise:

[0062] subsequently receiving an access request received from a task regarding a virtual memory block, the access request including an identifier of the task and an offset of a storage area in the virtual memory block; and

[0063] serving the access request by accessing a memory portion corresponding to a partition of the virtual memory block indicated by the offset.

[0064] Tasks may be grouped together in a workgroup for execution at a SIMD processor, a task may be a first received task of the workgroup, and allocating consecutive virtual memory blocks to the tasks may include:

[0065] reserving a sufficient memory portion for the workgroup for each task of the workgroup to receive a memory resource equivalent to a first memory resource; and

[0066] allocating the requested memory resource from the memory portion reserved for the workgroup to the first received task.

[0067] Reserving a memory portion for the workgroup may include reserving virtual memory blocks for each task of the workgroup such that the virtual memory blocks allocated to the tasks of the workgroup together represent virtual memory blocks of a consecutive superblock.

[0068] The method may further comprise: in response to a subsequent receipt of a memory resource request for a second task of the workgroup, allocating consecutive virtual memory blocks from the superblock to the second task.

[0069] The memory subsystem may be implemented in hardware on an integrated circuit.

[0070] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided.

[0071] A method of manufacturing the memory subsystem described herein using an integrated circuit manufacturing system is provided, the method comprising:

[0072] processing a computer-readable description of a graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and

[0073] manufacturing the graphics processing system according to the circuit layout description using an integrated circuit generation system.

[0074] An integrated circuit definition data set is provided that, when processed in an integrated circuit manufacturing system, configures the system to manufacture the memory subsystem described herein.

[0075] A non-transitory computer-readable storage medium is provided, having stored thereon a computer-readable description of an integrated circuit, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture the memory subsystem described herein.

[0076] A non-transitory computer-readable storage medium is provided, having stored thereon a computer-readable description of the memory subsystem described herein, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to:

[0077] Process a computer-readable description of a graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit implementing the graphics processing system; and

[0078] Manufacture a graphics processing system using an integrated circuit generation system based on the circuit layout description.

[0079] An integrated circuit manufacturing system is provided, configured to manufacture the memory subsystem described herein.

[0080] An integrated circuit manufacturing system is provided, comprising:

[0081] A non-transitory computer-readable storage medium having stored thereon a computer-readable integrated circuit description of the memory subsystem described herein;

[0082] A layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit implementing the memory subsystem; and

[0083] An integrated circuit generation system configured to manufacture the memory subsystem based on the circuit layout description.

[0084] A memory subsystem is provided, configured to execute the method described herein. A computer program code is provided for executing the method described herein. A non-transitory computer-readable storage medium is provided, having stored thereon computer-readable instructions, which, when executed at a computer system, cause the computer system to execute the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The present invention is described by way of example with reference to the accompanying drawings. In the drawings:

[0086] Figure 1 A conventional allocation of shared memory for tasks executed at a SIMD processor is shown.

[0087] Figure 2 is a schematic diagram of a first memory subsystem including a resource allocator configured according to the principles described herein.

[0088] Figure 3Illustrates a memory subsystem in the context of a computer system having multiple processing cores.

[0089] Figure 4 Illustrates a shared memory including blocks of tasks allocated on a per - slice basis to workgroups.

[0090] Figure 5 Illustrates the allocation of shared memory by a resource allocator to tasks executed at a SIMD processor as workgroups.

[0091] Figure 6 Illustrates the allocation of shared memory by a resource allocator configured according to a particular embodiment described herein.

[0092] Figure 7 Is a schematic diagram of a second memory subsystem including a resource allocator configured according to the principles described herein.

[0093] Figure 8 Illustrates the correspondence between virtual memory blocks and the memory portions of underlying shared memory.

[0094] Figure 9 Is a flowchart illustrating the operation of a first memory subsystem configured according to the principles described herein.

[0095] Figure 10 Is a flowchart illustrating the operation of a second memory subsystem configured according to the principles described herein.

[0096] Figure 11 Is a schematic diagram of an integrated circuit manufacturing system. Detailed Description

[0097] The following description is given by way of example so that those skilled in the art can make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art. The embodiments are described only by way of example.

[0098] The term "task" is used herein to refer to a set of data items and the work to be performed on those data items. For example, in addition to a set of data to be processed according to a program (e.g., the same sequence of ALU instructions or references to those ALU instructions), a task may include a program or be associated with a program or may also include a reference to a program, where the set of data may include one or more data elements (or data items, e.g., multiple pixels or vertices).

[0099] The term "program instance" is used here to refer to an individual instance of code walking through. Thus, a program instance refers to a single data item and a reference to a program to be executed on the data item (e.g., a pointer). Thus, a task can be considered to include multiple program instances (e.g., up to 32 program instances), although in practice each task only requires a single instance of the common program (or reference). A group of tasks sharing a common purpose share local memory and can execute the same program (although they can execute different parts of the program) or compatible programs on different data blocks can be connected by a group ID. A group of tasks with the same group ID can be called a "workgroup" (thus, the group ID can be called the "workgroup ID"). Thus, there is a hierarchy of terms, where a task includes multiple program instances and a group (or workgroup) includes multiple tasks.

[0100] Some processing systems that provide shared memory can include one or more SIMD (Single Instruction Multiple Data) processors, each SIMD processor being configured to execute the tasks of a workgroup in parallel. Each task of the workgroup may require an allocation of shared memory. However, generally, there are certain interdependencies among the processing of the tasks executed at the SIMD processor. For example, for some workgroups, a barrier synchronization point is defined where all tasks must arrive in order to continue processing any of these tasks beyond the barrier synchronization point and to complete the processing of the workgroup as a whole. Sometimes more than one barrier synchronization point can be defined for a workgroup, where all tasks need to arrive at a given barrier in order to continue processing any of these tasks beyond the barrier.

[0101] For traditional allocation mechanisms of shared memory, the use of barrier synchronization points can lead to deadlocks in the processing of workgroups. This is shown in Figure 1 a simple example where Figure 1It is shown that the shared memory 100 is divided into multiple memory sections 101 and a SIMD processor 102 having five processing elements 109 - 113 is configured to process a workgroup 'A' including five tasks 103 - 107 in parallel. In this example, only the processing of four tasks is enabled and all these tasks have reached the barrier synchronization point 108. This is because each task requires two memory sections of a contiguous block in the shared memory (marked with the letter 'A'), but once the blocks in the shared memory are allocated for the first four tasks, it is impossible to allocate two memory sections of a contiguous block for the fifth task. The memory sections belonging to the second workgroup are marked with the letter 'B'. A number of memory sections 114 in the shared memory are available, but these memory sections are fragmented and cannot provide a contiguous block of memory. As a result, a deadlock has occurred: the processing of four tasks 103 - 106 has stopped, waiting for the fifth task 107 to reach the barrier synchronization point 108, but the processing of the fifth task 107 cannot be started because its request to allocate the shared memory cannot be satisfied.

[0102] Such a deadlock can cause significant delays and generally can only be resolved when the processing of the deadlocked workgroup reaches a predetermined timeout. This not only wastes the time spent waiting for the timeout, but also partial processing has been performed on the tasks that have reached the barrier synchronization point (this processing will need to be repeated). The situation of deadlocks is particularly problematic in systems including multiple SIMD processors and all these SIMD processors sharing a common memory, because when the available shared memory is limited or when tasks require relatively large contiguous blocks of memory, more than one workgroup can be locked up at a time.

[0103] Avoid Fragmentation

[0104] Figure 2FIG. is a schematic diagram showing a memory subsystem 200 configured to allocate shared memory in a manner that addresses the above problems. The memory subsystem 200 includes a resource allocator 201 and a shared memory 202, which is divided into a plurality of memory portions 203. The resource allocator may be configured to receive memory requests from a plurality of requesters 205. For example, each task running at a processing unit may be a requester, each processing unit may be a requester, or each type of processing unit may be a requester. Different requesters may have different hardwired input terminals to the resource allocator, and the allocator needs to arbitrate between these hardwired input terminals. The resource allocator may service memory requests from the requesters in any suitable manner. For example, memory requests may be serviced based on simple polling or according to a set of arbitration rules for selecting the next requester to be serviced. In other embodiments, memory requests from a plurality of requesters may be aggregated at a request queue, where memory allocation requests are received into the request queue, and the resource allocator may receive memory requests to be serviced from the request queue.

[0105] The processing unit may access the allocated portion of the shared memory through an input / output arbiter 204. The I / O arbiter 204 arbitrates access to the shared memory among a plurality of units (e.g., processing units or other units capable of accessing the shared memory) that submit access requests (e.g., read / write) through an interface 211, where the interface 211 may include one or more hardwires, and the I / O arbiter arbitrates between the one or more hardwires.

[0106] Figure 3 FIG. shows the memory subsystem 200 in the context of a computer system 300 that includes a plurality of processing units 301 and a job scheduler 302, where the job scheduler 302 is configured to schedule tasks it receives for processing at the processing units. One or more of the processing units may be SIMD processors. The computer system may be, for example, a graphics processing unit (GPU). The processing units may include one or more units (all of which may be SIMD processors) for performing integer operations, floating-point operations, compound arithmetic, texture address calculations, sampling operations, etc. The processing units may together form a set of parallel pipelines 305, where each processing unit represents one pipeline in the set of pipelines. The computer system may be, for example, a vector processor.

[0107] Each processing unit may include one or more arithmetic logic units (ALUs), and each SIMD processor may include one or more different types of ALUs. For example, each type of ALU is optimized for a specific type of computation. In the example where the GPU provides the processing unit 301, the processing block 104 may include a plurality of shader cores, and each shader core includes one or more ALUs.

[0108] The work scheduler 302 may include a scheduler unit 303 and an instruction decoder 304. The scheduler unit is configured to perform scheduling of tasks to be executed on the processing unit 301, and the instruction decoder is configured to decode the tasks into a form suitable for execution on the processing unit 301. It will be appreciated that the specific decoding of the tasks performed by the instruction decoder will be determined by the type and configuration of the processing unit 301.

[0109] The resource allocator 201 is operable to receive requests for allocating memory resources to tasks executed at the processing unit 301 of the computer system. Each task or processing unit may be a requester 205. One or more processing units may be SIMD processors configured to execute tasks of a workgroup in parallel. Each task of the workgroup may request the resource allocator for memory allocation (depending on the specific architecture of the system), and such requests may be made on behalf of the tasks by, for example, the work scheduler when assigning tasks to the parallel processors or by the tasks themselves.

[0110] In some architectures, tasks related to various workgroups or for execution on different types of processing units may be queued for processing in one or more common work queues, such that tasks for parallel processing as a workgroup are assigned to appropriate processing units by the work scheduler 302 over a period of time. Tasks related to other workgroups and / or processing units may be interleaved with the tasks of the workgroup. Requests for memory resources of the resource allocator may generally be made for tasks as the tasks are scheduled at the processor. The tasks themselves may not need to make requests for memory resources. For example, in some architectures, the work scheduler 302 may make requests for memory resources on behalf of the tasks it schedules at the processing unit. Many other architectures are possible.

[0111] The resource allocator is configured to allocate memory resources to the entire workgroup when it first receives a request for the allocation of a task for the workgroup. The allocation request generally indicates the size of the memory resources required. For example, a task may require a certain number of bytes or memory segments. Since the SIMD processor executes the same instructions in parallel on different source data, each task of the workgroup for execution at the SIMD processor has the same memory requirement. The resource allocator is configured to, in response to the first request for memory for a task of the workgroup that it receives, allocate a contiguous block of memory for the entire workgroup. The size of the block is at least N times the memory resources requested in the first request, where N is the number of tasks in the workgroup (e.g., 32). The number of tasks in the workgroup can be fixed for a given system and is thus known in advance to the resource allocator, and the number of tasks in the workgroup can be provided to the resource allocator together with the first request for memory resources, and / or the number of tasks in the workgroup is available to the resource allocator (e.g., as a parameter held at a data store accessible to the resource allocator).

[0112] If there is sufficient contiguous space in the shared memory for the entire workflow block, the resource allocator is configured to respond to the first received request for the memory resources for the task of the workflow by utilizing the memory resources requested by the tasks of the workflow. This enables the processing of the first task of the workgroup to begin. The resource allocator also reserves the remaining portion of the shared memory block (which will be allocated to the other tasks of the workgroup) to prevent the block from being allocated to the tasks of other workgroups. If there is not sufficient contiguous space in the shared memory for the entire workflow block, the first received memory request is rejected. This process ensures that once the first task of the workgroup receives the allocation of memory, the system can guarantee that all tasks of the workgroup will receive their memory allocations. This avoids the deadlock scenario described above.

[0113] When the resource allocator receives a request for memory resources for a subsequently received task for the same workgroup, the resource allocator allocates the required resources to these tasks from the reserved block. The resource allocator can then sequentially allocate adjacent "slices" of the reserved shared memory block to the tasks in the order in which the corresponding memory requests are received, where the first received task receives the first slice of the block. This is shown in Figure 4 shown Figure 4It is shown that the shared memory 202 includes multiple memory portions 203 and a block 402 reserved for a workgroup after receiving a first memory request for a task of the workgroup. In this example, each slice includes two memory portions. The first slice 403 of the block is assigned to the first task, and the subsequent slices of the block are reserved in the shared memory so that when the next memory request for a task of the workgroup is received, the next slice 404 is assigned, and so on until the slices 403 - 407 of the block are assigned to tasks and the workgroup can be processed to completion.

[0114] In cases where the tasks of a workgroup have a desired or predetermined order, the resource allocator can alternatively assign the memory slices of the reserved memory block to the tasks of the workgroup according to that order. In some architectures, memory requests for tasks may be received out of order. For example, even if a memory request for a third task is received before a memory request for a second task, the third slice of the block can be assigned to the third task among five tasks. In this way, a predetermined slice of the block reserved for the workgroup can be assigned to each task of the workgroup. Since the first received task may not be the first task of the workgroup, it is not necessary to assign the first slice of the block reserved for the workgroup to the first task.

[0115] In some architectures, a task may make further requests for memory during processing. These memory requests can be processed in the same way as the initial request for memory, where when the first request regarding the workgroup is received, based on the fact that each task to be processed by the SIMD processor will require the same memory resources, a block of memory resources is assigned to all tasks of the workgroup. The size of the assigned block will be at least N times the memory resources requested in the first request, where N is the number of tasks in the workgroup.

[0116] Once a certain task has been scheduled at the processing unit, the task can access its assigned memory resources at the shared memory 202 as indicated by the data path 210 (e.g., by performing reads and / or writes).

[0117] The resource allocator can be configured to immediately release the slice of the shared memory assigned to a task when the processing of each task is completed at the processing unit (i.e., it is not necessary to wait until the entire workgroup is completed). This ensures that memory portions that are no longer needed are released as soon as possible for use by other tasks. For example, the release of the last "slice" can create a block of sufficient size to be assigned to the subsequent consecutive workgroup to be processed next.

[0118] Figure 5Illustrates the allocation of shared memory to tasks according to the principles described herein. The shared memory 202 is shown to be divided into a plurality of memory sections 203, and the SIMD processor 502 has five processing elements 503 - 507. Contrary to the traditional way of shared memory allocation, this arrangement corresponds to Figure 1 the arrangement shown. The shared memory includes a memory section 508 of an existing block assigned to workgroup B (the corresponding memory section in the figure is labeled 'B'). Consider the point at which a first memory request for a task 509 of a second workgroup A is received. When the memory request for task 509 is received, a memory section of block 514 (labeled 'A') is allocated to the shared memory, as shown. Thus, as described above, in response to receiving a first memory request for a task of a workgroup, the allocation of the memory section of that block for the entire workgroup is performed. A slice 515 of two memory sections of block 514 is allocated to the first task. Slices 516 - 519 of the block are reserved for other tasks of workgroup A.

[0119] When memory requests for subsequent tasks 510 - 513 are received, each memory request can be serviced by allocating to the corresponding task and scheduling one of the reserved slices 516 - 519 for the task to be executed on the SIMD processor 502. Since all the memory required by the workgroup has been allocated in advance, the last task 513 receives slice 519 of the block. Since all tasks have received their required memory allocations, the processing of all tasks can proceed to the barrier synchronization point 520. Once all tasks have been processed to the barrier 520, the processing of all tasks can be permitted to continue to completion, so that the workgroup can be processed as a whole. Due to the method of allocating contiguous space to the entire block, the shared memory is more robust to fragmentation, and contiguous space in the shared memory is more available for allocation to new blocks.

[0120] If there is not sufficient contiguous space available for the entire block 514, the memory request for the first task of workgroup A is rejected. The rejected memory request can be handled in various ways, but generally, it can be waited until the next time the resource allocator expects to service the corresponding requester according to its defined mechanism (e.g., based on polling or according to a set of rules for arbitration between requesters). Once sufficient memory becomes available, subsequent memory allocation requests can succeed (e.g., by tasks being completed on other processing units and the corresponding memory sections being released).

[0121] The resource allocator can maintain a data structure 206 that identifies which memory blocks have been allocated to which workgroups. For example, when reserving a memory block for a workgroup, the resource allocator can add an entry to the data structure that associates the memory block with the workgroup, so that when a task for the workgroup is subsequently received, the task allocator can identify which memory slice of the block will be allocated. Similarly, the resource allocator can use the data structure 206 to identify which tasks do not belong to workgroups that already have memory blocks allocated, so as to identify the tasks for which new block allocations will be performed. The data structure 206 can have any suitable configuration, for example, it can be a simple lookup table or register.

[0122] To allocate slices of reserved blocks to subsequently received tasks of a workgroup, the resource allocator needs to know which workgroup the memory requests it receives belong to. The memory requests can include an identifier of the workgroup to which the corresponding task belongs, to allow the resource allocator to identify memory requests that belong to workgroups for which blocks have been reserved. Alternatively, the resource allocator can access a data store that identifies which tasks belong to which workgroups (e.g., a lookup table maintained by the work scheduler 302). Generally, the memory request will identify the task for which the memory request is being made, and such a task identifier can be used to look up the corresponding workgroup in the data store.

[0123] Memory requests received for workgroups that already have memory blocks allocated can preferably be processed with priority. This helps to minimize latency in the system by ensuring that once a memory has been allocated to a workgroup and resources have been reserved in the system for the completion of processing of that workgroup, all tasks of the workgroup can start processing as soon as possible and reach any defined barrier synchronization points of the workgroup. In other words, ensuring that memory can be allocated to tasks as soon as possible after reserving the memory as a block for the workgroup helps to avoid tasks that are subsequently allocated their memory slices from delaying the processing of other tasks in the workgroup that already have other slices allocated and are ready to start processing.

[0124] Any suitable mechanism for processing memory requests with priority can be implemented. For example, the resource allocator can be configured to process requests from the same requester 205 from which it received the initial request that caused the block to be allocated with priority. The requester can be made to have priority over other requesters by increasing the frequency at which the requester's memory requests are processed. The requester can be made to have priority by repeatedly processing memory requests received from the same requester after the block is allocated until a memory request for a task belonging to a different workgroup is received. The resource allocator can use the above data structure 206 to identify which memory request corresponds to a workgroup for which a memory block has been reserved, so as to identify which memory request should be processed with priority.

[0125] Note that, in some embodiments of the systems described herein, some of the tasks targeted by the memory requests received by the resource allocator will not belong to a workgroup. For example, Figure 3 One or more processing units 301 used to process tasks in Figure 3 may not be SIMD processors. The resource allocator may be configured to identify, in the memory requests it receives, tasks that do not belong to a workgroup from one or more identifiers, or to identify the lack of one or more identifiers (e.g., no workgroup identifier). Alternatively or additionally, the resource allocator may determine whether a received memory request is made for a task that does not belong to a workgroup by looking up the identifier received in the memory request in a data store (e.g., a lookup table maintained by the work scheduler 302) that identifies which tasks belong to which workgroups.

[0126] A method for allocating memory portions at the resource allocator 201 will now be described.

[0127] The resource allocator may be configured to further divide the shared memory 202 into a plurality of non - overlapping windows, each non - overlapping window including a plurality of memory portions. The length of each window may be a power of two to enable an efficient implementation of the resource allocator in binary logic circuitry. The windows may all be of the same length. The resource allocator may maintain a fine - grained status array 207 indicating which memory portions are allocated and which are not. For example, the fine - grained status array may include a one - dimensional array of bits, where each bit corresponds to a memory portion of the shared memory and the value of each bit indicates whether the corresponding memory portion is allocated (e.g., '1' indicates the portion is allocated; '0' indicates the portion is not allocated). The size of each window may be chosen such that the resource allocator can search for an unallocated memory portion in the window that is large enough to contain a contiguous block of a given size in a single clock cycle of the digital hardware implementing the resource allocator. Generally, the size of the shared memory 202 is too large for the resource allocator to search the entire memory space in one clock cycle.

[0128] The resource allocator may maintain a window pointer 208 (e.g., a register internal to the resource allocator) that identifies in which window of the shared memory the resource allocator will start a new memory allocation when receiving a first memory request for a workgroup for which memory has not yet been allocated. When receiving a memory request for the first task of a workgroup, the resource allocator checks in the current window where a contiguous memory block to be reserved for the workgroup can fit. This is referred to as a "fine check". The resource allocator uses a fine state array 207 to identify unallocated portions of memory. The resource allocator may start at the lowest location in the current window and scan in the direction of increasing memory addresses to identify the lowest possible start location of the block and minimize fragmentation of the shared memory. Any suitable algorithm for identifying a contiguous set of memory portions that can provide the memory block may be used.

[0129] When allocating a block to a workgroup, the fine state array is updated to mark all corresponding memory portions of the block as allocated. In this way, the block can be reserved for use by the tasks of the workgroup.

[0130] The resource allocator may also maintain a coarse state array 209 that indicates whether each window of the shared memory is completely unallocated, in other words, whether all memory portions of the window are unallocated. The coarse state array may be a one-dimensional bit array where each bit corresponds to a window of the shared memory and the value of each bit indicates whether the corresponding window is completely allocated. The bit value of a window in the coarse state array may be calculated as the OR reduction of all bits in the state array corresponding to the memory portions located in that window.

[0131] The resource allocator may be configured to perform a coarse check in combination with the fine check. The coarse check includes identifying from the coarse state array 209 whether the next one or more windows following the current window indicated by the window pointer 208 can provide a new memory block required by the workgroup, i.e., if the coarse state array 209 indicates that the next window is completely unallocated, then all these memory portions can be used by the block. One or more windows following the current window may be checked because in some systems the block required may be larger than the size of a window.

[0132] The resource allocator may be configured to allocate shared memory blocks to workgroups using the following fine check and coarse check:

[0133] 1. If the fine-grained check for the memory block allocated to the workgroup is successful and the memory block can be provided in the current window, the memory block is reserved for use by the workgroup (e.g., by marking the corresponding memory portion as allocated in the fine-grained status array and adding the workgroup to the data structure 206) and a slice (e.g., the first slice) of the memory block is allocated to the task for which the memory request regarding it was made. If the coarse-grained check of the memory indicates that the current window of the shared memory is completely unallocated, the coarse-grained status array can be updated to indicate that the current window is now no longer completely unallocated.

[0134] 2. If the fine-grained check for the memory block allocated to the workgroup is unsuccessful but the coarse-grained check of the memory is successful, the memory block can be provided by one or more unallocated windows following the current window and the memory block is reserved for the workgroup. The memory block can be allocated starting from the first unallocated memory portion of the unallocated memory portion that forms a contiguous partition with the next window. The allocated memory block extends into one or more next windows identified as available by the coarse-grained check.

[0135] If both the fine-grained check and the coarse-grained check fail, the memory request can be postponed and the memory request can be attempted again later (possibly on the next clock cycle or after a predetermined number of clock cycles). The resource allocator can service other memory requests before returning the failed memory request (e.g., the resource allocator can move to service the memory request from the next requester according to some rules defined for arbitrating between requesters).

[0136] Figure 6 An example allocation according to the above method is shown, where the shared memory is divided into 128 memory portions and 4 windows each including 32 memory portions. The '1' bit 601 in the fine-grained status array 207 indicates that the first 24 memory portions are allocated to the workgroup, but all higher memory portions are unallocated. The coarse-grained status array indicates that only the first window includes the allocated memory portion (see the '1' bit 602 corresponding to the first window). In this example, the workgroup needs a 48-memory-portion block of memory 603. The resource allocator can determine the need for such a block when receiving a memory request for 3 memory portions for the first task of the workgroup: it is known that the memory request has been received from a requester that is a SIMD processor for a workgroup of 16 tasks, and the resource allocator determines that the workgroup as a whole needs 48 memory portions.

[0137] Assume that the current window indicated by the window pointer of the resource allocator is the first window '0' (possibly because the first unallocated memory portion after the allocated memory portion is memory portion 24). The detailed check of the current window will fail because there are no remaining 48 memory portions in the first window '0'. However, the coarse check will succeed because the second and third windows are indicated by the coarse status array 209 as not allocated at all. Therefore, a block of 48 will be allocated from memory portion 24 of the first window until memory portion 71 of the third window. The window pointer can be updated to identify the third window as the current window because the next available unallocated memory portion is the number 72.

[0138] The resource allocator can be configured to compute an overflow metric in parallel with the detailed check and the coarse check. The overflow metric is given by the number of memory blocks required by the workgroup minus the number of unallocated memory portions at the top of the current window, so as to give a measure of the number of memory portions that will overflow the current window in the case of allocating the block starting from the lowest memory portion of the unallocated memory of consecutive blocks at the top of the current window.

[0139] The resource allocator can be configured to use the overflow metric in an additional step after steps 1 and 2 above to attempt to allocate a memory block to the workflow:

[0140] 3. If the detailed check and the coarse check fail but there are still some free memory portions at the top of the current window, use the overflow metric as the requested size and make a new attempt to allocate the memory block starting from the lowest address of the next window of the shared memory at the next opportunity (e.g., the next clock cycle). For the initial allocation attempt, the detailed check and the coarse check as described above can be performed. This additional new attempt is beneficial because although the coarse check on the next window fails, the lower partition of the next window and the free memory portions of the current window are sufficient to provide enough contiguous unallocated space to receive the memory block.

[0141] If step 3 fails, the memory request can be postponed and then the memory request can be attempted again as described above.

[0142] Once the allocation is successful, the resource allocator immediately updates the window pointer to the window containing the memory portion (which can be the current window or the next window) after the successful allocation of the memory portion block ends.

[0143] Using fine-grained and coarse-grained checks that are executed in parallel (optionally, combined with an overflow metric that is executed in parallel with the fine-grained and coarse-grained checks), enables a resource allocator to efficiently attempt to allocate shared memory to a workgroup in a minimum number of clock cycles. On appropriate hardware, the above method can enable block allocation to be performed in two clock cycles (in the first clock cycle, the fine-grained and coarse-grained checks are executed and optionally the overflow metric is calculated; in the second clock cycle, an attempt is made to allocate according to steps 1-2 above and optional step 3).

[0144] Figure 9 A flowchart showing a resource allocator allocating shared memory is shown. The resource allocator receives a shared memory resource request 901 for a first task of a workgroup (i.e., the first task among the tasks of the workgroup for which a shared memory request has been received). In response, a shared memory block is allocated to the workgroup, the size of the shared memory block being sufficient to provide the resources requested by the first task to all tasks of the workgroup. Generally, the shared memory block will be a contiguous shared memory block. A portion 903 of the shared memory block is allocated for the first task, and this step may or may not be considered part 902 of the memory allocation to the workgroup, and thus may be performed before, simultaneously with, or after allocating memory resources to the workgroup. In some embodiments, the remaining portion of the shared memory may be reserved for other tasks of the workgroup.

[0145] When a subsequent shared memory request 904 for another task of the same workgroup is received, a portion 905 of the shared memory block allocated to the workgroup is allocated to that task. In some embodiments, in response to receiving a first memory request 901 from a task of a workgroup, memory allocation for all tasks may be performed. In such embodiments, further allocation may include responding by allocating memory resources to tasks established in response to the first received memory request for the workgroup.

[0146] The allocation 907 of memory resources from the block is repeated until all tasks of the workgroup have received the required memory resources. Then, the execution 906 of the workgroup may be performed, where each task uses the shared memory resources allocated to that task in the block allocated to the workgroup.

[0147] Fragmentation Insensitivity

[0148] A second configuration of a memory subsystem that also addresses the above limitations of the prior art for allocating shared memory will now be described.

[0149] Figure 7Memory subsystem 700 is shown including resource allocator 701, shared memory 702, input / output arbiter 703, and translation unit 704. The shared memory is divided into a plurality of memory sections 705. The resource allocator may be configured to receive memory requests from a plurality of requesters 205. For example, each task running on a processing unit may be a requester, each processing unit may be a requester, or each type of processing unit may be a requester. Different requesters may have different hardwired input terminals to the resource allocator, and the resource allocator needs to arbitrate among these hardwired input terminals. The resource allocator may service memory requests from requesters on any basis. For example, memory requests may be serviced based on simple polling or according to a set of arbitration rules for selecting the next requester to be serviced. In other embodiments, memory requests from a plurality of requesters may be aggregated at a request queue, memory allocation requests may be received into this request queue, and the resource allocator may receive memory requests to be serviced from this request queue. The allocated portions of the shared memory may be accessed by processing units using input / output arbiter 703.

[0150] Memory subsystem 700 may be arranged in a computer system in the Figure 3 manner shown and described above.

[0151] In this configuration, the resource allocator is configured to allocate shared memory 702 to tasks using translation unit 704, and input / output arbiter 703 is configured to access shared memory 702 using translation unit 704. Translation unit 704 is configured to associate each task regarding a received memory request with a plurality of memory sections 705 of the shared memory. Each task generally has an identifier provided in the memory request made regarding the task. The plurality of memory sections associated with a task need not be contiguous sections in the shared memory. However, from the perspective of the task, the task will receive an allocation of a contiguous memory block. This may be achieved by the resource allocator configured to allocate a contiguous virtual memory block to the task when the same task is actually associated at the translation unit with a potentially non-contiguous set of memory sections.

[0152] When a task accesses the memory region allocated to it (e.g., via a read or write operation), the translation unit 704 is configured to translate between the logical address provided by the task and the physical address where its data is stored or to which its data will be written. Note that the virtual address provided to the task by the resource allocator does not need to belong to a constant virtual address space, and in some examples, the virtual address will correspond to a physical address at the shared memory. It is advantageous for the logical address used by the task to have an offset relative to some base address (which itself can be a zero offset or some other predetermined value) with respect to a predetermined position (typically the start position of the block) in the virtual block. Since the virtual blocks are contiguous, only the relative positions of the data in the block that needs to be accessed need to be provided to the translation unit. The memory portion in the shared memory that makes up a particular set of virtual blocks of a task can be determined by the translation unit based on the association between the task identifier in the data storage area 706 and the memory portion.

[0153] In response to receiving a request for memory resources, the resource allocator is configured to allocate to the task a set of memory portions sufficient to satisfy the request. The memory portions do not need to be contiguous or stored sequentially in the shared memory. The resource allocator is configured to cause the translation unit to associate the task with each memory portion in the set of memory portions. The resource allocator can utilize a data structure 707 indicating which memory portions are available to keep track of the unallocated memory portions at the shared memory. For example, the data structure can be a status array (e.g., the aforementioned fine-grained status array) whose bits identify whether each memory portion at the shared memory is allocated.

[0154] The translation unit can include a data storage area 706 and can be configured to associate the task with a set of memory portions by correspondingly storing in the data storage area the identifier of the task (e.g., task_id) and the physical address of each memory portion in the set of memory portions, as well as an indication of which partition of the contiguous virtual memory block each memory portion allocated to the task corresponds to (e.g., the offset of the partition in the virtual memory block or the logical address of the partition). In this way, the translation unit can, in response to the task providing its identifier and an indication of which partition of its virtual memory block it needs (e.g., the offset in its virtual block), return the appropriate physical address to the task. The offset can be the number of memory portions of the memory region relative to the base address (e.g., in the case where the memory portion is a basic allocation interval), the number of bytes or bits, or any other position indicator. The base address can be but does not have to be the memory address where the virtual memory block starts; the base address can be a zero offset. The data storage area can be an associative memory (e.g., a content-addressable memory), whose inputs can be the task identifier and an indication of the desired memory partition in its block (e.g., the memory offset relative to the base address of its block).

[0155] More generally, a translation unit can be provided at any suitable point along the data paths between the resource allocator 701 and the shared memory 702, and between the I / O arbiter 703 and the shared memory 702. More than one translation unit can be provided. Multiple translation units can share a common data storage area such as a content addressable memory.

[0156] The resource allocator can be configured to assign logical base addresses to different tasks of a workgroup to identify contiguous virtual memory blocks within a contiguous virtual memory superblock for the workgroup as a whole, where each task of the workgroup is assigned.

[0157] Figure 8 The correspondence between the memory portion 805 of the shared memory 702 assigned to a task and the virtual block 801 assigned to the task is shown. The memory portions 805 of the shared memory labeled A - F provide a lower-level storage area for the partitions 806 of the virtual contiguous memory block 801. The memory blocks A - F are non-contiguous or are stored sequentially in the shared memory and are located among the memory portions 804 assigned to other tasks. The virtual block 801 has a base address 802, and the partitions can be identified by the offset 803 of the partition within the block relative to the base address.

[0158] In one example, in response to receiving a request for memory resources, the resource allocator is configured to assign to the task a contiguous range of memory addresses starting from the physical address of the memory portion corresponding to the first partition of the block. In other words, the base address 802 of the virtual block (and Figure 8 the base address of partition 'A' or 809 in Figure 8 is the physical address of memory portion 'A' or 808 in

[0159] Figure 8 ). Assuming there is sufficient unallocated memory portion in the shared memory to service the memory request, the resource allocator assigns to the task a virtual memory block of the requested size starting from the physical memory address of the memory partition 809. Thus, from the perspective of the task, it is assigned a contiguous block of six memory portions starting from memory portion 808 in the shared memory; in reality, the task is assigned six non-contiguous memory portions labeled A, B, C, D, E, and F in the shared memory 702. Thus, the actual allocation of the memory portion for the task is labeled by the virtual block 801 it receives. In this example, the virtual block has a physical address and there is no such virtual address space.

[0159] Once virtual memory blocks have been allocated for a task, it can access the memory using the I / O arbiter 703, where the I / O arbiter generally arbitrates access to the shared memory among multiple units (such as processing units or other units capable of accessing the shared memory) that submit access requests (e.g., read / write) via an interface 708, which may include one or more hardwired connections and the I / O arbiter arbitrates among these hardwired connections.

[0160] Continuing with the above example, consider a task requesting access to a region of block 801 located in partition 'D'. The task can submit a request to the I / O arbiter 703 that includes an indication of its task identifier (e.g., task_id) and the region of the block it needs to access. The indication can be, for example, an offset 803 relative to a base address 802 (e.g., number of bits, bytes, memory partial units), or a physical address formed based on the physical base address 802 assigned to the task and the offset 803 in the block where the requested memory region is found. The translation unit will look up the task unit identifier and the desired offset / address in the data store 706 to identify the corresponding physical address that identifies the memory portion 811 in the shared memory 702. The translation unit will service the access request from the memory portion 811 in the shared memory. In the case where the memory region in the block spans more than one memory partition, the data store will return more than one physical address for the corresponding data portions.

[0161] In an implementation where the task provides the physical address of the storage area, the physical address submitted in this case refers to a storage area in the memory portion 810 assigned to another task. Thus, different tasks can ostensibly access overlapping blocks in the shared memory. However, since the translation unit 704 translates each access request, the access to the physical address submitted by the task is redirected by the translation unit to the correct memory portion (811 in this example). Thus, in this case, the logical address that references the virtual memory block is actually a physical address, but not necessarily the physical address where the corresponding data of the block is stored in the shared memory.

[0162] In a second example, in response to receiving a request for memory resources, a resource allocator is configured to allocate to a task logical memory addresses starting from a logical base address and representing a contiguous range of virtual memory blocks. The logical addresses can be allocated in any suitable manner since a given logical address will be mapped to a physical address of a corresponding underlying memory portion. For example, logical addresses in a virtual address space that is consistent across tasks can be allocated to tasks such that the virtual blocks allocated to different tasks do not share overlapping logical address ranges; a logical base address in a set of logical addresses and an offset relative to the base address that is used to refer to a location within a virtual block can be allocated to a task; or a logical base address that is a zero offset can be allocated.

[0163] The translation unit responds to an access request (e.g., read / write) received from a task via the I / O arbiter 703. Each access request can include an identifier of the task and an offset or logical address that identifies a region of the virtual block that it needs to access. When an access request is received from a task via the I / O arbiter 703, the translation unit is configured to identify a desired physical address in the data structure based on the identifier of the task unit and the offset / logical address. The offset / logical address is used to identify which memory portions associated with the task will be accessed.

[0164] For example, again consider the case where a task requests access Figure 8 to the storage area in partition 'D' of the virtual block 801 shown. The translation unit will receive from the task an access request that includes an identifier of the task and an offset 803 that points to a location in partition 'D'. Using the identifier of the task and the offset, the translation unit looks up in the data store 706 the physical address of the corresponding memory portion 811 in the shared memory (i.e., the physical address in the data store that is associated with this task and offset). The translation unit can then proceed to service the access request in the corresponding memory portion 811 (e.g., by performing the requested read / write to / from the memory portion 811).

[0165] It will be appreciated that the "fragmentation-insensitive" method described herein allows the underlying shared memory 702 to become fragmented while the tasks themselves receive contiguous memory blocks. In cases where there are a sufficient number of memory portions in the shared memory, this enables memory blocks to be allocated to tasks and workgroups that include multiple tasks even when there is not sufficient contiguous space in the shared memory to allocate a contiguous block of the required size.

[0166] Figure 10A flowchart showing a resource allocator allocating memory resources to tasks is shown. When a shared memory resource request 1001 for a task is received, a contiguous block of virtual memory 1002 is allocated for the task. The virtual memory block is associated with a physical memory address of the shared memory 1003 such that each portion of the virtual block corresponds to a lower-level portion of the physical shared memory. Once a task receives its allocated virtual memory, the task can be executed on a SIMD processor 1004, where the task uses the lower-level shared memory resources at the physical address corresponding to its allocated virtual memory block. Figure 2 , 3 The memory subsystem of 7 is shown as including a plurality of functional blocks. This is merely illustrative and is not intended to define a strict partitioning between the different logic elements of these entities. Each functional block can be provided in any suitable manner. It will be understood that the intermediate values formed by each functional block described herein need not be physically generated by the functional block at any given time, but can merely represent logical values that conventionally describe the processing performed by the functional block between its inputs and outputs.

[0167] The memory subsystem described herein can be implemented in hardware on an integrated circuit. The memory subsystem described herein can be configured to perform any of the methods described herein. In general, any of the functions, methods, techniques, or components described above can be implemented in software, firmware, hardware (e.g., fixed logic circuitry), or any combination thereof. The terms “module,” “function,” “component,” “element,” “unit,” “block,” and “logic” can be used herein generally to denote software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, unit, block, or logic represents program code that, when executed on a processor, performs a specified task. The algorithms and methods described herein can be executed by one or more processors executing code that causes the one or more processors to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disks, flash memory, hard disk memory, and other memory devices that can use magnetic, optical, or other technologies to store instructions or other data that can be accessed by a machine.

[0168] The terms "computer program code" and "computer-readable instructions" are used herein to refer to any kind of executable code for a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, byte code, code defining an integrated circuit (e.g., a hardware description language or a netlist), and code expressed in a programming language code (e.g., C, Java, or OpenCL). The executable code can be, for example, any kind of software, hardware, script, module, or library that, when properly executed, processed, interpreted, assembled, or executed in a virtual machine or other software environment, causes a processor of a computer system that supports the executable code to perform the tasks specified by the code.

[0169] A processor, computer, or computer system can be any kind of device, machine, or dedicated circuit, or a collection or part thereof, having processing capabilities such that it can execute instructions. The processor can be any kind of general-purpose or special-purpose processor, e.g., a CPU, GPU, system-on-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system can include one or more processors.

[0170] There is also a desire to cover software that defines the hardware configurations described herein, e.g., HDL (hardware description language) software for designing an integrated circuit or for configuring a programmable chip to implement a desired function. That is, a computer-readable storage medium can be provided having encoded thereon computer-readable program code in the form of an integrated circuit definition data set that, when processed in an integrated circuit manufacturing system, configures the system to manufacture a memory subsystem configured to execute any of the methods described herein, or to manufacture a memory subsystem including any of the devices described herein. The integrated circuit definition data set can be, for example, an integrated circuit description.

[0171] A method of manufacturing a memory subsystem described herein in an integrated circuit manufacturing system is provided. An integrated circuit definition data set can be provided that, when processed by the integrated circuit manufacturing system, causes the method of manufacturing the memory subsystem to be executed.

[0172] An integrated circuit definition dataset can be in the form of computer code, for example, as a netlist, code for configuring a programmable chip, as a hardware description language (including register transfer level (RTL) code) that defines an integrated circuit at any level, as a high-level circuit representation (e.g., Verilog or VHDL), and as a low-level circuit representation (e.g., OASIS(RTM) and GDSII). A computer system configured to generate a manufacturing definition of an integrated circuit in the context of a software environment can process a higher-level representation (e.g., RTL) that logically defines the integrated circuit, where the manufacturing definition includes definitions of circuit elements and rules for joining these elements to generate a manufacturing definition representing the defined integrated circuit. In the general case where a computer system executes software to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to cause the computer system configured to generate a manufacturing definition of an integrated circuit to execute the execution code that defines the integrated circuit and thereby generate the manufacturing definition of the integrated circuit.

[0173] Now refer to Figure 11 , and describe an example of an integrated circuit manufacturing system processing an integrated circuit definition dataset to configure the system to manufacture a memory subsystem.

[0174] Figure 11 An example of an integrated circuit (IC) manufacturing system 1102 that can be configured to manufacture the memory subsystems described in any of the examples herein is shown. Specifically, the IC manufacturing system 1102 includes a layout processing system 1104 and an integrated circuit generation system 1106. The IC manufacturing system 1102 is configured to receive an IC definition dataset (e.g., defining the memory subsystems described in any of the examples herein), process the IC definition dataset, and generate an IC according to the IC definition dataset (e.g., an IC definition dataset embodying the memory subsystems described in any of the examples herein). The processing of the IC definition dataset configures the IC manufacturing system 1102 to manufacture an integrated circuit embodying the memory subsystems described in any of the examples herein.

[0175] The layout processing system 1104 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout from an IC definition data set are known in the art and may include, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated (e.g., in terms of logic components such as NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). The circuit layout may be determined from the gate-level representation of the circuit by determining the location information of the logic components. This may be achieved automatically or under the control of a user to optimize the circuit layout. When the layout processing system 1104 has determined the circuit layout, it may output a circuit layout definition to the IC generation system 1106. The circuit layout definition may be, for example, a circuit layout description.

[0176] The IC generation system 1106 generates an IC from the circuit layout definition, as is known in the art. For example, the IC generation system 1106 may perform a semiconductor device assembly process to generate the IC, which may include a multi-step sequential lithography process and chemical processing steps during which the electronic circuit is gradually formed on a wafer made of semiconductor material. The circuit layout definition may be in the form of a mask, which may be used in the lithography process for generating the IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 1106 may be in the form of computer-readable code, and the IC generation system 1106 may use the computer-readable code to form an appropriate mask for generating the IC.

[0177] The different processes that may be performed by the IC manufacturing system 1102 may all be in one location, e.g., implemented by one party. Alternatively, the IC manufacturing system 1102 may be a distributed system, so that some processes may be performed at different locations and may be performed by different parties. For example, some of the following stages may be performed at different locations and / or by different parties: (i) synthesizing the RTL code representing the IC definition data set to form a gate-level representation of the circuit to be generated, (ii) generating a circuit layout based on the gate-level representation, (iii) forming a mask according to the circuit layout, and (iv) assembling an integrated circuit using the mask.

[0178] In other examples, the processing of the IC definition data set at the integrated circuit manufacturing system may configure the system to manufacture a memory subsystem without processing the IC definition data set to determine a circuit layout. For example, the IC definition data set may define the configuration of a reconfigurable processor (e.g., an FPGA), and the processing of the data set may configure the IC manufacturing system to generate a reconfigurable processor with the defined configuration (e.g., by loading the configuration data into the FPGA).

[0179] In some embodiments, when the integrated circuit manufacturing definition dataset is processed in an integrated circuit manufacturing system, it can cause the integrated circuit manufacturing system to generate the devices described herein. For example, by configuring the integrated circuit manufacturing system in the manner described for Figure 11 the devices described herein can be manufactured.

[0180] In some examples, the integrated circuit definition dataset can include software that runs on hardware defined at the dataset or interfaces with the hardware defined at the dataset. In Figure 11 the example shown, the IC generation system can be further configured by the integrated circuit definition dataset to load hardware into the integrated circuit or provide program code for the integrated circuit to use according to the program code defined at the integrated circuit definition dataset when manufacturing the integrated circuit.

[0181] The graphics processing system described herein can be implemented in hardware on an integrated circuit. The graphics processing system described herein can be configured to perform any of the methods described herein.

[0182] The implementation of the concepts set forth in this application in devices, apparatuses, modules, and / or systems (and the methods implemented herein) can result in performance improvements compared to known implementations. Performance improvements can include one or more of increased computing performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of these devices, apparatuses, modules, and systems (e.g., in an integrated circuit), a balance can be struck between performance improvements and physical implementation to improve the manufacturing method. For example, a balance can be struck between performance improvements and layout area to match the performance of known implementations but use less silicon. This can be achieved, for example, by reusing functional blocks in a sequential manner or sharing functional blocks among the elements of the device, apparatus, module, and / or system. Conversely, the concepts of this application that result in improvements in the physical implementation of the device, apparatus, module, and system (e.g., reduced silicon area) can be traded for improved performance. This can be achieved, for example, by manufacturing multiple instances of the module within a predetermined area budget.

[0183] The applicant separately discloses each individual feature described herein, as well as any combination of two or more such features, such that these features or combinations can be implemented based on the common general knowledge of those skilled in the art from the specification, regardless of whether these features or combinations of features solve any of the problems disclosed herein. In view of the above description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present invention.

Claims

1. A memory subsystem for a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units for processing workgroups including SIMD tasks, the memory subsystem including: A shared memory divided into a plurality of memory portions for allocation to tasks to be processed by the processor; A translation unit configured to associate a task with one or more physical addresses of the shared memory; And A resource allocator configured to, in response to receiving a memory resource request for a first memory resource for a first received task of a workgroup, reserve a contiguous virtual memory block of sufficient memory portions for the workgroup such that each task of the workgroup receives a memory resource equivalent to the first memory resource, allocate the requested memory resource from the memory portions reserved for the workgroup to the first received task, and cause the translation unit to associate the first received task and subsequent tasks of the workgroup with a plurality of physical addresses of memory portions corresponding to partitions of the contiguous virtual memory block, the memory portions together implementing a complete contiguous virtual memory block.

2. The memory subsystem according to claim 1, wherein, At least some of the memory portions of the contiguous virtual memory block are non - contiguous in the shared memory.

3. The memory subsystem according to claim 1 or 2, wherein The contiguous virtual memory block includes a base address that is a logical address of the contiguous virtual memory block.

4. The memory subsystem according to claim 1 or 2, wherein, The translation unit is operable to subsequently service an access request received from the first received task regarding the contiguous virtual memory block, the access request including an identifier of the first received task and an offset of a storage area in the contiguous virtual memory block, the translation unit being configured to service the access request in a memory portion corresponding to the partition of the contiguous virtual memory block indicated by the offset.

5. The memory subsystem according to claim 1 or 2, wherein, The translation unit includes: a content - addressable memory configured to, in response to receiving an identifier of an item and an offset in the contiguous virtual memory block, return one or more corresponding physical addresses.

6. The memory subsystem according to claim 1 or 2, wherein The resource allocator is configured to allocate contiguous virtual memory blocks to each task of the workgroup such that the contiguous virtual memory blocks allocated to the tasks of the workgroup together represent a contiguous virtual memory block of a contiguous superblock.

7. The memory subsystem according to claim 6, wherein, The resource allocator is configured to, in response to subsequently receiving a memory resource request for a second task of the workgroup, allocate a contiguous virtual memory block from the contiguous superblock to the second task.

8. The memory subsystem according to claim 1 or 2, wherein, The resource allocator maintains a data structure that identifies which of the one or more workgroups are currently allocated contiguous virtual memory blocks.

9. The memory subsystem according to claim 8, wherein, The data structure includes a fine - grained status array arranged to indicate whether each memory portion of the shared memory is allocated to a task.

10. A method for allocating shared memory resources to tasks to be executed in a single instruction multiple data (SIMD) processor, the SIMD processor including a plurality of processing units respectively configured to process workgroups including a plurality of SIMD tasks, the method including: Receive a memory resource request for a first memory resource for a first receiving task of a workgroup; Reserve a contiguous virtual memory block of a sufficient memory portion for the workgroup, such that each task of the workgroup receives a memory resource equivalent to the first memory resource; Allocate the requested memory resource from the memory portion reserved for the workgroup to the first receiving task; And Associate the first receiving task and subsequent tasks of the workgroup with a plurality of physical addresses of memory portions of a shared memory, each memory portion corresponding to a partition of the contiguous virtual memory block, the memory portions together implementing the complete contiguous virtual memory block.

11. The method according to claim 10, wherein, At least some of the memory portions in the memory portion of the contiguous virtual memory block are non - contiguous in the shared memory.

12. The method according to claim 10 or 11, wherein The contiguous virtual memory block includes a base address that is a logical address of the contiguous virtual memory block.

13. The method according to claim 10 or 11, further comprising: Subsequently receive an access request received from the first receiving task regarding the contiguous virtual memory block, the access request including an identifier of the first receiving task and an offset of a storage area in the contiguous virtual memory block; And Service the access request by accessing a memory portion corresponding to the partition of the contiguous virtual memory block indicated by the offset.

14. The method according to claim 10 or 11, wherein Reserving for the workgroup includes: reserving a contiguous virtual memory block for each task of the workgroup such that the contiguous virtual memory blocks allocated to the tasks of the workgroup together represent a contiguous virtual memory block of a contiguous superblock.

15. The method according to claim 14, further comprising: In response to subsequently receiving a memory resource request for a second task of the workgroup, allocate a contiguous virtual memory block from the contiguous superblock to the second task.

16. The method according to claim 10 or 11, further comprising: Maintain a data structure that identifies which of the one or more workgroups are currently allocated contiguous virtual memory blocks.

17. The method according to claim 16, wherein, The data structure includes a fine - grained status array that is arranged to indicate whether each memory portion of the shared memory is allocated to a task.

18. A non - transitory computer - readable storage medium having stored thereon a computer - readable description of an integrated circuit, which when processed in an integrated - circuit manufacturing system causes the integrated - circuit manufacturing system to manufacture the memory subsystem of claim 1 or claim 2.

19. An integrated - circuit manufacturing system, comprising: A non - transitory computer - readable storage medium having stored thereon a computer - readable integrated - circuit description of the memory subsystem of claim 1 or claim 2; A layout processing system configured to process the integrated - circuit description to generate a circuit layout description of an integrated circuit implementing the memory system; And An integrated - circuit generation system configured to manufacture the memory subsystem according to the circuit layout description.

20. A non - transitory computer - readable storage medium having stored thereon computer - readable instructions that, when executed at a computer system, cause the computer system to perform the method of claim 10 or claim 11.