Multi-channel memory for expanding local memory

The integration of external global memory with auxiliary DRAM scratch pads addresses bandwidth limitations in memory-intensive systems, enhancing capacity and efficiency in neural network accelerators.

JP2025522956APending Publication Date: 2025-07-17INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025500825
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-22
Filing Date
2023-07-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing memory-intensive systems, such as those in neural network designs, face bandwidth constraints due to limited capacity of on-chip scratchpads, leading to increased replenishment frequency and processing overheads.

Method used

A memory system architecture that includes a global memory device external to the chip, coupled with multiple auxiliary scratch pads configured as integrated multi-channel DRAM devices, expanding the scratchpad capacity and managing data flow through software, thereby reducing latency and improving processor utilization.

Benefits of technology

The solution enhances memory capacity and reduces bandwidth constraints, improving processing performance and energy efficiency by optimizing data movement and utilization of processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522956000001_ABST
    Figure 2025522956000001_ABST
Patent Text Reader

Abstract

Memory system, method of assembling a memory system, and computer system. The memory system includes a global memory device coupled to a plurality of processing elements. The global memory device is disposed outside a chip in which a plurality of processing devices exist. The memory system also includes at least one main scratch pad coupled to at least one processing element among the plurality of processing devices and the global memory device. The memory system further includes a plurality of auxiliary scratch pads coupled to the plurality of processing elements and the global memory device. One or more of the auxiliary scratch pads are configured to store static tensors. At least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a random access memory (RAM) device, and more particularly, to an expansion of the capacity of a main on-chip scratchpad for a processing core.

Background Art

[0002] At least some known dedicated accelerators, such as, but not limited to, deep learning accelerators, include a set of interconnected processing cores within a multi-core processing device, where each core can quickly access its working set using local memory or a scratchpad, and access global memory to periodically or continuously replenish the contents of the scratchpad. The scratchpad is typically implemented as a static random access memory (SRAM) array with low latency and high bandwidth performance, while global memory is typically implemented as a dynamic random access memory (DRAM) array external to the chip that includes multiple processor cores. The period of scratchpad data replenishment is determined by the relationship between the size of the application's working set and the physical capacity of the scratchpad. The larger the scratchpad, the lower the replenishment frequency.

Summary of the Invention

[0003] A system and method are provided for expanding the capacity of a main on-chip scratchpad for a processing core.

[0004] In one aspect, a memory system is presented that is configured to expand the capacity of a plurality of main scratch pads for each of a plurality of processing cores. The memory system includes a global memory device coupled to a plurality of processing elements. The global memory device is disposed external to the chip on which the plurality of processing elements reside. The memory system also includes at least one main scratch pad coupled to at least one of the plurality of processing elements and the global memory device. The memory system further includes a plurality of auxiliary scratch pads coupled to the plurality of processing elements and the global memory device. At least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device. Thus, bandwidth constraint relaxation is implemented through the device, overcoming the capacity constraints of existing memory-intensive systems such as those implemented in neural network designs.

[0005] In another aspect, a method of assembling a memory system configured to expand the capacity of a plurality of main scratch pads for each of a plurality of processing cores is presented. The method includes placing a plurality of processing elements on a chip and coupling a global memory device to the plurality of processing elements. The global memory device is disposed external to the chip. The method also includes coupling at least one main scratch pad to at least one of the plurality of processing elements and the global memory device. The method further includes coupling a plurality of auxiliary scratch pads to the plurality of processing elements and the global memory device. At least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device. Thus, bandwidth constraint relaxation is implemented through the device, overcoming the capacity constraints of existing memory-intensive systems such as those implemented in neural network designs.

[0006] In yet another aspect, a computer system is presented that is configured to expand the capacity of a plurality of main scratch pads for each of a plurality of processing cores. The computer system includes a plurality of processing devices disposed on a chip. Each of the one or more processing devices includes one or more processing elements. The computer system also includes at least one main scratch pad coupled to the one or more processing elements. The computer system further includes a global memory device coupled to the plurality of processing devices. The global memory device is disposed external to the chip, and the global memory device is coupled to at least one main scratch pad. The computer system also includes one or more auxiliary scratch pads coupled to the plurality of processing devices and the global memory device. At least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device. Thus, bandwidth constraint relaxation is implemented through the device, and through the device that overcomes the capacity constraints of existing memory-intensive systems such as those implemented in neural network designs, bandwidth constraint reduction is implemented.

[0007] This summary is not intended to represent every aspect of every implementation and / or every embodiment of the present disclosure. These and other features and advantages will become apparent from the following detailed description of the embodiments in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0008] The drawings included in this application are incorporated herein and form a part hereof. They illustrate embodiments of the disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are exemplary of particular embodiments and do not limit the disclosure.

[0009]

Figure 1A

[0010]

Figure 1B

[0011]

Figure 2A

[0012]

Figure 2B

[0013]

Figure 3A

[0014]

Figure 3B

[0015]

Figure 4

[0016]

Figure 5

[0017]

Figure 6

[0018]

Figure 7

[0019] The present disclosure is subject to various modifications and alternative forms, and specific details thereof are shown by way of example in the drawings and will be described in detail below. However, it should be understood that the intention is not to limit the present disclosure to the specific embodiments described. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the present disclosure are intended to be included.

Embodiments for Carrying Out the Invention

[0020] Aspects of the present disclosure relate to expanding the capacity of the main on-chip scratchpad for processing cores. The present disclosure is not necessarily limited to such applications, but various aspects of the present disclosure can be understood through discussion of various examples using this context.

[0021] It will be readily understood that the components of the embodiments generally described herein and shown in each figure may be arranged and designed in a wide variety of different configurations. Accordingly, the following detailed description of the embodiments of the apparatus, system, method, and computer program product of the present embodiment, as presented in each figure, is not intended to limit the scope of the claimed embodiments, but merely represents selected embodiments.

[0022] Throughout this specification, references to "a select embodiment", "at least one embodiment", "one embodiment", "another embodiment", "other embodiments", or "an embodiment", and similar language, mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, the phrases "a select embodiment", "at least one embodiment", "in one embodiment", "another embodiment", "other embodiments", or "an embodiment" that appear in various places throughout this specification are not necessarily referring to the same embodiment.

[0023] The illustrated embodiments are best understood by reference to the drawings, wherein like parts are designated by like reference numerals throughout. The following description is intended only by way of example and merely shows certain selected embodiments of devices, systems, and processes corresponding to the embodiments claimed herein.

[0024] As used herein, "facilitating" an action includes performing the action, making the action easier, assisting in performing the action, or causing the action to be performed. Thus, by way of example and not limitation, instructions executed on a processor may facilitate an action performed by a semiconductor processing apparatus by sending appropriate data or commands to cause the action to be performed or to assist in performing the action. Even when an agent facilitates an action other than by performing the action, the action is performed by some entity or combination of entities.

[0025] Generally, as the robustness of recent computing systems has increased, when performance is limited by the speed and efficiency of operations using a memory bus, additional memory strategies are implemented to reduce access time. The use of a memory controller reduces access time somewhat, but typically does not overcome all of the limitations associated with a memory bus that is relatively distant from the processing device. Simply increasing the amount of accessible memory does not reduce long access times. Further, defining multiple separate channels between the memory and the processing device tends to reduce latency. However, using software features to manage the data flow through the various channels may tend to increase latency again. Other known solutions include using three-dimensional stacked DRAM (3D DRAM) to localize communication between columns of memory elements and their underlying processor cores, thereby improving access time. However, some of these solutions use multiplexing devices that increase access time somewhat for some operations.

[0026] More specifically, at least some known dedicated accelerators, such as, but not limited to, deep learning accelerators, include a set of interconnected processing cores within a multi-core processing device, where each core can quickly access its working set using local memory, or a scratch pad, and can also access a global memory device to periodically or continuously replenish the contents of the scratch pad. The working set is the memory that a process needs during a given period. For example, with respect to the size of the working set, in a 100 gigabyte (GB) main memory system, a particular application may only require 50 megabytes (MB) per second of operation. Thus, in such a scenario, 50MB is the size of the working set.

[0027] The scratch pad is typically implemented as a static random access memory (SRAM) array directly placed on each chip or processing core, which has low latency and high bandwidth performance. Here, the proximity of the main scratch pad and the processing elements facilitates the low-latency aspect of the on-chip device. In contrast, global memory is typically implemented as a dynamic random access memory (DRAM) array outside the chip containing multiple processor cores. The period of data replenishment of the scratch pad is determined by the relationship between the working set size of the application and the physical capacity of the scratch pad. Generally, the scratch pad is distinguished from a memory cache. The scratch pad has a more deterministic tendency with respect to the data contents present therein, thereby facilitating fetch and prefetch operations for only specific data. Thus, in some processing requirements, the scratch pad is preferred over the cache.

[0028] To reduce these overheads and improve the utilization of processing resources, at least some known efforts typically involve creating a memory hierarchy where additional tiers of memory are inserted directly between the main scratchpad and the global memory device. These tiers are typically implemented as caches of the main global memory, and replenishment is achieved by continuously moving content from one memory tier to the next. If the movement between tiers is implemented in hardware, special circuitry is provided to determine which locations within each tier should be moved and when they should be moved to maximize the utilization of processing resources. Additional hardware is used to both implement the memory tiers on the same chip as the processor and control the transmission between the tiers, and this hardware can become quite complex. Furthermore, the additional hardware components typically do not recognize the intent of the programs running on the processor and thus cannot maximize processor utilization. They also consume valuable real estate on the processor chip that could otherwise be used to increase the processing power of the chip.

[0029] Furthermore, some known efforts to mitigate these overheads and improve utilization of processing resources typically involve expanding the size of the main scratch pad. In general, the larger the main scratch pad, the less frequent data replenishment is required. More frequent replenishment promotes increased processing overhead, both in terms of the power required to transmit data back and forth and in terms of the time required for these transmissions. Furthermore, the time required for this transmission inevitably results in idling of processing resources and therefore underutilization of the processor. However, simply expanding the size of the main scratch pad tends to counter efforts to reduce the size of the chip on which the main scratch pad resides. Thus, the opportunity cost of implementing the additional tiers on-chip discussed above is a similar drawback to simply expanding the size of the local scratch pad memory. Furthermore, expanding the size of the main scratch pad also increases the latency of accessing the main scratch pad, even for applications with smaller working sets that may not require larger scratch pad memory resources.

[0030] Therefore, there is a need to alleviate bandwidth constraints through devices that overcome the capacity limitations of existing memory-intensive systems such as those implemented in neural network designs.

[0031] Referring to FIG. 1A, a block schematic diagram is presented showing a memory system architecture 100 (which may be referred to herein as memory system 100) including an array multi-channel direct random access memory (DRAM) feature according to some embodiments of the present disclosure. In some embodiments, the memory system architecture 100 is integrated on a single die. In some embodiments, at least a portion of the memory system architecture 100 includes at least a portion of a neural network accelerator chip 101. However, the memory system architecture 100 is not limited to such an implementation and is configured to be utilized in any implementation and device architecture that enables the operation of the memory system architecture 100 described herein.

[0032] In some embodiments, the memory system architecture 100 includes a global memory device 102. In at least some embodiments, the global memory device 102 is implemented as a dynamic random access memory (DRAM) device. In some embodiments, the global memory device 102 is implemented as a DRAM array. In some embodiments, the global memory device 102 is on the same die as the neural network accelerator chip 101. In some embodiments, the global memory device 102 is integrated on a separate chip from the neural network accelerator chip 101 (as shown in FIG. 1A).

[0033] In one or more embodiments, the neural network accelerator chip 101 includes a memory controller 104 communicatively and operably coupled to a global memory device 102. The memory controller 104 is a device that manages the flow of data transmitted to and from the global memory device 102. In at least some embodiments, the neural network accelerator chip 101 includes an on-chip interconnect 106 communicatively and operably coupled to the memory controller 104 through an interconnect conduit 108. Further, in some embodiments, the neural network accelerator chip 101 includes a plurality of processing cores 110 ranging from processing core 110-1 to 110-n, where processing core 110-n-1 is read as the "n minus first" processing core, and where the variable "n" has any value that enables the operation of the neural network accelerator chip 101, including, but not limited to, values from 1 to 16. Each processing core 110 is communicatively coupled to the on-chip interconnect 106 through an on-chip interconnect channel 112 (conceptually shown by dashed lines and only one is labeled in FIG. 1A for clarity). The processing cores 110 are further discussed with respect to FIG. 1B.

[0034] Further, neural network accelerator chip 101 includes a multi-channel memory array 120, referred to herein as a multi-channel auxiliary DRAM array 120. The multi-channel auxiliary DRAM array 120 is configured to extend global memory device 102. In some embodiments, the multi-channel auxiliary DRAM array 120 is a monolithic device (as shown in FIG. 1A). However, in some embodiments, the multi-channel auxiliary DRAM array 120 is distributed across multiple auxiliary DRAM devices to enable optimal placement of other chip architecture elements. The multi-channel auxiliary DRAM array 120 includes a plurality of auxiliary DRAM elements or channels, referred to herein as auxiliary scratch pads 122, ranging from auxiliary scratch pad 122-1 to auxiliary scratch pad 122-n. Thus, neural network accelerator chip 101 includes a plurality of multi-lateral segments in the form of auxiliary scratch pads 122, where at least a portion of the plurality of auxiliary scratch pads 122 is configured as a monolithic multi-channel device, i.e., one or more multi-channel auxiliary DRAM arrays 120.

[0035] In some embodiments, each auxiliary scratch pad 122 is directly, communicatively, and operably coupled to a corresponding processing core 110 through auxiliary scratch pad channels 124 (only two are shown and labeled for clarity, i.e., 124-1 and 124-n), for example, auxiliary scratch pad 122-1 is coupled to processing core 110-1 and auxiliary scratch pad 122-n is coupled to processing core 110-n. In some embodiments, in contrast to the one-to-one relationship between auxiliary scratch pads 122 and processing cores 110, at least some of the processing cores 110 are coupled to multiple auxiliary scratch pads 122. In some embodiments, each auxiliary scratch pad channel 124 cooperates with an associated auxiliary scratch pad 122 to define each auxiliary DRAM channel. In some embodiments, the global memory device 102 and the multi-channel auxiliary DRAM array 120 are physically mounted side by side in a parallel configuration. Accordingly, the multi-channel auxiliary DRAM array 120 and the individual auxiliary scratch pads 122 are located outside each processing core (or chip) 110, thereby enabling the designer of the processing core chip 110 to use the unused portions of the processing core 110 for other features. Further, by using DRAM for the auxiliary scratch pads 122, the utilization of the high density of DRAM technology is promoted as compared to lower density SRAM or hybrid DRAM technologies. In some embodiments, such use of DRAM technology enables the auxiliary scratch pads 122 to include, for example but not limited to, a working set for large-scale deep learning inference applications targeted at natural language understanding.

[0036] Referring to FIG. 1B, a block schematic diagram is presented showing an enlarged view of a portion 150 of a memory system architecture 100, and more particularly a portion of a neural network accelerator chip 101 (shown in FIG. 1A), according to some embodiments of the present disclosure, where the reference to FIG. 1A is continued. In one or more embodiments, each processing core 110 includes a main on-chip scratch pad 152, referred to herein as the main scratch pad 152, which is coupled to the on-chip interconnect 106 through each on-chip interconnect channel 112. In some embodiments, the main scratch pad 152 at least partially defines the on-chip interconnect channels 112. In some embodiments, the main scratch pad 152 is configured as a static random access memory (SRAM) array disposed directly on each chip or processing core 110 having low latency and high bandwidth performance, where the proximity of the main scratch pad 152 and the processing element 158 facilitates the low latency aspect of the on-chip devices of the memory system architecture 100. In some embodiments, the main scratch pad 152 has any memory architecture that enables the operation of the memory system architecture 100, including but not limited to a neural network accelerator chip 101 including a DRAM device.

[0037] Furthermore, in some embodiments, each processing core 110 includes a channel controller 154 coupled to each auxiliary scratchpad 122 through each auxiliary scratchpad channel 124. The channel controller 154 is configured to manage the transmission of signals from each auxiliary scratchpad 122 to the processing element 158. In some embodiments, the channel controller 154 at least partially defines the auxiliary scratchpad channel 124. Also, in some embodiments, each processing core 110 includes a multiplexor (MUX) 156 configured to direct selected signals from each main scratchpad 152 and each auxiliary scratchpad 122 (through the channel controller 154) to the processing element 158 for a desired processing operation. Thus, each auxiliary scratchpad 122 is physically located outside the chip having the processing core 110 and is mapped to the processing element 158 in parallel with each main scratchpad 152.

[0038] In some embodiments, by implementing the memory features of the auxiliary scratchpad 122 to extend the memory features of the main scratchpad 152, the movement of the contents of the memory and the contents coming and going from each processing core 110 can be fully orchestrated within the software, and thus it is possible to approach the theoretical maximum value of the utilization rate of the processing resources. In some embodiments, as a result of the expansion of the memory system described herein, the performance of the neural network accelerator chip 101 is improved, and thus a reduction in the total energy consumption of the processing core 110 can be realized.

[0039] In at least some embodiments, the auxiliary scratch pad 122 is configured to facilitate the processing of static tensors and dynamic tensors within a neural network, such as a neural network found within a deep learning platform that includes, but is not limited to, the neural network accelerator chip 101 shown at least partially in FIGS. 1A and 1B. Further, the auxiliary scratch pads described herein are suitable for use in a broader artificial intelligence platform that includes, but is not limited to, a machine learning platform. Generally, a tensor represents a multi-dimensional array that includes elements of a single data type configured for use in any numerical computation. Thus, in at least some embodiments, a deep learning tensor is configured as one or more matrices represented using an n-dimensional array (where the term "n-dimensional" as used with respect to the tensor is not associated with the variables of the processing core 110 and the auxiliary scratch pad 122 described herein). In one non-limiting example, numerical data in the form of 27 values is sequentially stored one value at a time in a single array implemented as a single contiguous block in memory as a 3×3×3 tensor, where 3 dimensions is also not limiting. In some embodiments, a tensor is static with respect to shape, i.e., the number of dimensions within the matrix and the extent of each dimension do not change over time. Some non-limiting examples of values for static tensors include read-only values that are utilized during the definition of each computational graph of the nodes within a neural network corresponding to its internal operations or variables (sometimes referred to as inference time, where the model remains fixed). In contrast, in some embodiments, a tensor is dynamic, i.e., the values within these dynamic tensors are not fixed, and these dynamic tensors are generated as intermediate or final outputs, where the duration of these dynamic tensors is not necessarily the entire execution of the program. In some embodiments, for example, but not limited to, the matrix configuration with respect to the number of dimensions within the array and the extent of each dimension is subject to change during the execution of each computational graph.

[0040] In at least some embodiments, the auxiliary scratch pad 122 is configured to store a static deep learning tensor that includes, but is not limited to, weight values for facilitating the operation of each neural network. In contrast, the main scratch pad 152 is also configured to hold a dynamic tensor that is utilized, for example, but not limited to, to facilitate node activation for the operation of each neural network. In some embodiments, the auxiliary scratch pad 122 is configured with a lower read bandwidth than the read bandwidth for the main scratch pad 152, where the more restricted static content of the auxiliary scratch pad 122 does not require the larger read bandwidth typically found in the main scratch pad 152. Further, at least partially due to the aforementioned characteristics of the data stored in the auxiliary scratch pad 122, the write bandwidth of the auxiliary scratch pad 122 is lower than the read bandwidth of the auxiliary scratch pad 122. Thus, in such embodiments, the read / write bandwidth of the auxiliary scratch pad 122 need not be as large as the read / write bandwidth of the main scratch pad 152.

[0041] In some embodiments, the read / write bandwidth of the auxiliary scratch pad 122 is scaled to accommodate the storage of dynamic tensors, as required by embodiments of the memory system 100 and the neural network accelerator chip 101 that require such a configuration.

[0042] In one or more embodiments, the neural network accelerator chip 101 includes a one-to-one relationship between the auxiliary scratchpad 122 and the processing core 110. In some embodiments, the neural network accelerator chip 101 includes a many-to-one relationship between the auxiliary scratchpad 122 and the processing core 110, where one of the limiting factors includes the available amount of remaining space with respect to the physical landscape of the neural network accelerator chip 101. Thus, in some embodiments, multiple auxiliary scratchpads 122 are coupled to each MUX 156 through one of the individual respective channel controllers 154 or an integrated multi-channel controller (not shown). In some embodiments, the processing element 158 is configured to accommodate any number of auxiliary scratchpads 122, including, but not limited to, from 1 to 16 auxiliary scratchpads 122.

[0043] Referring to FIG. 2A, a block schematic diagram is presented showing a memory system architecture 200 (which may be referred to herein as memory system 200) including multi-channel DRAM features, according to some embodiments of the present disclosure. In many embodiments, the memory system 200 is similar to the memory system 100 (shown in FIGS. 1A and 1B), with one difference being that a global channel 230 (only one is labeled) has been added. Also referring to FIG. 2B, a block schematic diagram is presented showing an enlarged view of a portion 250 of the memory system architecture 200 shown in FIG. 2A, according to some embodiments of the present disclosure. Components that are similarly numbered in FIGS. 1A, 1B, 2A, and 2B have similar functionality and are similarly named.

[0044] The global channel 230 is communicatively coupled to the on-chip interconnect 206, which is communicatively coupled to the global memory device 202. Further, the global memory channel 230 is communicatively and operably coupled to each of the processing cores 210, thereby directly coupling the global memory device 202 to the processing cores 210. Such direct coupling facilitates instances where there are data calls in which each processing element 258 seeks data present in other parts of the memory system 200.

[0045] In at least some embodiments, the global channel 230 is coupled to a channel controller 254 to facilitate management of the flow of information entering and leaving the MUX 256 for the processing elements 258. In some embodiments, the global channel 230 is coupled to the processing elements 258 through any mechanism that enables the operation of the memory system 200 described herein.

[0046] Referring to FIG. 3A, a block schematic diagram is presented showing a memory system architecture 300 (which may be referred to herein as memory system 300) including multi-channel DRAM features, according to some embodiments of the present disclosure. Components similarly numbered in FIGS. 1A, 1B, 2A, and 2B have similar functionality and are similarly named, and reference to FIGS. 1A, 1B, 2A, and 2B is continued.

[0047] In at least some embodiments, at least a portion of the memory system 300 includes one or more neural network accelerator chips in a manner similar to the network accelerator chips 101 / 201. In some embodiments, the memory system 300 includes an on-chip interconnect 306 that is substantially similar to the on-chip interconnects 106 / 206. In at least some embodiments, the memory system 300 includes a global memory device (substantially similar to the global memory devices 102 / 202), and the global memory device is communicatively coupled to the on-chip interconnect 306 through a memory controller (substantially similar to the memory controllers 104 / 204) and through an interconnect conduit 308 (substantially similar to the interconnect conduits 108 / 208), where the global memory device and the memory controller are not shown in FIG. 3A for clarity.

[0048] In one or more embodiments, the memory system 300 includes a plurality of dielets 340, namely the first dielet 340-1, the second dielet 340-2, etc. up to the m-th dielet 340-m, where the variable “m” has any value that enables the operation of the memory system 300, including values from 1 to 16, although not limited thereto. In some embodiments, the dielets 340 are present on the neural network accelerator chip (not shown in FIG. 3A) described above. Each dielet 340 includes one or more processing cores 310, where, for example, two processing cores 310-1-1 and 310-1-2, which are non-limiting numbers, are shown for the first dielet 340-1 in FIG. 3A. Similarly, the second dielet 340-2 includes, for example, two processing cores 310-2-1 and 310-2-2, which are non-limiting numbers, and the m-th dielet 340-m includes, for example, two processing cores 310-m-1 and 310-m-2, which are non-limiting numbers. In some embodiments, each dielet 340 includes any number of processing cores 310 that enables the operation of the memory system 300 described herein, including from 1 to 16 processing cores 310, although not limited thereto. Each processing core 310 is communicatively coupled to the on-chip interconnect 306 through an on-chip interconnect channel 312 (conceptually shown by dashed lines and only one is labeled in FIG. 3A for clarity). The processing cores 310 will be further discussed with respect to FIG. 3B.

[0049] Further, in at least some embodiments, each die 340 includes one or more auxiliary DRAM elements, or channels, referred to herein as auxiliary scratch pads 322. As shown in FIG. 3A, each die 340 includes one or more auxiliary scratch pads 322, where, for example, two non-limiting auxiliary scratch pads 322-1-1 and 322-1-2 are shown for the first die 340-1 in FIG. 3A. Similarly, the second die 340-2 includes, for example, two non-limiting auxiliary scratch pads 322-2-1 and 322-2-2, and the m-th die 340-m includes, for example, two non-limiting auxiliary scratch pads 322-m-1 and 322-m-2. In some embodiments, each die 340 includes any number of auxiliary scratch pads 322 that enable the operation of the memory system 300 described herein, including, but not limited to, from 1 to 16 auxiliary scratch pads 322.

[0050] In some embodiments, the plurality of auxiliary scratch pads 322 for each die 340 are individual elements disposed on each die 340, and thus each die 340 is an integral device. In contrast, in some embodiments, the plurality of auxiliary scratch pads 322 are formed as a separate multi-channel auxiliary DRAM array 320 (shown in FIG. 3A by the dashed line) coupled across the set of dies 340. Thus, in such embodiments, the multi-channel auxiliary DRAM array 320 includes a plurality of multi-lateral segments in the form of auxiliary scratch pads 322, where at least a portion of the plurality of auxiliary scratch pads 322 are configured as an integral multi-channel device, for example, each die 340 and the multi-channel auxiliary DRAM array 320 are such devices.

[0051] In some embodiments, each auxiliary scratch pad 322 is directly, communicably, and operably coupled to a corresponding processing core 310 through an auxiliary scratch pad channel 324 (only one is shown and labeled for clarity), e.g., auxiliary scratch pad 322-1-1 is coupled to processing core 310-1-1 and auxiliary scratch pad 322-m-2 is coupled to processing core 310-m-2. In some embodiments, in contrast to the one-to-one relationship between the auxiliary scratch pads 322 and the processing cores 310, at least some of the processing cores 310 are coupled to a plurality of auxiliary scratch pads 322. In some embodiments, each auxiliary scratch pad channel 324, in cooperation with the associated auxiliary scratch pad 322, defines each auxiliary DRAM channel.

[0052] Referring to FIG. 3B, a block schematic diagram is presented showing an enlarged view of a portion 350 of the memory system architecture 300 shown in FIG. 3A, according to some embodiments of the present disclosure. In many embodiments, portion 350 of memory system 300 is substantially similar to portion 250 of memory system 200. Components that are similarly numbered in FIGS. 2B and 3B have similar functionality and are similarly named.

[0053] Referring to FIG. 4, a block schematic diagram is presented showing a memory system architecture 400 including multi-channel DRAM features according to some embodiments of the present disclosure. Referring also to FIGS. 1A, 1B, 2A, 2B, 3A, and 3B, in at least some embodiments, the memory system architecture 400 includes at least one logical processor die 402, where the number 1 is non-limiting. In some embodiments, the logical processor die 402 includes a plurality of processing cores 410 that are substantially similar to the processing cores 110, 210, and 310. The logical processor die 402 includes any number of processing cores 410 that enable the operation of the memory systems 100, 200, 300, and 400 described herein, including but not limited to from 1 to 16. Here, 16 units of processing cores 410 are shown in FIG. 4, and eight processing cores 410 are labeled from processing cores 410-1 to 410-8.

[0054] Further, in some embodiments, the memory system architecture 400 includes a plurality of DRAM dies 420. In some embodiments, the DRAM die 420 is similar to the multi-channel auxiliary DRAM array 120, 220, or 320. The memory system architecture 400 includes any number of DRAM dies 420 that enable the operation of the memory systems 100, 200, 300, and 400 described herein, including but not limited to from 1 to 16. Here, four units of DRAM dies 420, namely DRAM dies 420-1, 420-1, 420-3, and 420-4, are shown in FIG. 4. As shown, the memory system architecture 400 defines a three-dimensional (3D) stack configuration. In some embodiments, the number of stacked DRAM dies 420 is affected by environmental conditions, including but not limited to thermal considerations regarding heat generation and removal. Further, the number of selected stacked DRAM dies 420 is affected by considerations in chip design regarding the placement of various components thereon and practical limitations based on recent chip manufacturing techniques and the structural strength requirements of the DRAM die 420.

[0055] In one or more embodiments, each DRAM die 420 includes a plurality of auxiliary scratch pads 422 that are substantially similar to the auxiliary scratch pad 122. In some embodiments, the plurality of auxiliary scratch pads 422 are substantially similar to the auxiliary scratch pads 222 and 322. In some embodiments, each DRAM die 420 includes any number of auxiliary scratch pads 422 that enable the operation of the memory systems 100, 200, 300, and 400 described herein, including but not limited to from 1 to 16. Here, in FIG. 4, 16 units of auxiliary scratch pads 422 are shown for each of the four DRAM dies 420, and eight processing cores 410 are labeled from 410-1 to 410-8. In some embodiments, the electrical connections through the stack are facilitated through a plurality of through silicon vias (TSVs) 430. The communication between the processing cores 410 and the auxiliary scratch pads 422 is described with reference to FIGS. 1A, 1B, 2A, 2B, 3A, and 3B.

[0056] As shown in FIG. 4, in some embodiments, each processor core 410 is communicatively and operably coupled to four auxiliary scratch pads 422. For example, processor core 410-1 is coupled to auxiliary scratch pads 422-1-1, 422-2-1, 422-3-1, and 422-4-1. In some embodiments, the processor core 410 is coupled to any number of auxiliary scratch pads 422 that enable the operation of the memory system architecture 400 described herein.

[0057] Referring to FIG. 5, a block schematic diagram is presented showing a memory system architecture 500 that includes multi-channel DRAM features, according to some embodiments of the present disclosure. Memory system architecture 500 is similar to memory system architecture 400 (shown in FIG. 4 and still being continuously referred to). However, memory system architecture 500 is in a parallel configuration rather than a stacked configuration. Thus, memory system architecture 500 includes at least one logic processor die 502 that is similar to logic processor die 402. Further, in some embodiments, memory system architecture 500 includes one or more DRAM dies 520, where two DRAM dies 520-1 and 520-2 are shown. DRAM die 520 is similar to DRAM die 420. Communication between processing core 510 and auxiliary scratch pad 522 is described with reference to FIGS. 1A, 1B, 2A, 2B, 3A, and 3B.

[0058] Referring now to FIG. 6, a block schematic diagram is provided showing a computing system 601 that may be used to implement one or more of the methods, tools, and modules, and any associated functions described herein (e.g., using one or more processor circuits or computer processors of a computer). In some embodiments, the main components of computer system 601 can include one or more CPUs 606, a memory subsystem 604, a terminal interface 612, a storage interface 616, an I / O (input / output) device interface 614, and a network interface 618, all of which can be communicatively coupled directly or indirectly via a memory bus 603, an I / O bus 608, and an I / O bus interface unit 610 for communication between components.

[0059] The computer system 601 can include one or more general-purpose programmable central processing units (CPUs) 606-1, 606-2, 606-3, 606-N, which are collectively referred to herein as CPU 606. In some embodiments, the computer system 601 may include multiple processors, which are typical of relatively large systems. However, in other embodiments, the computer system 601 can alternatively be a single-CPU system. Each CPU 606 can execute instructions stored in the memory subsystem 604 and can include one or more levels of on-board cache.

[0060] The system memory 604 can include a computer system readable medium in the form of volatile memory, such as random access memory (RAM) 622 or cache memory 624. The computer system 601 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 626 can be provided for reading from and writing to a non-removable non-volatile magnetic medium, such as a "hard drive". Although not shown, a magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), or an optical disk drive for reading from or writing to a removable non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media, can be provided. Further, the memory 604 can include flash memory, such as a flash memory stick drive or a flash drive. Further, the global memory devices, main scratch pads, and auxiliary scratch pads described herein are included as part of the set of memory devices described. The memory devices can be connected to the memory bus 603 by one or more data media interfaces. The memory 604 can include at least one program product having a set of (e.g., at least one) program modules configured to perform the functions of various embodiments.

[0061] Memory bus 603 is shown in FIG. 6 as a single bus structure providing a direct communication path between CPU 606, memory subsystem 604, and I / O bus interface 610. However, in some embodiments, memory bus 603 may include a plurality of different buses or communication paths, which may be arranged in any of various forms, such as point-to-point links in a hierarchical, star, or web configuration, multiple hierarchical buses, parallel and redundant paths, or any other suitable type of configuration. Further, although I / O bus interface 610 and I / O bus 608 are each shown as a single unit, in some embodiments, computer system 601 may include a plurality of I / O bus interface units 610, a plurality of I / O buses 608, or both. Further, a plurality of I / O interface units separating I / O bus 608 from various communication paths extending to various I / O devices are shown, but in other embodiments, some or all of the I / O devices may be directly connected to one or more system I / O buses.

[0062] In some embodiments, computer system 601 may be a multi-user mainframe computer system, a single-user system, or a server computer or similar device that has little or no direct user interface but receives requests from other computer systems (clients). Further, in some embodiments, computer system 601 may be implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, a network switch or router, or any other suitable type of electronic device.

[0063] Note that FIG. 6 is intended to show the representative main components of an exemplary computer system 601. However, in some embodiments, the individual components may have a higher or lower complexity than shown in FIG. 6, there may be components other than or in addition to those shown in FIG. 6, and the number, type, and configuration of such components may vary.

[0064] One or more programs / utilities 628, each having at least one set of program modules 630, may be stored in memory 604. Programs / utilities 628 may include a hypervisor (also referred to as a virtual machine monitor), one or more operating systems, one or more application programs, other program modules, and program data. Each of the operating systems, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a networking environment. Programs 628 and / or program modules 630 generally implement the functionality or methodology of various embodiments.

[0065] Referring to FIG. 7, a flowchart is provided showing a process 700 for assembling a computer system, such as computer system 600 (see FIG. 6). Referring also to FIGS. 1A, 1B, 2A, 2B, 3A, 3B, 4, 5, and 6, process 700 includes a step 702 of placing a plurality of processing elements 158 / 258 / 358 on chip 101 / 201. Process 700 also includes a step 704 of coupling a global memory device 102 / 202 to the plurality of processing elements 158 / 258 / 358, where the global memory device 102 / 202 is located external to chip 101 / 201. Process 700 further includes a step 706 of coupling at least one main scratch pad 152 / 252 / 352 to at least one of the plurality of processing elements and the global memory device 102 / 202. Process 700 also includes a step 708 of coupling a plurality of auxiliary scratch pads 122 / 222 / 322 to each processing element 158 / 258 / 358 and the global memory device 102 / 202. At least a portion of the plurality of auxiliary scratch pads 122 / 222 / 322 is configured as either an integrated multi-channel device, namely a multi-channel auxiliary DRAM array 120 / 220 / 320 or a chiplet 340.

[0066] Process 700 also includes a stage 712 of coupling 710 each auxiliary scratch pad 122 / 222 / 322 to each processing core 110 / 210 / 310, thereby defining a plurality of auxiliary scratch pad channels 124 / 224 / 324. In some embodiments, each auxiliary scratch pad channel 124 / 224 / 324 is further defined 712 by coupling channel controllers 154 / 254 / 354 to each auxiliary scratch pad 122 / 222 / 322. In some embodiments, a neural network accelerator chip 101 / 201 is fabricated 714. In some embodiments, a plurality of dielets 340 are fabricated 716. Both the neural network accelerator chip 101 / 201 and the dielets 340 include main scratch pads 152 / 252 / 352, processing elements 158 / 258 / 358, and auxiliary scratch pads 122 / 222 / 322.

[0067] The embodiments disclosed and described herein are configured to provide improvements in computer technology. The materials, operative structures, and techniques disclosed herein may provide substantial and beneficial technical effects. Some embodiments may not have all of these potential advantages, and these potential advantages are not necessarily required in all embodiments. By way of example only, and not limitation, one or more embodiments may provide enhanced operation of the memory system by adding dedicated auxiliary scratch pad memory to individual processor cores.

[0068] In at least some of the embodiments described herein, enhancing a memory system includes increasing bandwidth, as off-chip memory bandwidth is typically limited by each packaging feature. Further, as computing performance of recent computing systems has increased, memory bandwidth has also scaled in proportion to computing gains to achieve balanced system performance. Further, the use of off-chip memory access is relatively power consuming compared to on-chip memory access. Accordingly, the embodiments described herein, including 3D stack embodiments, facilitate reducing power consumption of an associated computer system by using the proximity of processor cores and auxiliary DRAM scratch pads on the same chip. Further, the 3D stack DRAM embodiments described herein facilitate a memory density larger than on-die SRAM. Accordingly, the embodiments described herein facilitate a much larger memory capacity in a smaller form factor, thereby facilitating a larger memory capacity. Further, with respect to form factor, the embodiments described herein facilitate reducing dependence on memory packages external to the chip, thereby reducing the need for external memory packages and, in some cases, eliminating them.

[0069] In one or more of the embodiments described herein, additional advantages are obtained when executing smaller batch cases more quickly through each computing system for operations where amortization of costs associated with the import of various weighting values is not achievable, i.e., such costs are not recoverable.

[0070] The present disclosure may be a system, method, and / or computer program product at any possible technical detail level of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present disclosure.

[0071] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other electromagnetic wave propagating freely, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0072] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0073] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0074] Computer-readable program instructions for performing the operations of this disclosure may be any combination of source code or object code written in one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or any combination of one or more programming languages such as object-oriented programming languages like Smalltalk®, C++, or the like, and procedural programming languages like the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer as a stand-alone software package, may be executed partially on the user's computer, may be executed partially on the user's computer and partially on a remote computer, or may be executed entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may utilize the state information of the computer-readable program instructions to personalize the electronic circuit in order to execute the computer-readable program instructions to implement aspects of this disclosure.

[0075] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0076] These computer-readable program instructions are provided to a computer's processor or other programmable data processing apparatus, generating a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium, and the instructions can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner. Thus, a computer-readable storage medium storing the instructions internally includes a product comprising instructions for implementing the mode of function / operation specified in one or more blocks of the flowchart and / or block diagram.

[0077] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to generate a computer-implemented process. Thus, the instructions executed on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks of the flowchart and / or block diagram.

[0078] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be performed in an order different from that noted in the figures. For example, two blocks shown in succession may in fact be implemented as one step, executed at the same time, substantially simultaneously, in a partially or wholly overlapping manner in time, or the blocks may, in some cases, be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations or by a combination of dedicated hardware and computer instructions.

[0079] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or a technical improvement over technologies found in the marketplace, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

**Claim 1** A memory system configured to expand the capacity of a plurality of main scratch pads for respective ones of a plurality of processing cores, a global memory device coupled to a plurality of processing elements, where the global memory device is disposed external to a chip on which the plurality of processing elements reside; at least one main scratch pad coupled to at least one of the plurality of processing elements and the global memory device; and a plurality of auxiliary scratch pads coupled to the plurality of processing elements and the global memory device, where at least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device, comprising a memory system. **Claim 2** The memory system of claim 1, wherein the plurality of auxiliary scratch pads are configured as dynamic random access memory (DRAM). **Claim 3** The memory system of claim 2, wherein the plurality of auxiliary scratch pads define a multi-channel auxiliary DRAM array. **Claim 4** The plurality of processing elements include a plurality of processing cores; each of the plurality of auxiliary scratch pads is coupled to each of the plurality of processing cores, thereby defining a plurality of auxiliary scratch pad channels, The memory system of claim 3. **Claim 5** The memory system of claim 4, further comprising a channel controller coupled to at least one of the plurality of auxiliary scratch pads, thereby further defining one of the plurality of auxiliary scratch pad channels. **Claim 6** Further comprising a plurality of dielets, each of the plurality of dielets including the at least one main scratch pad; the at least one processing element; and one or more of the plurality of auxiliary scratch pads, including The memory system of claim 1. **Claim 7** The memory system of claim 1, further comprising an on-chip interconnect coupled to each of the global memory device, at least one of the plurality of processing elements, the at least one main scratch pad, and the plurality of auxiliary scratch pads comprising a memory system.

8. A method of assembling a memory system configured to expand the capacity of a plurality of main scratch pads for each of a plurality of processing cores, comprising: placing a plurality of processing elements on a chip; coupling a global memory device to the plurality of processing elements, wherein the global memory device is disposed external to the chip; coupling at least one main scratch pad to at least one of the plurality of processing elements and the global memory device; and coupling a plurality of auxiliary scratch pads to the plurality of processing elements and the global memory device, wherein at least a portion of the plurality of auxiliary scratch pads is configured as an integrated multi-channel device, The method comprising.

9. The method of claim 8, further comprising configuring the plurality of auxiliary scratch pads as dynamic random access memory (DRAM).

10. The method of claim 9, further comprising assembling the plurality of auxiliary scratch pads to define a multi-channel auxiliary DRAM array.

11. coupling each of the plurality of auxiliary scratch pads to each of the plurality of processing cores, thereby defining a plurality of auxiliary scratch pad channels The method of claim 10, further comprising.

12. The method of claim 11, further comprising coupling a channel controller to at least one of the plurality of auxiliary scratch pads, thereby further defining one of the plurality of auxiliary scratch pad channels.

13. further comprising assembling a plurality of chiplets, each of the plurality of chiplets comprising: the at least one main scratch pad; the at least one processing element; and at least one of the plurality of auxiliary scratch pads The method of claim 8, including.

14. A computer system comprising: a plurality of processing devices disposed on a chip, each of the plurality of processing devices having one or more processing elements and at least one main scratch pad coupled to the one or more processing elements; and a memory system according to any one of claims 1 to 7.