Hierarchical Method and System for Storing Data
By adopting hierarchical memory arrangement and data selection devices in deep neural network training, selectively storing or discarding data instances, the problem of increasing off-chip memory transactions is solved, computing efficiency is improved and energy consumption is reduced.
Patent Information
- Application Number
- CN202080098193.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-05-06
AI Technical Summary
During deep neural network training, off-chip memory transactions increase latency, reduce computing efficiency and increase energy consumption, and the prior art is difficult to effectively reduce the number of memory transactions.
The hierarchical memory arrangement and data selection devices are adopted to determine their storage location by selecting and pruning data instances, reducing the amount of off-chip storage, including multiple processing engines, memory and data selection devices within a system on chip (SOC), and selecting data instances to be stored or discarded using threshold, random selection and sorting methods.
The total amount of memory transactions, especially off-chip memory transactions, improve computing efficiency, and reduce training time and energy consumption.
Smart Images

Figure CN115244521B_ABST
Abstract
Description
Background Art
[0001] During the training of a deep neural network that performs artificial intelligence tasks, the backpropagation (BP) algorithm is widely used. Generally, BP is used to determine and fine-tune the weights of the neurons (or nodes) of the neural network based on the error levels calculated in previous (e.g., iterative) epochs. BP is well-known in the art.
[0002] Training is performed in "mini-batches", where each mini-batch consists of dozens to hundreds of samples. Each sample passes through the neural network in the "forward" direction from the input layer through the intermediate (or hidden) layer to the output layer. Each layer can include multiple nodes that can communicate with one or more other nodes in other layers along different paths.
[0003] Each sample generates a set of signals, which are referred to herein as units or instances of data. The activation function defines the output of a node based on the inputs to the node. Each input to a node can be referred to as an activation. When training a neural network, the activations propagate forward through the network as described above, and the error signals propagate backward through the network. The forward activations can be referred to as forward activations (fw_a), and the error signals can be referred to as the backward derivatives of the activations (bw_da), both of which are also referred to herein as data instances. The weight changes (weight gradients) of the nodes are determined as a function of fw_a and bw_da.
[0004] The amount of data generated per epoch (e.g., fw_a, bw_da) and the total amount of data can be substantial and can exceed the on-chip memory capacity. Thus, some data must be stored off-chip. However, accessing data stored outside the chip is slower than accessing data stored on-chip. Therefore, off-chip memory transactions increase latency and reduce computational efficiency, increase the time required to train a neural network, and also increase the energy consumed by the computer system performing the training. Summary of the Invention
[0005] Accordingly, a system or method that generally reduces the number of memory transactions and particularly reduces the number of off-chip memory transactions is advantageous. Systems and methods are disclosed herein that provide these advantages, thereby also reducing latency, increasing computational efficiency, reducing the time spent training a neural network, and reducing energy consumption.
[0006] In various embodiments, the disclosed systems and methods are used to determine whether individual data instances (e.g., forward activations (fw_a), backward derivatives of the activations (bw_da)) are to be stored on-chip or off-chip. The disclosed systems and methods are also used to "prune" the data (discard or delete selected data instances) to generally reduce the total amount of data, particularly the amount of data stored off-chip.
[0007] In one embodiment, the system according to the present invention includes a hierarchical arrangement of various memories and also includes a hierarchical arrangement of various data selection devices for determining whether to discard data and where to discard the data in the system. In one embodiment, the system includes a System-on-Chip (SOC), the SOC including a plurality of processing engines, a plurality of memories, and a plurality of data selection devices. Each processing engine includes a processing unit, a memory (referred to herein as a first memory), and a data selection device (referred to herein as a first data selection device). The SOC also includes a plurality of other data selection devices (referred to herein as second data selection devices) coupled to the first data selection devices. The SOC also includes a number of other memories (referred to herein as second memories). In one embodiment, each second memory is coupled to a corresponding second data selection device. The SOC also includes other data selection devices (referred to herein as third data selection devices) coupled to the second data selection devices. The SOC is coupled to an off-chip memory (referred to herein as a third memory). The third memory is coupled to the third data selection devices.
[0008] In the present embodiment, each first data selection device operates to select various data instances (e.g., fw_a, bw_da) received from its corresponding processing unit. The various data instances not selected by the first data selection device are discarded. At least some of the data instances selected by the first data selection device are stored in the corresponding first (on-chip) memory. The various data instances selected by the first data selection device and not stored in the first memory are forwarded to any one of the second data selection devices.
[0009] In the present embodiment, each second data selection device operates to select various data instances received from the first data selection devices. The various data instances received by the second data selection device and not selected are discarded. At least some of the data instances selected by the second data selection device are stored in the corresponding second (on-chip) memory. The various data instances selected by the second data selection device and not stored in the second memory are forwarded to the third data selection device.
[0010] In the present embodiment, the third data selection device operates to select various data instances received from the second data selection devices. The various data instances received by the third data selection device and not selected are discarded. The various data instances selected by the third data selection device are stored in the third (off-chip) memory.
[0011] The first data selection device, the second data selection device, and the third data selection device can use various tests or methods to select (or discard) the various data instances. In one embodiment, each first data selection device compares the value of each data instance with a threshold and selects the data instances having corresponding values that satisfy the threshold. In other words, for example, if the value of a data instance is less than a minimum value, this data instance is discarded.
[0012] In one embodiment, each second data selection device randomly selects each data instance to be stored and each data instance to be discarded.
[0013] In one embodiment, the third data selection device ranks each data instance from the highest value to the lowest value and selects a certain percentage of data instances based on the ranking. In other words, for example, the third data selection device selects the top X percent of the sorted values.
[0014] Therefore, according to the embodiments of the present invention, the amount of data stored off-chip is reduced, thereby reducing the number of off-chip memory transactions. Generally, the amount of data is reduced, thereby reducing the total number of memory transactions (on-chip and off-chip) during the training of the deep neural network.
[0015] After reading the specific implementations of the embodiments shown in the various drawings, those of ordinary skill in the art will recognize the above and other objects and advantages of the various embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The various drawings are incorporated into and form a part of this specification, in which like reference numerals depict like elements, and these drawings illustrate embodiments of the present disclosure and, together with the detailed description, are used to explain the principles of the present disclosure.
[0017] Figure 1 is a block diagram showing a process for managing data during the training of a deep neural network for performing an artificial intelligence task according to an embodiment of the present invention.
[0018] Figure 2 is a block diagram showing a system for managing data and reducing memory transactions according to an embodiment of the present invention.
[0019] Figure 3 is a block diagram showing a system for managing data and reducing memory transactions according to an embodiment of the present invention.
[0020] Figure 4 is a flowchart showing the operations of an implementation of a system for managing data and reducing memory transactions according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] Reference will now be made in detail to various embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Although described in connection with these embodiments, it is to be understood that they are not intended to limit the present disclosure to these embodiments. On the contrary, the present disclosure is intended to cover alternative forms, modifications, and equivalents, which may be included within the spirit and scope of the present disclosure as defined by the appended claims. In addition, in the following detailed description of the present disclosure, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it should be understood that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the present disclosure.
[0022] Some portions of the detailed descriptions which follow are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. In the present application, a procedure, logic block, processing, etc. is conceived to be a self-consistent sequence of steps or instructions leading to a desired result. These steps are those utilizing physical operations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system. Sometimes, for the sake of common usage, it has proven convenient to refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, etc.
[0023] However, it should be borne in mind that all such and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless clearly stated otherwise explicitly from the following discussion, it should be understood that throughout the present disclosure, discussions using terms such as "receive", "send", "forward", "access", "determine", "use", "store", "select", "discard", "delete", "read", "write", "execute", "compare", "satisfy", "sort", "order", "train", etc. refer to the actions and processes of a device or a computer system or similar electronic computing device or system (e.g., Figure 2 system 200 and Figure 3 system 300). A computer system or similar electronic computing device manipulates and transforms data represented as physical (electronic) quantities in a memory, register, or other such information storage, transmission, or display device.
[0024] The various embodiments described herein can be discussed in the general context of computer-executable instructions, such as program modules, that are executed by one or more computers or other devices and reside on some form of computer-readable storage medium. By way of example, and not limitation, computer-readable storage media can include non-transitory computer storage media and communication media. In general, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the functionality of program modules can be combined or distributed as desired.
[0025] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory (e.g., SSD) or other storage technology, compact disc read only memory (CD-ROM), digital versatile disks (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed to retrieve that information.
[0026] Communication media can embody computer-executable instructions, data structures, and program modules, and includes any information delivery media. By way of example, and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. Combinations of any of the above are also included within the scope of computer-readable media.
[0027] Figure 1 is a block diagram illustrating a process 100 for managing data during the training of a deep neural network (DNN) for performing an artificial intelligence (AI) task according to an embodiment of the present invention. The DNN includes multiple layers, e.g., layer i and i + 1. Each layer includes one or more neurons or nodes. During training, individual data instances (forward activation (fw_a)) pass through the DNN in the forward direction (e.g., from layer i to layer i + 1), and individual data instances (backward activation of derivatives (bw_da)) also pass through the DNN in the backward direction (e.g., from layer i + 1 to layer i). The weight change (weight gradient (w)) at each node can be determined as a function of fw_a and bw_da, as Figure 1 shown in the upper part (above the dotted line).
[0028] In an embodiment according to the present invention, in blocks 102 and 104, decisions are made regarding whether to prune (delete, discard, not store) data instances of fw_a and / or bw_da data. For example, these decisions can be made based on the amount of data relative to the available computer system storage capacity. For example, if the available memory capacity is below a corresponding threshold, or the amount of data exceeds a corresponding threshold, the data can be pruned. If the decision is not to prune the data, the process 100 proceeds to block 110; otherwise, the process 100 proceeds to block 106 and / or block 108.
[0029] If the decision is to prune the data, the fw_a data can be pruned in block 106, and the bw_da data can be pruned in block 108. Systems and methods for pruning data are described further below. It should be noted that it can be decided to prune only the fw_a data, or only the bw_da data, or both types of data. As will be described further below, pruning the data can result in fewer memory transactions (individual writes to memory, individual reads from memory).
[0030] In block 110, the weight gradient Δw is calculated using the fw_a data and the bw_da data. The weight gradient is output in block 112.
[0031] Figure 2 is a block diagram showing a system 200 for managing data and reducing memory transactions in an embodiment according to the present invention.
[0032] In Figure 2 an embodiment, the system 200 includes a system-on-chip (SOC) 250. The SOC 250 includes a plurality of on-chip processing engine cores 202_1 to 202_N, where N can be any actual integer value (greater than or equal to 1). The processing engine cores can be collectively referred to as the processing engine cores 202 or individually referred to as the processing engine core 202. In one embodiment, each processing engine core is connected to a bus 210.
[0033] Each processing engine core 202 includes a corresponding on-core processing engine or processing unit 204_1 to 204_N. One or more processing engines can be collectively referred to as the processing unit 204 or individually referred to as the processing unit 204. For example, the processing unit 204 can be a computing unit such as a central processing unit, a vector processing unit (VPU), a matrix multiplication unit (MMU), an AI training processor, or an AI accelerator, but the present invention is not limited thereto.
[0034] Each processing engine core 202 also includes corresponding on-core memories 206_1 to 206_N. The on-core memories can be collectively referred to as the first memory 206 or individually as the first memory 206. For example, the first memory 206 can be implemented as a cache, buffer, static random access memory (SRAM), dynamic random access memory (DRAM), or high bandwidth memory (HBM), but the present invention is not limited thereto.
[0035] In an embodiment, each processing engine core 202 also includes corresponding on-core data selection devices 208_1 to 208_N. The on-core data selection devices can be collectively referred to as the first data selection device 208 or individually as the first data selection device 208. The first data selection device 208 may not be on-core but may be an independent device coupled to the processing engine core 202. Generally, each first data selection device 208 is coupled to a corresponding processing unit 204.
[0036] In various embodiments, the SOC 250 also includes multiple other on-chip data selection devices 212_1 to 212_M, where M can be any actual integer value (greater than or equal to 1). These data selection devices can be collectively referred to as the second data selection device 212 or individually as the second data selection device 212. The second data selection device 212 is coupled to the bus 210 and thus coupled to the processing engine core 202, specifically, to the first data selection device 208.
[0037] The SOC 250 also includes multiple other on-chip memories 214_1 to 214_K, where K can be any actual integer value (greater than or equal to 1). These other memories can be collectively referred to as the second memory 214 or individually as the second memory 214. Generally, transactions between the processing unit 204 and the first memory 206 are executed faster than transactions between the processing unit and the second memory 214. That is, in the hierarchy of each memory, the first memory 206 is physically closer to the processing unit 204 than the second memory 214.
[0038] In one embodiment, each second memory 214 is coupled to a corresponding second data selection device 212. That is, in one embodiment, there is a one-to-one correspondence between the second data selection device 212 and the second memory 214 (M equals K); however, the present invention is not limited thereto. For example, the second memory 214 can be implemented as SRAM or DRAM, but the present invention is not limited thereto.
[0039] In various embodiments, the SOC 250 also includes yet another on-chip data selection device, called the third data selection device 216. The third data selection device 216 is coupled to the bus 210 and thus coupled to the second data selection device 212.
[0040] The SOC 250 is coupled to an off-chip memory referred to as the third memory 230. The third memory 230 is coupled to a third data selection device 216. For example, the third memory 230 may be implemented as a DRAM, but the present invention is not limited thereto. Generally, transactions between the processing unit 204 and the second memory 214 are executed faster than transactions between the processing unit and the third memory 230, because the second memory is physically closer to the processing unit 204 and also because the third memory is off-chip.
[0041] In one embodiment, the SOC 250 includes other components or devices coupled between the third data selection device 216 and the third memory 230, including but not limited to a sparse engine 218 and a double data rate (DDR) interface 220. The sparse engine 218 can be used to determine that a small number of data instances output by the third data selection device 216 are non-zero.
[0042] The first data selection device 208, the second data selection device 212, and the third data selection device 216 can use various tests or methods to select (or discard) individual data instances. The type of data selection device can be selected according to the type of test or mechanism type performed by the device. Other factors for selecting the type of data selection device include the impact on the overall system performance, such as accuracy, latency, and speed, power consumption, and hardware overhead (e.g., the space occupied by the device). The type of data selection device can also depend on the position of the device in the data selection device hierarchy in the system 200. For example, the first data selection device 208 is higher in the hierarchy; thus, there are more first data selection devices 208. In addition, the first data selection device 208 receives and evaluates more data instances than the second data selection device 212 and the third data selector 216. Thus, the type of the first data selection device 208 has more constraints than the types of data selection devices lower in the hierarchy. For example, compared to other data selection devices lower in the hierarchy, the first data selection device 208 will advantageously occupy less space and be superior in terms of latency, speed, and power consumption. Similarly, the second data selection device 212 has more constraints than the third data selection device 216.
[0043] In various embodiments, each first data selection device 208 operates to select individual data instances received from its corresponding processing unit 204 (e.g., fw_a, bw_da). Individual data instances not selected by the first data selection device 208 are discarded. At least some of the data instances selected by the first data selection device 208 are stored in the corresponding first (on-chip) memory 206. Individual data instances selected by the first data selection device 208 and not stored in the first data memory 206 are forwarded to either of the second data selection devices 212.
[0044] In one embodiment, each first data selection device 208 compares the values (specifically, absolute values) of respective data instances with a threshold, and selects the data instances having corresponding values that satisfy the threshold. In other words, for example, if the value of a data instance is less than a minimum value, this data instance is discarded. In one embodiment, the first data selection device 208 is implemented as a comparator in hardware. The threshold used by the first data selection device 208 is programmable. For example, the threshold can be different for different epochs depending on factors such as desired accuracy, desired training time, desired computational efficiency, available on-chip memory capacity, and other factors mentioned herein.
[0045] In various embodiments, each second data selection device 212 operates to select respective data instances received from the first data selection device 208. The respective data instances received by the second data selection device 212 and not selected by the second data selection device 212 are discarded. At least some of the data instances selected by the second data selection device 212 are stored in the corresponding second (on-chip) memory 214. The respective data instances selected by the second data selection device 212 and not stored in the second memory 214 are forwarded to the third data selection device 216.
[0046] In one embodiment, each second data selection device 212 randomly selects the respective data instances to be stored and the respective data instances to be discarded. In one embodiment, the second data selection device 212 is implemented in hardware using a random number generator and a register that associates a random number (e.g., 0 or 1) with each data instance. Then, for example, the data instances associated with the value 1 are discarded. The random number generator can be programmed such that a certain percentage of the data instances are associated with one value (e.g., 1) and a certain percentage of the data instances are discarded. The percentage value used by the random number generator is programmable. For example, the percentage value can be different for different epochs depending on factors such as desired accuracy, desired training time, desired computational efficiency, available on-chip memory capacity, and other factors mentioned herein.
[0047] In various embodiments, each third data selection device 216 operates to select respective data instances received from the second data selection device 212. The respective data instances received by the third data selection device 216 and not selected by the third data selection device 216 are discarded. The respective data instances selected by the third data selection device are stored in the third (off-chip) memory 230.
[0048] In one embodiment, the third data selection device 216 sorts each data instance from the highest value to the lowest value (e.g., using the absolute value of each data instance), and selects a certain percentage of the data instances based on the ranking. In other words, for example, the third data selection device 216 selects the top X percentage of the sorted values or the top Y values (Y is an integer). In one embodiment, the third data selection device 216 is implemented in hardware as a comparator and a register. The X value or Y value used by the third data selection device 216 is programmable. For example, depending on factors such as desired accuracy, desired training time, desired computational efficiency, and other factors mentioned in the text, the X value or Y value at different times can be different.
[0049] Figure 3 is a block diagram showing a system 300 for managing data and reducing memory transactions in various embodiments of the present invention. The system 300 is similar to Figure 2 the system 200, and is presented in a specific manner of Figure 2 the system 200. The same reference numerals depict the same elements.
[0050] In Figure 3 the example, the SOC 350 is coupled to the host CPU 310, which in turn is coupled to the host memory 312 (e.g., DRAM). The SOC 350 includes a scheduler 314, a direct memory access (DMA) engine 316, and a command buffer 318.
[0051] The SOC 350 also includes processing engine cores 302 and 303, similar to Figure 2 the processing engine cores 202. In one embodiment, the processing engine core 302 is a VPU including a plurality of processing cores 304_1 to 304_N, similar to Figure 2 the first memory 206, and the plurality of processing cores 304_1 to 304_N are coupled to the corresponding buffers 306_1 to 306_N. In one embodiment, the processing engine core 303 is a MMU and includes a buffer 307 similar to the first memory 206. The processing engine cores 302 and 303 also include a first data selection device 208 and a second data selection device 212, which are coupled to the second memory 214 (e.g., SRAM). The SOC 350 also includes a third data selection device 216 coupled to the second memory 214. The third data selection device 350 is coupled to an off-chip memory (e.g., DRAM) accelerator 320, and the accelerator 320 is in turn coupled to a third (off-chip) memory 230.
[0052] System 300 operates in a similar manner to system 200, evaluating each data instance and trimming each data instance accordingly, thereby reducing the number of memory transactions overall, and in particular the number of off-chip memory transactions.
[0053] Figure 4 is a flowchart 400 showing operations in a method for managing data and reducing memory transactions in various embodiments of the present invention. Flowchart 400 can be implemented and used on Figure 2 system 200 of Figure 3 and used on system 300 of Figure 2 The flowchart 400 is described with reference to the elements of Figure 3 but such a description can be easily extended to system 300 according to the above discussion of
[0054] In Figure 4 block 402, each data instance is received from the processing unit 204 by the first data selector 208.
[0055] In block 404, the first data selector 208 selects each data instance received from the processing unit 204.
[0056] In block 406, each data instance not selected by the first data selector 208 is discarded.
[0057] In block 408, at least some of the data instances selected by the first data selector 208 are stored in the first memory 206.
[0058] In block 410, each data instance selected by the first data selector 208 and not stored in the first memory 206 is forwarded to the second data selector 212. Blocks 402, 404, 406, 408, and 410 are similarly executed for each processing engine core 202.
[0059] In block 412, the second data selector 212 selects each data instance received from the first data selector 208.
[0060] In block 414, each data instance not selected by the second data selector 212 is discarded.
[0061] In block 416, at least some of the data instances selected by the second data selector 212 are stored in the second memory 214.
[0062] In block 418, each data instance selected by the second data selector 212 and not stored in the second memory 214 is forwarded to the third data selector 216. Blocks 412, 414, 416, and 418 are similarly executed with respect to the second data selector 212 and the second memory 214.
[0063] In block 420, a third data selection device 216 selects respective data instances received from a second data selection device 212.
[0064] In block 422, respective data instances not selected by the third data selection device 216 are discarded.
[0065] In block 424, respective data instances selected by the third data selection device 216 are stored in a third memory 230.
[0066] The process parameters and the order of steps described and / or illustrated herein are given only as examples and may vary as needed. For example, although the steps shown and / or described herein may be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order shown or discussed. The various example methods described and / or illustrated herein may also omit one or more of the steps described or illustrated herein, or include additional steps other than those disclosed.
[0067] In addition, although the foregoing disclosure has set forth various embodiments using specific block diagrams, flowcharts, and examples, each block diagram component, flowchart step, operation, and / or component described and / or illustrated herein can be implemented individually and / or in various hardware, software, or firmware (or any combination thereof) configurations. In addition, any disclosure of components contained within other components should be considered as an example, as many other architectures can be implemented to achieve the same functionality.
[0068] Although various embodiments have been described and / or illustrated herein in the context of a full-featured computing system or device, one or more of these example embodiments can be distributed in various forms as a program product, regardless of the specific type of computer-readable medium used for actual execution of the distribution. The embodiments disclosed herein can also be implemented using software modules that perform certain tasks. These software modules can include scripts, batch files, or other executable files that can be stored on a computer-readable storage medium or a computing system. These software modules can configure the computing system or device to perform one or more of the example embodiments disclosed herein.
[0069] One or more software modules can be implemented in a cloud computing environment. A cloud computing environment can provide various services and applications over the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) can be accessed via a web browser or other remote interface. The various functions described herein can be provided via a remote desktop environment or any other cloud-based computing environment.
[0070] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in this disclosure is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are disclosed as example forms of implementing this disclosure.
[0071] Embodiments in accordance with the present invention are thus described. Although the present disclosure has been described in specific embodiments, it is to be understood that the present invention should not be construed as being limited by these embodiments, but rather is to be interpreted in accordance with the appended claims.
Claims
1. A system, comprising: A processing engine core, the processing engine core including a processing unit and a first memory coupled to the processing unit; A second memory, coupled to the processing unit, wherein transactions between the processing unit and the first memory are performed faster than transactions between the processing unit and the second memory; A first data selection device, coupled to the processing unit; and A second data selection device, coupled to the first data selection device; Wherein the first data selection device is operable to select each data instance received from the processing unit, wherein each data instance not selected by the first data selection device is discarded, wherein at least some of the data instances selected by the first data selection device are stored in the first memory, wherein each data instance selected by the first data selection device and not stored in the first memory is forwarded from the first data selection device to the second data selection device, wherein the second data selection device is operable to select each data instance received from the first data selection device, and wherein at least some of the data instances selected by the second data selection device are stored in the second memory.
2. The system according to claim 1, wherein The first data selection device compares the value of each data instance received from the processing unit with a threshold, and wherein each data instance selected by the first data selection device has a corresponding value that meets the threshold.
3. The system according to claim 1, further comprising: A third data selection device, coupled to the second data selection device; And A third memory, coupled to the processing unit, wherein transactions between the processing unit and the second memory are performed faster than transactions between the processing unit and the third memory; Wherein each data instance received by the first data selection device but not selected by the second data selection device is discarded, wherein each data instance selected by the second data selection device and not stored in the second memory is forwarded from the second data selection device to the third data selection device, wherein the third data selection device is operable to select each data instance received from the second data selection device, and wherein each data instance selected by the third data selection device is stored in the third memory.
4. The system according to claim 3, wherein The second data selection device associates a corresponding randomly selected value with each data instance received by the first data selection device, and wherein each data instance selected by the second data selection device is randomly selected based on the respective randomly selected values.
5. The system according to claim 3, wherein, Each data instance received by the second data selection device and not selected by the third data selection device is discarded.
6. The system according to claim 5, wherein, The third data selection device ranks the data instances received by the second data selection device from highest to lowest value, and wherein the third data selection device selects a certain percentage of the data instances based on the ranking.
7. The system according to claim 5, wherein The processing unit, the first memory, the second memory, the first data selection device, and the second data selection device are on the same chip, and wherein the third memory is outside the chip.
8. The system according to claim 1, wherein Each data instance from the processing unit is selected from a set consisting of a forward activation signal and a reverse derivative of an activation signal for determining a weight gradient of a node during training of an artificial neural network.
9. A method for managing storage transactions between a processing unit and a plurality of memories, the method comprising: At a first data selection device, receiving each data instance from the processing unit; By the first data selection device, selecting each data instance received from the processing unit; Discarding each data instance not selected by the first data selection device; Storing at least some of the data instances selected by the first data selection device in a first memory of the plurality of memories, wherein a transaction between the processing unit and the first memory is executed faster than a transaction between the processing unit and other memories in the plurality of memories; And Forwarding to a second data selection device each data instance selected by the first data selection device and not stored in the first memory.
10. The method according to claim 9, further comprising: By the first data selection device, comparing the value of each data instance received from the processing unit with a threshold, wherein the value of each data instance selected by the first data selection device has a corresponding value that satisfies the threshold.
11. The method according to claim 9, further comprising: By the second data selection device, selecting each data instance received by the first data selection device; Discarding each data instance received by the first data selection device and not selected by the second data selection device; Storing at least some of the data instances selected by the second data selection device in a second memory of the plurality of memories, wherein a transaction between the processing unit and the second memory is executed slower than a transaction between the processing unit and the first memory; Forwarding to a third data selection device each data instance selected by the second data selection device and not stored in the second memory.
12. The method according to claim 11, further comprising: By the second data selection device, associating a corresponding randomly selected value with each data instance received by the first data selection device, wherein each data instance selected by the second data selection device is randomly selected based on respective randomly selected values.
13. The method according to claim 11, further comprising: By the third data selection device, selecting each data instance received from the second data selection device; Discarding each data instance received by the second data selection device and not selected by the third data selection device; And Store each data instance selected by the third data selection device in a third memory of the plurality of memories, wherein transactions between the processing unit and the third memory are performed more slowly than transactions between the processing unit and the second memory.
14. The method according to claim 13, further comprising: Sort, by the third data selection device, the ranking of each data instance received from the second data selection device from the highest value to the lowest value, wherein the third data selection device selects a certain percentage of data instances based on the ranking.
15. The method according to claim 9, wherein, The data instances from the processing unit are selected from the set consisting of forward activation signals and reverse derivatives of activation signals for determining the weight gradients of nodes during the training of an artificial neural network.
16. A system, comprising: A system-on-chip, comprising: A plurality of processing engines, each of the plurality of processing engines including a corresponding artificial intelligence (AI) training processor, a corresponding first memory, and a corresponding first data selection device among a plurality of first data selection devices; A plurality of second data selectors, coupled to the plurality of first data selectors; A plurality of second memories, each of the plurality of second memories being coupled to a corresponding second data selection device among the plurality of second data selection devices; and A third data selection device, coupled to the second data selection devices; and A third memory, coupled to the system-on-chip; Wherein the corresponding first data selection device is operable to select each data instance received from the corresponding AI training processor, wherein each data instance not selected by the corresponding first data selection device is discarded, wherein at least some of the data instances selected by the corresponding first data selection device are stored in the corresponding first memory, and wherein each data instance selected by the corresponding first data selection device and not stored in the corresponding first memory is forwarded from the first data selection device to any one of the second data selectors; Wherein each of the plurality of second data selection devices is further operable to select each data instance received by the plurality of first data selection devices, wherein each data instance received and not selected by each second data selector is discarded, wherein at least some of the data instances selected by each second data selection device are stored in the corresponding second memory among the plurality of second memories, and wherein each data instance selected by each second data selection device and not stored in the second memory is forwarded from each second data selection device to the third data selection device; and Wherein the third data selection device is further operable to select each data instance received by each second data selection device, wherein each data instance received and not selected by the third data selection device is discarded, and wherein each data instance selected by the third data selection device is stored in the third memory.
17. The system according to claim 16, wherein, The corresponding first data selection device compares the value of each data instance received from the corresponding AI training processor with a threshold value, and among them, the value of each data instance selected by the corresponding first data selection device has a corresponding value that satisfies the threshold value.
18. The system according to claim 16, wherein, Each of the second data selection devices associates each randomly selected value with each data instance received from the first data selection device, and among them, each data instance selected by each of the second data selection devices is randomly selected based on each randomly selected value.
19. The system according to claim 16, wherein, The third data selection device sorts the received data instances from the second data selection device in descending order of value, and among them, the third data selection device selects a certain percentage of data instances based on the ranking.
20. The system according to claim 16, wherein, Each data instance is selected from a set consisting of a forward activation signal and a reverse derivative of an activation signal for determining a weight gradient of a node during training of an artificial neural network.
Citation Information
Patent Citations
Memory controller
CN110678853A
Parallel Access to Volatile Memory by Processing Device for Machine Learning
CN110888826A