A Caching Method, System, Device and Storage Medium for a Neural Network
By configuring the working mode of the cache and adopting appropriate cache mapping and data multiplexing strategies, the problem of supporting neural networks in different dimensions and sizes on the hardware platform is solved, and high concurrency and high throughput computing efficiency is achieved.
Patent Information
- Application Number
- CN202210299126.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The prior art is difficult to efficiently support neural networks of different dimensions and sizes on hardware platforms, resulting in waste of computing resources and data congestion, and the inability to achieve high concurrency and high throughput writes and outputs.
By obtaining the dimension information of the neural network, configuring the working mode of the cache, and using cache mapping and data multiplexing strategies of different dimensions, the cache mapping for neural networks in different dimensions is realized to avoid data congestion.
The write and output of caches under high concurrency and high throughput are realized, reducing hardware cache resource overhead and improving computing efficiency.
Smart Images

Figure CN114742214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network algorithms, and in particular to a cache method, system, device and storage medium for a neural network. Background Art
[0002] Neural networks of different dimensions and sizes are different. Networks of different dimensions need to allocate additional resources to calculate the dimensional differences, resulting in waste of computing resource distribution; for networks of different sizes, caches that cannot be flexibly configured cannot meet the high-performance computing requirements and become bottlenecks. The implementation and deployment of convolutional neural networks on hardware platforms are becoming increasingly different, and the hardware design lacks the flexibility to support multiple network dimensions and sizes.
[0003] In summary, the problems existing in the related technologies need to be solved urgently. Summary of the Invention
[0004] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0005] To this end, an object of an embodiment of the present invention is to provide a cache method, system, device and medium for a neural network, which can enable the cache to write and output with high concurrency and high throughput.
[0006] To achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:
[0007] On the one hand, an embodiment of the present invention provides a cache method for a neural network, including the following steps:
[0008] Obtain configuration information of the cache, where the configuration information includes dimensional information of the neural network to be processed;
[0009] Set the working mode of the cache according to the dimensional information;
[0010] Obtain target data to be processed through the configured cache;
[0011] Process the target data to be processed through the configured cache according to the configuration information.
[0012] Further, the configuration information includes dimensional information, computing size information and computing step information.
[0013] Further, the step of setting the working mode of the cache includes:
[0014] Obtain the dimensional information from the configuration information;
[0015] Set the cache mapping scheme of the cache according to the dimension information.
[0016] Furthermore, the cache mapping scheme includes a one-dimensional cache mapping scheme, a two-dimensional cache mapping scheme, and a three-dimensional cache mapping scheme.
[0017] Furthermore, the cache that has been set to obtain the target data to be processed specifically includes the following steps:
[0018] Obtain the target data to be processed from the data to be processed according to the configuration information;
[0019] Write the target data to be processed into the cache.
[0020] Furthermore, the cache that has been set to process the target data to be processed specifically includes the following steps:
[0021] Determine the corresponding data reuse strategy according to the calculation size information and the calculation step information;
[0022] Process the target data to be processed according to the data reuse strategy.
[0023] Furthermore, the data reuse strategy includes a one-dimensional data reuse strategy, a two-dimensional data reuse strategy, and a three-dimensional data reuse strategy.
[0024] On the other hand, an embodiment of the present invention provides a cache system for a neural network, including:
[0025] A first module for obtaining the configuration information of the cache;
[0026] A second module for setting the working mode of the cache according to the configuration information;
[0027] A third module for obtaining the target data to be processed through the cache that has been set;
[0028] A fourth module for processing the target data to be processed through the cache that has been set according to the configuration information.
[0029] On the other hand, an embodiment of the present invention provides a cache device for a neural network, including:
[0030] At least one processor;
[0031] At least one memory for storing at least one program;
[0032] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the cache method of the neural network described above.
[0033] On the other hand, an embodiment of the present invention provides a storage medium storing instructions executable by a processor, and the instructions executable by the processor are used to implement the cache method of the neural network when executed by the processor.
[0034] The present invention discloses a cache method for a neural network, having the following beneficial effects:
[0035] In this embodiment, configuration information including neural network dimensions is obtained; according to the configuration information, the working mode of the cache is set; the target data to be processed is obtained through the configured cache; according to the configuration information, the target data to be processed is processed by the configured cache. By configuring the cache, the cache determines the target data to be processed, and then according to the configuration information, different data processing schemes are adopted for different data to be processed, realizing cache mapping for neural networks of different dimensions, so that the cache can avoid congestion, and thus write and output with high concurrency and high throughput. Moreover, efficient mapping can be achieved for a unified fixed computing array, thereby improving computing efficiency. At the same time, supporting different computing sizes of convolutional neural networks can effectively reduce the redundant practice of padding zeros for storage in the cache. Finally, different reuse strategies are selected for data of different dimensions and different sizes, reducing the amount of data accessed and stored in the cache, and reducing the hardware cache resource overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings introduced below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.
[0037] Figure 1 It is a schematic flowchart of a cache method for a neural network provided by an embodiment of the present invention;
[0038] Figure 2 It is a schematic structural diagram of a cache system for a neural network provided by an embodiment of the present invention;
[0039] Figure 3 It is a schematic structural diagram of a cache device for a neural network provided by an embodiment of the present invention;
[0040] Figure 4 A direct mapping schematic diagram provided by an embodiment of the present invention;
[0041] Figure 5 A fully associative mapping schematic diagram provided by an embodiment of the present invention;
[0042] Figure 6 A set associative mapping schematic diagram provided by an embodiment of the present invention;
[0043] Figure 7 A schematic diagram of three-dimensional data in data reuse provided by an embodiment of the present invention;
[0044] Figure 8 A schematic diagram of two-dimensional data in data reuse provided by an embodiment of the present invention;
[0045] Figure 9 A schematic diagram of one-dimensional data in data reuse provided by an embodiment of the present invention. Detailed implementation manners
[0046] This part will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the drawings. The role of the drawings is to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it cannot be understood as a limitation on the protection scope of the present invention.
[0047] In the description of the embodiments of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the number itself, above, below, within, etc. are understood as including the number itself, "at least one" means one or more, and "at least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. If there is a description of "first", "second", etc., it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0048] It should be noted that the terms such as setting, installing, and connecting in the embodiments of the present invention should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above terms in the embodiments of the present invention in combination with the specific content of the technical solution. For example, the term "connection" can be a mechanical connection, an electrical connection, or can communicate with each other; it can be directly connected or indirectly connected through an intermediate medium.
[0049] In the description of the embodiments of the present invention, the descriptions referring to terms such as "one embodiment / implementation", "another embodiment / implementation", "certain embodiments / implementations", "in the above embodiments / implementations", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least two embodiments or implementations of the present disclosure. In the present disclosure, the schematic expressions of the above terms do not necessarily refer to the same exemplary embodiments or implementations. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or implementations.
[0050] It should be noted that the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0051] A neural network is an algorithmic mathematical model that mimics the behavioral characteristics of an animal neural network and performs distributed parallel information processing. This network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information. This mathematical model is composed of a large number of directly interconnected nodes (or neurons); each node (except for the input nodes) represents a specific output function (or is considered an operation), called an activation function; the connection between every two nodes represents the proportion of the signal in the transmission (i.e., the proportion of the "memory value" of the node being passed on), called a weight; the output of the network varies due to different activation functions and weights, and is an approximation of a certain function or an approximate description of the mapping relationship.
[0052] There are differences in neural networks of different dimensions and sizes. Networks of different dimensions require additional resources to calculate the dimensional differences, resulting in a waste of the distribution of computing resources; for networks of different sizes, the inflexible cache cannot meet the high-performance computing requirements and becomes a bottleneck. The implementation and deployment of convolutional neural networks on hardware platforms are becoming increasingly different. In related technologies, the hardware design lacks the flexibility to support multiple network dimensions and sizes, resulting in data congestion, inability to work with high concurrency, and low computing efficiency.
[0053] To this end, the present application proposes a cache method, system, device and storage medium for a neural network. The method includes obtaining configuration information including the dimensions of the neural network; setting the working mode of the cache according to the configuration information; obtaining target data to be processed through the configured cache; and processing the target data to be processed through the configured cache according to the configuration information. By configuring the cache, the cache determines the target data to be processed, and then according to the configuration information, different data processing schemes are adopted for different data to be processed, realizing cache mapping for neural networks of different dimensions, so that the cache can avoid congestion and perform high-concurrency and high-throughput writing and output. The present invention can be widely applied in the technical field of neural network algorithms.
[0054] Figure 1 It is a flowchart of a cache method for a neural network provided by an embodiment of the present application. Referring to Figure 1 , the cache method for the neural network includes but is not limited to steps S110 to S140.
[0055] S110. Obtain the configuration information of the cache, where the configuration information includes the dimension information of the neural network to be processed.
[0056] In this step, in order to set the working mode of the cache, it is necessary to obtain the configuration information of the cache, and the configuration information includes the dimension information of the neural network to be processed. It can be understood that the dimension information of the neural network includes one-dimensional dimension information, two-dimensional dimension information and three-dimensional dimension information, and different cache configurations are implemented for neural networks of different dimensions. Among them, the dimension information is obtained externally and is determined by the neural network to be processed. Exemplarily, if the neural network to be processed is a one-dimensional neural network, the dimension information includes one-dimensional dimension information; if the neural network to be processed is a two-dimensional neural network, the dimension information includes two-dimensional dimension information.
[0057] Among them, the cache refers to a cache memory (Cache), whose original meaning is a kind of RAM with a faster access speed than ordinary random access memory (RAM). Generally speaking, it does not use DRAM technology like the system main memory, but uses expensive but faster SRAM technology, and also has the name of cache memory.
[0058] The cache memory is a primary memory between the main memory and the CPU, composed of static memory chips (SRAM). It has a relatively small capacity but a much higher speed than the main memory, approaching the speed of the CPU. In the hierarchical structure of the computer storage system, it is a high-speed small-capacity memory between the central processing unit and the main memory. It forms a primary-level memory together with the main memory. The scheduling and transfer of information between the cache memory and the main memory are automatically carried out by hardware.
[0059] S120. Set the working mode of the cache according to the dimension information.
[0060] In this step, set the working mode of the cache according to the obtained dimension information. Specifically, according to different dimension configuration information, corresponding cache mapping schemes are adopted for one-dimensional, two-dimensional, and three-dimensional data respectively. The cache structure includes 4 group caches, and each group cache includes 4 slice caches. Different cache mappings are required in the convolutional neural network calculations of different dimensions (one-dimensional, two-dimensional, three-dimensional) to achieve high-parallelism and high-throughput outputs.
[0061] S130. Obtain the target data to be processed through the configured cache.
[0062] In this step, after the configuration of the cache is completed, the cache can determine and read the data mode to be calculated. The cache can select the data addresses of different cache groups and cache slices to obtain the corresponding-sized data for calculation. When reading out the data, according to the configured output calculation size, the continuously distributed data is read and output.
[0063] S140. Process the target data to be processed through the configured cache according to the configuration information.
[0064] In this step, the cache selects the corresponding data readout multiplexing strategy according to the calculation size and calculation step of the convolutional neural network. Because during the convolutional neural network calculation process, overlapping data will be generated when cutting the activation features, that is, a certain data may exist in two different convolutional processes at the same time. In the related art, such the same data will be read and stored multiple times, resulting in waste of resources. However, in this application, by controlling the cache address pointer, it is possible to store the overlapping data once and read and use it multiple times, thereby achieving data multiplexing and reducing the amount of stored data.
[0065] A cache method, system, device and storage medium for a neural network proposed in this application obtain configuration information including neural network dimensions; set the working mode of the cache according to the configuration information; obtain target data to be processed through the configured cache; and process the target data to be processed through the configured cache according to the configuration information. By configuring the cache, the cache determines the target data to be processed, and then according to the configuration information, different data processing schemes are adopted for different data to be processed, realizing cache mapping for neural networks of different dimensions, so that the cache can avoid congestion, and thus write and output with high concurrency and high throughput. Moreover, efficient mapping can be achieved for a unified fixed computing array, thereby improving computing efficiency. At the same time, supporting different computing sizes of convolutional neural networks can effectively reduce the redundant practice of padding zeros for storage in the cache. Finally, different reuse strategies are selected for data of different dimensions and different sizes, reducing the amount of data accessed and stored in the cache, and reducing the hardware cache resource overhead.
[0066] Further as an optional implementation manner, the configuration information includes dimension information, computing size information, and computing step information.
[0067] Further as an optional implementation manner, the step of setting the working mode of the cache includes:
[0068] Obtain the dimension information from the configuration information;
[0069] Set the cache mapping scheme of the cache according to the dimension information.
[0070] Further as an optional implementation manner, the cache mapping scheme includes a one-dimensional cache mapping scheme, a two-dimensional cache mapping scheme, and a three-dimensional cache mapping scheme.
[0071] Specifically, there are three mapping methods for the cache memory, namely direct mapping, fully associative mapping, and set associative mapping
[0072] Direct mapping, the Cache organization of direct mapping is as Figure 4 shown. A block in the main memory can only be mapped to a specific block in the Cache. For example, block 0, block 16,... of the main memory, block 2032 can only be mapped to block 0 of the Cache; and block 1, block 17,... of the main memory, block 2033 can only be mapped to block 1 of the Cache...
[0073] Direct mapping is the simplest address mapping method. Its hardware is simple, with low cost, fast address transformation speed, and no issues related to replacement algorithms. However, this method is not flexible enough. The storage space of the Cache cannot be fully utilized. Each main memory block has only one fixed position for storage, which easily leads to conflicts, reducing the efficiency of the Cache. Therefore, it is only suitable for use with large-capacity Caches. For example, if a program needs to repeatedly reference block 0 and block 16 in the main memory, it is best to copy both main memory block 0 and block 16 to the Cache simultaneously. But since they can only be copied to block 0 of the Cache, even if other storage spaces in the Cache are empty, they cannot be occupied. As a result, these two blocks will be continuously loaded into the Cache alternately, leading to a decrease in the hit rate.
[0074] The fully associative mapping method, Figure 5 is a Cache organization with fully associative mapping. Any block in the main memory can be mapped to any position in the Cache.
[0075] The fully associative mapping method is relatively flexible. Each block in the main memory can be mapped to any block in the Cache. The utilization rate of the Cache is high, and the probability of block conflicts is low. As long as a certain block in the Cache is replaced, any block in the main memory can be loaded. However, due to the difficulty in the design and implementation of the Cache comparison circuit, this method is only suitable for use with small-capacity Caches.
[0076] Set associative mapping. Set associative mapping is actually a compromise between direct mapping and fully associative mapping. Its organizational structure is as Figure 6 shown. Both the main memory and the Cache are grouped. The number of blocks in a group in the main memory is the same as the number of groups in the Cache. Direct mapping is used between groups, and fully associative mapping is used within groups. That is to say, the Cache is divided into u groups, with v blocks in each group. Which group the main memory block is stored in is fixed, while which block in that group it is stored in is flexible. For example, the main memory is divided into 256 groups, with 8 blocks in each group, and the Cache is divided into 8 groups, with 2 blocks in each group.
[0077] Further as an optional implementation method, obtaining the target data to be processed by the configured cache specifically includes the following steps:
[0078] Obtain the target data to be processed from the data to be processed according to the configuration information;
[0079] Write the target data to be processed into the cache.
[0080] In this step, different from the configuration information which is obtained externally and used to set the working state of the cache, the data to be processed is the information that the cache needs to calculate and process. After configuring the calculation size of the cache, the cache can determine the mode of reading the data that needs to be calculated. The cache can select the data addresses of different cache groups and cache slices to obtain the corresponding size of data for calculation. When reading out the data, the continuously distributed data is read and output according to the configured output calculation size.
[0081] Further as an optional implementation manner, the cache that has been set processes the target data to be processed, specifically including the following steps:
[0082] Determine the corresponding data reuse strategy according to the calculation size information and the calculation step information;
[0083] Process the target data to be processed according to the data reuse strategy.
[0084] Further as an optional implementation manner, the data reuse strategy includes a one-dimensional data reuse strategy, a two-dimensional data reuse strategy, and a three-dimensional data reuse strategy.
[0085] Because during the calculation process of the convolutional neural network, overlapping data will be generated when cutting the activation features, that is, a certain data may exist in two different convolutional processes at the same time. In the related art, such the same data will be read and stored multiple times, resulting in waste of resources. However, in this application, by controlling the cache address pointer, it is possible to store the overlapping data once and read and use it multiple times, thereby achieving data reuse and reducing the amount of stored data.
[0086] One-dimensional data reuse only exists in the same cache slice, corresponding to the one-dimensional row (W) dimension, that is, the column data update in convolutional calculation; two-dimensional data reuse not only exists in the same cache slice, but also exists between different groups, corresponding to the two-dimensional row and column (WxH) dimension. When a line break update of the activation data occurs during the two-dimensional convolution process, only need to adjust the pointer of the previously used cache group to zero and add the data reading of the new cache group, then the updated data can be formed, reusing the data of the previous two cache groups and reducing the number of caches; three-dimensional data reuse exists simultaneously in the same slice, different groups, and different slices, corresponding to the three-dimensional row, column, and frame (WxHxF) dimension. The reuse situations also exist in the column update, row update, and frame update processes of activation respectively. By controlling the cache slice address pointer and the group address pointer, the pointers of the data that needs to be reused are maintained to achieve the function of retaining the overlapping calculation data.
[0087] The cache structure contains 4 set caches, and each set cache contains 4 slice caches. Different cache mappings are required for convolutional neural network calculations in different dimensions (one-dimensional, two-dimensional, three-dimensional) to achieve high parallelism and high throughput output.
[0088] In addition to the channel (C), the three-dimensional data has three dimensions: row, column, and frame (HxWxF). Its cache mapping stores data in units of 4 frames (F) into the 4 slice caches of a set. The data in the same row is mapped to consecutive addresses in the same slice of the cache. The data in the next column is stored in the cache of another set to output multiple frames, multiple rows, and multiple columns of data simultaneously for three-dimensional convolution calculation;
[0089] Exemplarily, Table 1 shows the transformation and control of the address pointer when three-dimensional data is reused. Refer to Figure 7 , if there is an overlap of two frames in the calculated data, the cache only needs to store the overlapping data of these two frames once, and then by repeating the read of the address pointer at the overlapping position twice, data reuse can be achieved, so as to obtain two data blocks with overlapping data for calculation, rather than storing all the data of the two data blocks.
[0090] Example of frame data update when reading 3×3×3 size data:
[0091]
[0092] Table 1
[0093] In addition to the channel (C), the two-dimensional data has two dimensions: row and column (HxW). The two-dimensional data can be regarded as a special case where the "frame" dimension of the three-dimensional data = 1. Its cache mapping caches data in units of 4 channels (C) in the 4 slice caches of the same set. The data in the same row is mapped to consecutive addresses in the same slice of the cache. The data in the next column is stored in the cache of another set to output multiple channels, multiple rows, and multiple columns of data simultaneously for two-dimensional convolution calculation;
[0094] Exemplarily, Table 2 shows the transformation and control of the address pointer when two-dimensional data is reused. Refer to Figure 8 , if there is an overlap of two rows in the calculated data, the cache only needs to store the overlapping data of these two rows once, and then by repeating the read of the address pointer at the overlapping position twice, data reuse can be achieved, so as to obtain two data blocks with overlapping data for calculation, rather than storing all the data of the two data blocks.
[0095] Example of row data update when reading 3×3×3 size data:
[0096]
[0097] Table 2
[0098] One-dimensional data has only the row (W) dimension in addition to the channel (C). The channel data is divided into channel (vertical) and channel (horizontal) according to the arrangement direction. Four channels (horizontal) are used as the basic unit and cached in four slice caches of the same group. The data of the same row is mapped to the consecutive addresses of the same slice in the cache. The next channel (vertical data) is stored in the cache of another group. The mapping methods of different dimensions can effectively reduce the vacancy of computing units for a fixed and unified computing circuit.
[0099] Exemplarily, Table 3 shows the transformation and control of the address pointer when one-dimensional data is reused. Refer to Figure 9 , if there is an overlap of two columns in the data to be calculated, the cache only needs to store the overlapping data of these two columns once, and then by repeatedly reading the address pointer at the overlapping position twice, data reuse can be achieved, so as to obtain two data blocks with data overlap for calculation, rather than storing all the data of the two data blocks.
[0100] Example of column data update when reading 3×3×3 size data:
[0101]
[0102] Table 3
[0103] Refer to Figure 2 , a cache system for a neural network proposed in an embodiment of the present invention includes:
[0104] A first module 210, configured to obtain configuration information of the cache;
[0105] A second module 220, configured to set the working mode of the cache according to the configuration information;
[0106] A third module 230, configured to obtain target data to be processed through the configured cache;
[0107] A fourth module 240, configured to process the target data to be processed through the configured cache according to the configuration information.
[0108] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0109] Refer to Figure 3 , an embodiment of the present invention provides a cache device for a neural network, including:
[0110] At least one processor 310;
[0111] At least one memory 320 for storing at least one program;
[0112] When the at least one program is executed by the at least one processor 310, the at least one processor 310 is caused to implement Figure 1 The cache method of the neural network shown.
[0113] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0114] The embodiments of the present invention further provide a storage medium, in which there are instructions executable by a processor, and the instructions executable by the processor are used to implement Figure 1 The cache method of the neural network shown.
[0115] It can be understood that, compared with the prior art, the embodiments of the present invention further have the following advantages:
[0116] 1) The cache mapping for neural networks of different dimensions is realized, so that the cache can avoid congestion, and thus write and output with high concurrency and high throughput.
[0117] 2) Efficient mapping can be achieved for a unified fixed computing array, thereby improving the computing efficiency.
[0118] 3) Supporting the computing sizes of different convolutional neural networks can effectively reduce the redundant practice of zero-padding storage of data in the cache.
[0119] 4) Selecting different reuse strategies for data of different dimensions and different sizes can reduce the amount of data accessed and stored in the cache, and can reduce the hardware cache resource overhead.
[0120] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A caching method for a neural network, characterized in that, it includes the following steps: Obtain the configuration information of the cache, and the configuration information includes the dimension information of the neural network to be processed; Set the working mode of the cache according to the dimension information; Obtain the target data to be processed through the configured cache; Process the target data to be processed through the configured cache according to the configuration information; The configuration information includes dimension information, calculation size information, and calculation step information; The processing of the target data to be processed through the configured cache specifically includes the following steps: Determine the corresponding data reuse strategy according to the calculation size information and the calculation step information; Process the target data to be processed according to the data reuse strategy; The data reuse strategy is used to store the overlapping data generated when slicing the activation features once and read and use it multiple times by controlling the cache address pointer.
2. The caching method for a neural network according to claim 1, characterized in that, The step of setting the working mode of the cache includes: Obtain the dimension information from the configuration information; Set the cache mapping scheme of the cache according to the dimension information.
3. The caching method for a neural network according to claim 2, characterized in that, The cache mapping scheme includes a one-dimensional cache mapping scheme, a two-dimensional cache mapping scheme, and a three-dimensional cache mapping scheme.
4. The caching method for a neural network according to claim 1, characterized in that, The obtaining of the target data to be processed through the configured cache specifically includes the following steps: Obtain the target data to be processed from the data to be processed according to the configuration information; Write the target data to be processed into the cache.
5. The caching method for a neural network according to claim 1, characterized in that, The data reuse strategy includes a one-dimensional data reuse strategy, a two-dimensional data reuse strategy, and a three-dimensional data reuse strategy.
6. A caching system for a neural network, characterized in that, it includes: A first module for obtaining the configuration information of the cache; A second module for setting the working mode of the cache according to the configuration information; A third module for obtaining the target data to be processed through the configured cache; A fourth module for processing the target data to be processed through the configured cache according to the configuration information; The configuration information includes dimension information, calculation size information, and calculation step information; The processing of the target data to be processed through the configured cache specifically includes the following steps: Determine the corresponding data reuse strategy according to the calculation size information and the calculation step information; Process the target data to be processed according to the data reuse strategy; The data reuse strategy is used to achieve one-time storage and multiple read uses of the overlapping data generated when slicing the activation features by controlling the cache address pointer.
7. A cache device for a neural network, characterized in that, it includes: at least one processor; at least one memory for storing at least one program; when the at least one program is executed by the at least one processor, the at least one processor implements the cache method for the neural network according to any one of claims 1-5.
8. A computer-readable storage medium storing processor-executable instructions, characterized in that, the processor-executable instructions are used to implement the cache method for the neural network according to any one of claims 1-5 when executed by a processor.
Citation Information
Patent Citations
Computer cache memory performance enhance using neural networks and machine learning
IN202041001789A